Skip to content

Opencode deepseek benchmark - #70

Closed
frankreyesgarcia wants to merge 30 commits into
chains-project:mainfrom
frankreyesgarcia:opencode-deepseek-benchmark
Closed

frankreyesgarcia wants to merge 30 commits into
chains-project:mainfrom
frankreyesgarcia:opencode-deepseek-benchmark

Conversation

@frankreyesgarcia

@frankreyesgarcia frankreyesgarcia commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Summary

yul against the same 60 dependency-addition tasks the paper's Claude Sonnet 5 sweep uses (benchmark/cases_top.json, 10 per ecosystem across Cargo, GitHub Actions, Go, Maven, npm, PyPI), driven through OpenCode against DeepSeek's hosted API (deepseek/deepseek-flash, model name DeepSeek V4.1 Flash) instead of claude -p against Claude — same repetition scheme, 60 cases × 2 conditions × 3 reps, 360 runs total.

Sandboxed via Docker (CONTAINER_RUNTIME=apptainer|docker auto-detects; Apptainer only runs on Linux and shares the host kernel, which a Docker host can't guarantee). Full details, methodology, and manual-review reasoning are in benchmark/runs-opencode-deepseek-docker/README.md.

360 runs, 0 failures, real cost $2.9045.

Table (same format as the paper's Already-latest/mitigation table)

Claude Sonnet 5 (bench/top-final-results, for reference):

Ecosystem Tasks (nohook) Already latest (nohook) Tasks (hook) Already latest (hook) Mitigated Rate
Cargo 29 29/29 29 24/29 5/5 100%
GitHub Actions 30 3/30 30 3/30 27/27 100%
Go 30 30/30 29 28/29 1/1 100%
Maven 30 19/30 30 16/30 14/14 100%
npm 26 23/26 27 18/27 9/9 100%
PyPI 30 23/30 30 14/30 16/16 100%
All 175 127/175 175 103/175 72/72 100%

DeepSeek V4.1 Flash (this PR):

Ecosystem Tasks (nohook) Already latest (nohook) Tasks (hook) Already latest (hook) Mitigated Rate
Cargo 25 25/25 27 27/27 0/0 —
GitHub Actions 29 22/29 30 4/30 22/26 85%
Go 23 22/23 23 21/23 2/2 100%
Maven 30 24/30 30 18/30 11/12 92%
npm 8 7/8 5 4/5 0/1 0%
PyPI 21 21/21 21 20/21 1/1 100%
All 136 121/136 136 94/136 36/42 86%

DeepSeek excludes/never-attempts far more reps than Claude (88/360, 24%, vs 5/180, 3%) — concentrated in npm (self-implements small single-purpose packages instead of depending on them) and PyPI's six/pandas cases (abandoned, or written to a file yul doesn't watch). Rate isn't uniformly 100% like Claude's row: GitHub Actions (85%) is the biggest real gap, npm has one unmitigated stale candidate (the model deleted the dependency instead of fixing its version), everything else is 92-100% or has no stale candidates left. Full manual-review breakdown and methodology caveats (ground truth via yul scan itself, an independent ecosyste.ms cross-check that only partially worked) are in the README linked above.

A credential leak was found and fixed while building this dataset (nohook runs had no protection against a model reading OpenCode's credential store) — documented in the README, redacted before ever reaching this branch's pushed history, all affected credentials rotated.

Test plan

  • All 360 runs completed with 0 failures (real API cost $2.9045)
  • Ground-truth analysis (benchmark/analyze_top_opencode.py) run against the full dataset via yul scan
  • Manual review of every rep where the target package was absent from the final manifest (88 of 360), reading generated source and full transcripts, not just final state
  • go build ./... and go test ./... pass after merging main
  • Full branch history scanned for leaked credentials (clean) before push

frankreyesgarcia and others added 25 commits September 13, 2026 23:06
Adds two pieces of harness support alongside the existing Claude Code plugin,
both validated live against real installs rather than just against docs.

Qwen Code (qwen-extension.json):
- Qwen Code's extension loader reads hooks/hooks.json via the same path,
  JSON shape, and ${CLAUDE_PLUGIN_ROOT} substitution Claude Code uses, and
  `qwen extensions install <marketplace>:<plugin>` already installs straight
  from this repo's existing chains-hooks marketplace listing.
- Found and fixed a real bug: Qwen Code's hook executor treats the "timeout"
  field as milliseconds, not seconds like Claude Code. hooks.json's
  "timeout": 60 (meant as 60s) became a 60ms timeout, so ensure-yul.sh's
  binary download failed instantly every session (silently, since it fails
  open). qwen-extension.json ships the same hooks with millisecond-scaled
  timeouts; Qwen Code prefers an extension's own manifest over the
  hooks/hooks.json fallback, so this fixes the bug without touching main.go.
- Verified end-to-end: installed from this repo's local path into a
  sandboxed Qwen Code config, confirmed both SessionStart hooks now succeed,
  confirmed the real release binary downloads and caches correctly, and fed
  the cached binary a synthetic PreToolUse payload in Qwen Code's exact
  documented shape (tool_name/tool_input.file_path/content) - it blocked an
  outdated pin with exit 2 as expected. No Go code changes were needed.

OpenCode (benchmark/run_case_opencode.sh):
- Parallel to benchmark/run_case.sh, but drives `opencode run` against a
  custom OpenAI-compatible provider (for a self-hosted vLLM endpoint) instead
  of `claude -p`, and wires up yul via a small JS plugin
  (.opencode/plugins/yul.js) implementing OpenCode's tool.execute.before
  hook, translating its write/edit args into yul's existing stdin JSON shape
  and shelling out to the same compiled binary.
- Live-tested against a real Qwen3.6-35B-A3B vLLM server: the block mechanism
  works end-to-end (plugin invokes yul, exit 2 becomes a thrown Error,
  OpenCode surfaces the exact reason as the tool result).

A follow-up change hardens this against a bash-based bypass found during
that testing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T2ZAeghvWbbLnsHQ9c9vA4
The OpenCode yul plugin only intercepted tool.execute.before for write/edit,
so a model blocked on an outdated pin could just write the same content via
the bash tool instead - completely skipping the check. Live-tested this
against a real Qwen3.6-35B-A3B model: it blocked cleanly on `edit` (yul
correctly flagged pyyaml 5.3.1 -> 6.0.3), then wrote the file via
`printf ... > requirements.txt` with the same outdated pin instead of
retrying with the suggested version.

This is a limitation inherited from yul's own PreToolUse matcher, which is
Write|Edit only in Claude Code and Qwen Code too - just newly visible here
since OpenCode's plugin API exposes bash as a distinctly interceptable tool.

Fix: the plugin now also intercepts bash calls, and blocks any command that
both mentions a known manifest filename (pom.xml, requirements.txt,
pyproject.toml, package.json, go.mod, Cargo.toml, .github/workflows/*.yml)
and contains a file-mutating construct (>, >>, tee, sed -i, perl -i, dd of=,
cp, mv), rather than trying to parse shell semantics. Validated the
detection regex against 14 synthetic commands (real bypass patterns and
plausible false positives like `cat requirements.txt`, `git diff pom.xml`,
`pip install -r requirements.txt`) - all matched as expected.

Re-ran the 10-case/20-run benchmark across all 6 ecosystems after this fix:
no bash bypass observed. Two cases showed the hook blocking an outdated pin
and the model self-correcting on retry through the proper channel (maven
lombok 1.18.34->1.18.48; a GitHub Actions checkout upgraded from @v4 to the
exact SHA for v7.0.1). One case (maven junit) showed the hook blocking
correctly twice but the model retrying identical rejected content instead of
using the suggested versions - a model self-correction gap, not a hook bug.
Remaining timeouts/no-ops in that run were unrelated to yul (no Rust/gcc
toolchain reachable in the sandbox, Node not on PATH, or the model choosing
to inline a feature instead of adding the dependency).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T2ZAeghvWbbLnsHQ9c9vA4
Brings in the merged Qwen Code/OpenCode PR (chains-project#46) plus other main
changes since this branch forked, while keeping the bash-bypass fix.

# Conflicts:
#	benchmark/run_case_opencode.sh
Runs benchmark/cases_top.json under both hook and nohook via OpenCode
against a self-hosted Qwen3.6-35B-A3B-FP8 vLLM server, on the
bash-bypass-fixed harness (f0cddd7). 21/60 cases triggered a real yul
block; 0 of 7 bash-bypass attempts succeeded (vs. 4/16 on a prior run
without the fix). See benchmark/runs-opencode-qwen-60/SUMMARY.md for
the full breakdown.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
transcript.jsonl/stderr.log were written via relative redirection after
cd "$WORKDIR" - the same directory the model operates in - so the model
could see and mutate its own run's live log files. Observed in practice:
go-top-02-go-difflib/hook ended up with no transcript at all, apparently
deleted mid-run by the model treating it as a build artifact. Now
captured to a tempfile outside WORKDIR and moved into place only after
the run finishes.

Re-ran go-top-02-go-difflib/hook with the fix; no block triggered
(go-difflib v1.0.0 is current), so this doesn't change any aggregate
numbers in SUMMARY.md.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The 60-case/120-run sweep and its SUMMARY.md now live on
opencode-benchmark-qwen60, keeping this branch scoped to the actual
bash-bypass and logging fixes (f0cddd7, 58eddf3) rather than mixing in
a large data dump.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
yul's PreToolUse hook only ever saw Write and Edit tool calls, so a
model blocked on an outdated pin could rewrite the same manifest via
bash (cat > file, sed -i, etc.) and skip the check entirely - the same
gap the OpenCode benchmark plugin already worked around (f0cddd7), but
inherited by Claude Code and Qwen Code too since they only wire yul's
hook to Write|Edit.

Adds a Bash branch to runHook: since yul can't simulate arbitrary
shell to know what content would land, it's a coarse block (any
command that both names a known manifest and contains a
content-mutating construct) rather than a real version check. Ported
the heuristic from the OpenCode plugin's JS regexes to Go, fixing a
boundary bug found in the process (dd of=go.mod wasn't matching
because "=" wasn't a recognized separator) and applying the same fix
back to the JS version for parity.

Updates hooks/hooks.json's PreToolUse matcher to Write|Edit|Bash, and
qwen-extension.json's matcher to match (kept untracked, as before).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Qwen Code support isn't something we're pursuing right now; drop the
extension manifest rather than carrying unused, unmaintained config.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Aggregates each run's step_finish events into a single summary file
alongside transcript.jsonl, so token/cost accounting doesn't require
re-parsing the full transcript every time. cost_usd is whatever
OpenCode's provider pricing table reports - 0 for the self-hosted
"local" provider used so far, real dollars for a priced provider like
DeepSeek.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
transcript.jsonl never records which provider/model served a run -
OpenCode's own runtime log does, per sessionID. Grep it (filtered by
this run's sessionID, so concurrent runs don't cross-contaminate) and
save the matching provider/model lines to model_used.log alongside
transcript.jsonl/usage.json, instead of having to dig through the
global log by hand to verify which model actually ran.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Mirrors benchmark/README.md's format (intro, summary table, narrative
findings) for the 200-run (10 cases x 2 conditions x 10 reps) pilot
against DeepSeek's hosted API instead of a single-run sweep.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Replaces the ad-hoc summary table with the same column set as
benchmark/README.md's top-60 table (Tasks/Versioned/Already latest/
Blocked/Rate), one row per case since we only have 10. "Already
latest" is now ground truth - every final_manifest replayed through
the real yul binary (before="") against live registries, not a regex
guess. Also fixes a narrative bug: the one urllib3 block was actually
on an unrelated pytest pin in a pyproject.toml the model also wrote,
not on urllib3 itself.

Adds a collapsible 200-row per-run detail table (case, condition,
rep, versioned, already latest, blocked, flagged dependency).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… plugin

main.go's runHook now handles tool_name "Bash" itself (7ca8985), so
the OpenCode plugin no longer needs its own copy of the
manifest/write-construct regexes - that was duplicate logic to keep
in sync by hand (already caused one bug: a boundary fix landed in Go
first and had to be remembered for the JS copy too). The plugin is
now a thin translation layer for all three tool kinds (write/edit/
bash), spawning yul and throwing on exit 2 - single source of truth
for the detection logic.

Verified end-to-end against DeepSeek: the block still fires and the
model still self-corrects through the refactored path.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
10 cases x 2 conditions x 10 repetitions against DeepSeek's hosted
API. See benchmark/runs-opencode-deepseek-pilot/README.md for the
full breakdown (37/100 hook reps blocked, 76% catch rate against the
nohook stale baseline, $1.1444 real cost).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A model run dumped its own environment via bash (env/printenv-style
command) during the pilot, and the real API key ended up verbatim in
3 committed transcript.jsonl files - cleaned up separately (amended
commit + git gc) since opencode's process env is inherited by every
bash command the model runs.

Two mitigations:
- opencode now runs under `env -i` with an explicit minimal
  allowlist (PATH/HOME/TERM/TMPDIR/DEEPSEEK_API_KEY/YUL_BIN) instead
  of the full launching shell's environment, so an env dump can only
  expose what's actually needed, not whatever else happens to be set
  (other API keys, SLURM tokens, etc.).
- transcript.jsonl/stderr.log get a redaction pass for any literal
  occurrence of the real key before being saved under their real
  names, regardless of how a run exposed it (env, cat .env,
  printenv, ...) - this is the actual backstop, since DEEPSEEK_API_KEY
  itself can't be kept out of opencode's own environment.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Drops the env var entirely - opencode now resolves its DeepSeek
credential from its own auth store (`opencode auth login`, saved to
~/.local/share/opencode/auth.json) instead of an env var this script
had to read/export/pass through env -i. Since the key is never in a
process environment during a run, a bash env dump has nothing to
expose.

Adds a generic auth-store-path block to the yul.js plugin (scans any
tool call's arguments - read, bash, grep, glob, whatever - for a
mention of the credential file path, not just the tools checked
before), closing the narrower residual risk of a model targeting that
file directly.

Redaction pass is now format-based (sk-[a-f0-9]{32}) instead of
value-based, so it still works after rotation without this script
ever needing to know the actual key.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Table 1 groups by ecosystem with an explicit note that "Tasks" counts
repetitions of the same task (not distinct tasks like the original
60-case sweep) - the rate denominator uses the without-yul baseline
instead of the with-yul columns, since a highly effective hook
(Maven, GitHub Actions) collapses that denominator toward zero after
correction despite real mitigation activity.

Table 2 mirrors the yul paper's per-case narrative format, adapted
for 10 repetitions per cell. Surfaced two things while verifying it
against real data:
- npm-top-10-fresh: not every mitigation is a clean fix - 4/10 reps
  evade the block by switching to a caret range on retry instead of
  adopting the suggested exact version, rather than actually
  correcting the pin.
- pypi-top-05-urllib3's one block was on an unrelated pytest
  dev-dependency pin, not urllib3 itself.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Answers a question raised in PR chains-project#58 review: only step-start, step-finish,
text, and tool part types ever appear in the JSON transcript. DeepSeek's
API does return a separate reasoning_content field, but OpenCode's
run --format json output doesn't forward it as a printable part - the
only surviving trace is step-finish's tokens.reasoning count, which is
billing accounting, not the actual reasoning text. Confirmed by
inspecting a real committed transcript (cargo-top-01-libc/hook/run-1).
…runs)

Same 10-case subset as runs-opencode-deepseek-pilot, but driven through the
new dual-runtime (apptainer|docker) run_case_opencode_deepseek.sh, 3 reps
instead of 10, and with --thinking's reasoning content actually captured
(the original pilot's README noted that gap; it's fixed here). 0 failures,
real cost $0.40.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Completes coverage of all 60 cases in cases_top.json (the first commit only
had the original 10-case pilot subset). These 50 are 1 rep each instead of
3 - reps are uneven across the full 60-case set until/unless topped up to
match Claude Sonnet 5's uniform 3-reps-per-case dataset.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…pilot

Rewrites the README against analyze_top_opencode.py's yul-scan-based
methodology (matching analyze_top.py, which produced the Claude Sonnet 5
numbers in the paper's already-latest/mitigated table) instead of the
earlier block/corrected-to-suggested framing, after a full manual review of
every rep where the target package was absent from the final manifest (63
of 160) - reading generated source, not just the manifest, to tell a
self-implementation or equivalent-package swap from a genuine miss.

Headline numbers: 0 failures, $1.238 real cost, but Rate is NOT uniformly
100% like Claude Sonnet 5's row (65% overall - 0% npm, 33% PyPI, 75%
Maven/GitHub Actions, no stale candidates in Cargo/Go) and npm's exclusion
rate is far higher than any other ecosystem (10/14 reps reimplemented
small single-purpose packages instead of depending on them). Also notes the
N caveat (uneven reps across the 60 cases) before this goes anywhere near
a direct comparison with the uniform 10-cases-per-ecosystem Claude sweep.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds reps 2-3 for the 50 cases that previously only had 1, bringing the
full cases_top.json set to a uniform 60 cases x 3 reps x 2 conditions
(360 runs total) - matching the paper's Claude Sonnet 5 methodology
exactly instead of the earlier uneven 10-cases-x-3reps + 50-cases-x-1rep
split.

Analysis (analysis_rows.json, README) not refreshed in this commit yet -
follows once the ground-truth yul-scan pass over the full 360 runs and its
manual review finish.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…aper

Recomputes Tasks/Already-latest/Mitigated/Rate per ecosystem now that all
60 cases have a uniform 3 reps x 2 conditions (360 runs), the same
methodology the paper's Claude Sonnet 5 sweep uses - directly comparable
now instead of the earlier uneven-N caveat.

Applies the paper's own exclusion criterion literally ("the coding agent
never declares the target dependency at all, implementing the required
feature itself") rather than a broader one: a rep that just failed to
complete the task (MANIFEST_NOT_WRITTEN, or wrote something unrelated)
stays in Tasks and counts against Already-latest/Rate, since the paper
doesn't define a separate bucket for that - only self-implementation and,
separately, package substitution (kept inside Tasks as "alternative" per
the paper's own Discussion section on this ambiguity) are handled
specially.

Headline: DeepSeek's overall Rate is 73% (36/49), not uniformly 100% like
Claude Sonnet 5's row - GitHub Actions/Maven come close (85%/92%), Cargo
and npm hit 0% (one stale candidate each, neither mitigated), PyPI is the
outlier at 14%. DeepSeek's exclusion rate (73/360, 20%) is also far above
Claude's (5/180, 3%), concentrated almost entirely in npm.

Also documents the credential leak found and fixed while building this
dataset (see the new README section) - one nohook run's `cat auth.json`
put four real credentials into a committed transcript, caught by GitHub's
push protection before it reached the remote, root-caused to the
auth-guard plugin only ever being installed for the hook condition.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
frankreyesgarcia and others added 4 commits September 29, 2026 16:42
…enchmark

# Conflicts:
#	main.go
#	main_test.go
Superseded entirely by runs-opencode-deepseek-docker-pilot (full 60-case
coverage, 3 reps each, matching the paper's methodology) - keeping both
just duplicated an earlier, smaller, less-analyzed run of the same idea
with no reason to reference it going forward.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… denominator

Renamed benchmark/runs-opencode-deepseek-docker-pilot to
benchmark/runs-opencode-deepseek-docker.

Also corrects the ground-truth table: Tasks should only count reps where
the model actually attempted to write the target dependency (on a manifest
yul watches), whether that attempt was blocked or not - not every rep
minus self-implementations. A second manual pass, re-reading full
transcripts instead of just final manifests, found 15 reps that never
touched the target dependency at all (mostly pypi-top-02-six/
pypi-top-10-pandas, plus a couple of zero-write sessions) that the
previous pass had left inside the Tasks denominator as "genuine misses,"
wrongly counting them against both Already-latest and Rate.

Corrected overall Rate: 86% (36/42), not 73% (36/49) - PyPI alone moves
from 14% to 100% once its abandoned-task reps are properly excluded rather
than counted as failed mitigations. Two reps also got reclassified from
"never attempted" to "alternative package" after the same re-read
surfaced real writes the pin-finder's regex didn't recognize
(cargo-top-01-libc/nohook/run-2's real Cargo.toml at a path the harness
never looks for, cargo-top-10-winapi-i686-pc-windows-gnu/hook/run-3's
windows_i686_gnu substitution).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
latest_versions_snapshot.json - package -> latest version, resolved via
yul scan against a synthetic per-ecosystem manifest seeding all 10 of that
ecosystem's target packages at once (independent of any single rep, unlike
the per-rep findings baked into analysis_rows.json). A quick reference
snapshot dated 2026-09-30, not something that stays current.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…k count

Found by direct comparison against latest_versions_snapshot.json:
docker/build-push-action and docker/setup-qemu-action both got a stale
suggested version (v7.3.0, v4.3.0) from yul's GitHub Actions resolver in
every run collected on 2026-09-27, and the model's retry wrote exactly
those suggested SHAs - a correct mitigation at the time. But the real
latest for both moved to v7.4.0/v4.4.0 by the time analyze_top_opencode.py
re-scanned on 2026-09-30, so these two reps show as "unmitigated" in the
table even though the model followed yul's own contemporaneous guidance
faithfully. Same class of resolver-lag issue the paper's own Discussion
section notes for PyPI's numpy/lombok in the Claude Sonnet 5 run.

Also makes npm's unusually low Task count (8/30 nohook, 5/30 hook -
smallest of any ecosystem by a wide margin) an explicit narrative finding
rather than something only visible by reading the Manual review list
closely.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@algomaster99

Copy link
Copy Markdown
Member

Will be done with latest release now: https://github.com/chains-project/yul/releases/tag/v0.0.18.
Has fixes for registry lag.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants