Opencode deepseek benchmark - #70
Closed
frankreyesgarcia wants to merge 30 commits into
Closed
frankreyesgarcia wants to merge 30 commits into
frankreyesgarcia wants to merge 30 commits into
Conversation
Adds two pieces of harness support alongside the existing Claude Code plugin,
both validated live against real installs rather than just against docs.
Qwen Code (qwen-extension.json):
- Qwen Code's extension loader reads hooks/hooks.json via the same path,
JSON shape, and ${CLAUDE_PLUGIN_ROOT} substitution Claude Code uses, and
`qwen extensions install <marketplace>:<plugin>` already installs straight
from this repo's existing chains-hooks marketplace listing.
- Found and fixed a real bug: Qwen Code's hook executor treats the "timeout"
field as milliseconds, not seconds like Claude Code. hooks.json's
"timeout": 60 (meant as 60s) became a 60ms timeout, so ensure-yul.sh's
binary download failed instantly every session (silently, since it fails
open). qwen-extension.json ships the same hooks with millisecond-scaled
timeouts; Qwen Code prefers an extension's own manifest over the
hooks/hooks.json fallback, so this fixes the bug without touching main.go.
- Verified end-to-end: installed from this repo's local path into a
sandboxed Qwen Code config, confirmed both SessionStart hooks now succeed,
confirmed the real release binary downloads and caches correctly, and fed
the cached binary a synthetic PreToolUse payload in Qwen Code's exact
documented shape (tool_name/tool_input.file_path/content) - it blocked an
outdated pin with exit 2 as expected. No Go code changes were needed.
OpenCode (benchmark/run_case_opencode.sh):
- Parallel to benchmark/run_case.sh, but drives `opencode run` against a
custom OpenAI-compatible provider (for a self-hosted vLLM endpoint) instead
of `claude -p`, and wires up yul via a small JS plugin
(.opencode/plugins/yul.js) implementing OpenCode's tool.execute.before
hook, translating its write/edit args into yul's existing stdin JSON shape
and shelling out to the same compiled binary.
- Live-tested against a real Qwen3.6-35B-A3B vLLM server: the block mechanism
works end-to-end (plugin invokes yul, exit 2 becomes a thrown Error,
OpenCode surfaces the exact reason as the tool result).
A follow-up change hardens this against a bash-based bypass found during
that testing.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T2ZAeghvWbbLnsHQ9c9vA4
The OpenCode yul plugin only intercepted tool.execute.before for write/edit, so a model blocked on an outdated pin could just write the same content via the bash tool instead - completely skipping the check. Live-tested this against a real Qwen3.6-35B-A3B model: it blocked cleanly on `edit` (yul correctly flagged pyyaml 5.3.1 -> 6.0.3), then wrote the file via `printf ... > requirements.txt` with the same outdated pin instead of retrying with the suggested version. This is a limitation inherited from yul's own PreToolUse matcher, which is Write|Edit only in Claude Code and Qwen Code too - just newly visible here since OpenCode's plugin API exposes bash as a distinctly interceptable tool. Fix: the plugin now also intercepts bash calls, and blocks any command that both mentions a known manifest filename (pom.xml, requirements.txt, pyproject.toml, package.json, go.mod, Cargo.toml, .github/workflows/*.yml) and contains a file-mutating construct (>, >>, tee, sed -i, perl -i, dd of=, cp, mv), rather than trying to parse shell semantics. Validated the detection regex against 14 synthetic commands (real bypass patterns and plausible false positives like `cat requirements.txt`, `git diff pom.xml`, `pip install -r requirements.txt`) - all matched as expected. Re-ran the 10-case/20-run benchmark across all 6 ecosystems after this fix: no bash bypass observed. Two cases showed the hook blocking an outdated pin and the model self-correcting on retry through the proper channel (maven lombok 1.18.34->1.18.48; a GitHub Actions checkout upgraded from @v4 to the exact SHA for v7.0.1). One case (maven junit) showed the hook blocking correctly twice but the model retrying identical rejected content instead of using the suggested versions - a model self-correction gap, not a hook bug. Remaining timeouts/no-ops in that run were unrelated to yul (no Rust/gcc toolchain reachable in the sandbox, Node not on PATH, or the model choosing to inline a feature instead of adding the dependency). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T2ZAeghvWbbLnsHQ9c9vA4
Brings in the merged Qwen Code/OpenCode PR (chains-project#46) plus other main changes since this branch forked, while keeping the bash-bypass fix. # Conflicts: # benchmark/run_case_opencode.sh
Runs benchmark/cases_top.json under both hook and nohook via OpenCode against a self-hosted Qwen3.6-35B-A3B-FP8 vLLM server, on the bash-bypass-fixed harness (f0cddd7). 21/60 cases triggered a real yul block; 0 of 7 bash-bypass attempts succeeded (vs. 4/16 on a prior run without the fix). See benchmark/runs-opencode-qwen-60/SUMMARY.md for the full breakdown. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
transcript.jsonl/stderr.log were written via relative redirection after cd "$WORKDIR" - the same directory the model operates in - so the model could see and mutate its own run's live log files. Observed in practice: go-top-02-go-difflib/hook ended up with no transcript at all, apparently deleted mid-run by the model treating it as a build artifact. Now captured to a tempfile outside WORKDIR and moved into place only after the run finishes. Re-ran go-top-02-go-difflib/hook with the fix; no block triggered (go-difflib v1.0.0 is current), so this doesn't change any aggregate numbers in SUMMARY.md. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
yul's PreToolUse hook only ever saw Write and Edit tool calls, so a model blocked on an outdated pin could rewrite the same manifest via bash (cat > file, sed -i, etc.) and skip the check entirely - the same gap the OpenCode benchmark plugin already worked around (f0cddd7), but inherited by Claude Code and Qwen Code too since they only wire yul's hook to Write|Edit. Adds a Bash branch to runHook: since yul can't simulate arbitrary shell to know what content would land, it's a coarse block (any command that both names a known manifest and contains a content-mutating construct) rather than a real version check. Ported the heuristic from the OpenCode plugin's JS regexes to Go, fixing a boundary bug found in the process (dd of=go.mod wasn't matching because "=" wasn't a recognized separator) and applying the same fix back to the JS version for parity. Updates hooks/hooks.json's PreToolUse matcher to Write|Edit|Bash, and qwen-extension.json's matcher to match (kept untracked, as before). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Qwen Code support isn't something we're pursuing right now; drop the extension manifest rather than carrying unused, unmaintained config. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Aggregates each run's step_finish events into a single summary file alongside transcript.jsonl, so token/cost accounting doesn't require re-parsing the full transcript every time. cost_usd is whatever OpenCode's provider pricing table reports - 0 for the self-hosted "local" provider used so far, real dollars for a priced provider like DeepSeek. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
transcript.jsonl never records which provider/model served a run - OpenCode's own runtime log does, per sessionID. Grep it (filtered by this run's sessionID, so concurrent runs don't cross-contaminate) and save the matching provider/model lines to model_used.log alongside transcript.jsonl/usage.json, instead of having to dig through the global log by hand to verify which model actually ran. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Mirrors benchmark/README.md's format (intro, summary table, narrative findings) for the 200-run (10 cases x 2 conditions x 10 reps) pilot against DeepSeek's hosted API instead of a single-run sweep. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Replaces the ad-hoc summary table with the same column set as benchmark/README.md's top-60 table (Tasks/Versioned/Already latest/ Blocked/Rate), one row per case since we only have 10. "Already latest" is now ground truth - every final_manifest replayed through the real yul binary (before="") against live registries, not a regex guess. Also fixes a narrative bug: the one urllib3 block was actually on an unrelated pytest pin in a pyproject.toml the model also wrote, not on urllib3 itself. Adds a collapsible 200-row per-run detail table (case, condition, rep, versioned, already latest, blocked, flagged dependency). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… plugin main.go's runHook now handles tool_name "Bash" itself (7ca8985), so the OpenCode plugin no longer needs its own copy of the manifest/write-construct regexes - that was duplicate logic to keep in sync by hand (already caused one bug: a boundary fix landed in Go first and had to be remembered for the JS copy too). The plugin is now a thin translation layer for all three tool kinds (write/edit/ bash), spawning yul and throwing on exit 2 - single source of truth for the detection logic. Verified end-to-end against DeepSeek: the block still fires and the model still self-corrects through the refactored path. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
10 cases x 2 conditions x 10 repetitions against DeepSeek's hosted API. See benchmark/runs-opencode-deepseek-pilot/README.md for the full breakdown (37/100 hook reps blocked, 76% catch rate against the nohook stale baseline, $1.1444 real cost). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
A model run dumped its own environment via bash (env/printenv-style command) during the pilot, and the real API key ended up verbatim in 3 committed transcript.jsonl files - cleaned up separately (amended commit + git gc) since opencode's process env is inherited by every bash command the model runs. Two mitigations: - opencode now runs under `env -i` with an explicit minimal allowlist (PATH/HOME/TERM/TMPDIR/DEEPSEEK_API_KEY/YUL_BIN) instead of the full launching shell's environment, so an env dump can only expose what's actually needed, not whatever else happens to be set (other API keys, SLURM tokens, etc.). - transcript.jsonl/stderr.log get a redaction pass for any literal occurrence of the real key before being saved under their real names, regardless of how a run exposed it (env, cat .env, printenv, ...) - this is the actual backstop, since DEEPSEEK_API_KEY itself can't be kept out of opencode's own environment. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Drops the env var entirely - opencode now resolves its DeepSeek
credential from its own auth store (`opencode auth login`, saved to
~/.local/share/opencode/auth.json) instead of an env var this script
had to read/export/pass through env -i. Since the key is never in a
process environment during a run, a bash env dump has nothing to
expose.
Adds a generic auth-store-path block to the yul.js plugin (scans any
tool call's arguments - read, bash, grep, glob, whatever - for a
mention of the credential file path, not just the tools checked
before), closing the narrower residual risk of a model targeting that
file directly.
Redaction pass is now format-based (sk-[a-f0-9]{32}) instead of
value-based, so it still works after rotation without this script
ever needing to know the actual key.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Table 1 groups by ecosystem with an explicit note that "Tasks" counts repetitions of the same task (not distinct tasks like the original 60-case sweep) - the rate denominator uses the without-yul baseline instead of the with-yul columns, since a highly effective hook (Maven, GitHub Actions) collapses that denominator toward zero after correction despite real mitigation activity. Table 2 mirrors the yul paper's per-case narrative format, adapted for 10 repetitions per cell. Surfaced two things while verifying it against real data: - npm-top-10-fresh: not every mitigation is a clean fix - 4/10 reps evade the block by switching to a caret range on retry instead of adopting the suggested exact version, rather than actually correcting the pin. - pypi-top-05-urllib3's one block was on an unrelated pytest dev-dependency pin, not urllib3 itself. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Answers a question raised in PR chains-project#58 review: only step-start, step-finish, text, and tool part types ever appear in the JSON transcript. DeepSeek's API does return a separate reasoning_content field, but OpenCode's run --format json output doesn't forward it as a printable part - the only surviving trace is step-finish's tokens.reasoning count, which is billing accounting, not the actual reasoning text. Confirmed by inspecting a real committed transcript (cargo-top-01-libc/hook/run-1).
…runs) Same 10-case subset as runs-opencode-deepseek-pilot, but driven through the new dual-runtime (apptainer|docker) run_case_opencode_deepseek.sh, 3 reps instead of 10, and with --thinking's reasoning content actually captured (the original pilot's README noted that gap; it's fixed here). 0 failures, real cost $0.40. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Completes coverage of all 60 cases in cases_top.json (the first commit only had the original 10-case pilot subset). These 50 are 1 rep each instead of 3 - reps are uneven across the full 60-case set until/unless topped up to match Claude Sonnet 5's uniform 3-reps-per-case dataset. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…pilot Rewrites the README against analyze_top_opencode.py's yul-scan-based methodology (matching analyze_top.py, which produced the Claude Sonnet 5 numbers in the paper's already-latest/mitigated table) instead of the earlier block/corrected-to-suggested framing, after a full manual review of every rep where the target package was absent from the final manifest (63 of 160) - reading generated source, not just the manifest, to tell a self-implementation or equivalent-package swap from a genuine miss. Headline numbers: 0 failures, $1.238 real cost, but Rate is NOT uniformly 100% like Claude Sonnet 5's row (65% overall - 0% npm, 33% PyPI, 75% Maven/GitHub Actions, no stale candidates in Cargo/Go) and npm's exclusion rate is far higher than any other ecosystem (10/14 reps reimplemented small single-purpose packages instead of depending on them). Also notes the N caveat (uneven reps across the 60 cases) before this goes anywhere near a direct comparison with the uniform 10-cases-per-ecosystem Claude sweep. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Adds reps 2-3 for the 50 cases that previously only had 1, bringing the full cases_top.json set to a uniform 60 cases x 3 reps x 2 conditions (360 runs total) - matching the paper's Claude Sonnet 5 methodology exactly instead of the earlier uneven 10-cases-x-3reps + 50-cases-x-1rep split. Analysis (analysis_rows.json, README) not refreshed in this commit yet - follows once the ground-truth yul-scan pass over the full 360 runs and its manual review finish. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…aper
Recomputes Tasks/Already-latest/Mitigated/Rate per ecosystem now that all
60 cases have a uniform 3 reps x 2 conditions (360 runs), the same
methodology the paper's Claude Sonnet 5 sweep uses - directly comparable
now instead of the earlier uneven-N caveat.
Applies the paper's own exclusion criterion literally ("the coding agent
never declares the target dependency at all, implementing the required
feature itself") rather than a broader one: a rep that just failed to
complete the task (MANIFEST_NOT_WRITTEN, or wrote something unrelated)
stays in Tasks and counts against Already-latest/Rate, since the paper
doesn't define a separate bucket for that - only self-implementation and,
separately, package substitution (kept inside Tasks as "alternative" per
the paper's own Discussion section on this ambiguity) are handled
specially.
Headline: DeepSeek's overall Rate is 73% (36/49), not uniformly 100% like
Claude Sonnet 5's row - GitHub Actions/Maven come close (85%/92%), Cargo
and npm hit 0% (one stale candidate each, neither mitigated), PyPI is the
outlier at 14%. DeepSeek's exclusion rate (73/360, 20%) is also far above
Claude's (5/180, 3%), concentrated almost entirely in npm.
Also documents the credential leak found and fixed while building this
dataset (see the new README section) - one nohook run's `cat auth.json`
put four real credentials into a committed transcript, caught by GitHub's
push protection before it reached the remote, root-caused to the
auth-guard plugin only ever being installed for the hook condition.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…enchmark # Conflicts: # main.go # main_test.go
Superseded entirely by runs-opencode-deepseek-docker-pilot (full 60-case coverage, 3 reps each, matching the paper's methodology) - keeping both just duplicated an earlier, smaller, less-analyzed run of the same idea with no reason to reference it going forward. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
… denominator Renamed benchmark/runs-opencode-deepseek-docker-pilot to benchmark/runs-opencode-deepseek-docker. Also corrects the ground-truth table: Tasks should only count reps where the model actually attempted to write the target dependency (on a manifest yul watches), whether that attempt was blocked or not - not every rep minus self-implementations. A second manual pass, re-reading full transcripts instead of just final manifests, found 15 reps that never touched the target dependency at all (mostly pypi-top-02-six/ pypi-top-10-pandas, plus a couple of zero-write sessions) that the previous pass had left inside the Tasks denominator as "genuine misses," wrongly counting them against both Already-latest and Rate. Corrected overall Rate: 86% (36/42), not 73% (36/49) - PyPI alone moves from 14% to 100% once its abandoned-task reps are properly excluded rather than counted as failed mitigations. Two reps also got reclassified from "never attempted" to "alternative package" after the same re-read surfaced real writes the pin-finder's regex didn't recognize (cargo-top-01-libc/nohook/run-2's real Cargo.toml at a path the harness never looks for, cargo-top-10-winapi-i686-pc-windows-gnu/hook/run-3's windows_i686_gnu substitution). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
latest_versions_snapshot.json - package -> latest version, resolved via yul scan against a synthetic per-ecosystem manifest seeding all 10 of that ecosystem's target packages at once (independent of any single rep, unlike the per-rep findings baked into analysis_rows.json). A quick reference snapshot dated 2026-09-30, not something that stays current. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
…k count Found by direct comparison against latest_versions_snapshot.json: docker/build-push-action and docker/setup-qemu-action both got a stale suggested version (v7.3.0, v4.3.0) from yul's GitHub Actions resolver in every run collected on 2026-09-27, and the model's retry wrote exactly those suggested SHAs - a correct mitigation at the time. But the real latest for both moved to v7.4.0/v4.4.0 by the time analyze_top_opencode.py re-scanned on 2026-09-30, so these two reps show as "unmitigated" in the table even though the model followed yul's own contemporaneous guidance faithfully. Same class of resolver-lag issue the paper's own Discussion section notes for PyPI's numpy/lombok in the Claude Sonnet 5 run. Also makes npm's unusually low Task count (8/30 nohook, 5/30 hook - smallest of any ecosystem by a wide margin) an explicit narrative finding rather than something only visible by reading the Manual review list closely. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Member
|
Will be done with latest release now: https://github.com/chains-project/yul/releases/tag/v0.0.18. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
yulagainst the same 60 dependency-addition tasks the paper's Claude Sonnet 5 sweep uses (benchmark/cases_top.json, 10 per ecosystem across Cargo, GitHub Actions, Go, Maven, npm, PyPI), driven through OpenCode against DeepSeek's hosted API (deepseek/deepseek-flash, model nameDeepSeek V4.1 Flash) instead ofclaude -pagainst Claude — same repetition scheme, 60 cases × 2 conditions × 3 reps, 360 runs total.Sandboxed via Docker (
CONTAINER_RUNTIME=apptainer|dockerauto-detects; Apptainer only runs on Linux and shares the host kernel, which a Docker host can't guarantee). Full details, methodology, and manual-review reasoning are inbenchmark/runs-opencode-deepseek-docker/README.md.360 runs, 0 failures, real cost $2.9045.
Table (same format as the paper's Already-latest/mitigation table)
Claude Sonnet 5 (
bench/top-final-results, for reference):DeepSeek V4.1 Flash (this PR):
DeepSeek excludes/never-attempts far more reps than Claude (88/360, 24%, vs 5/180, 3%) — concentrated in npm (self-implements small single-purpose packages instead of depending on them) and PyPI's
six/pandascases (abandoned, or written to a fileyuldoesn't watch). Rate isn't uniformly 100% like Claude's row: GitHub Actions (85%) is the biggest real gap, npm has one unmitigated stale candidate (the model deleted the dependency instead of fixing its version), everything else is 92-100% or has no stale candidates left. Full manual-review breakdown and methodology caveats (ground truth viayul scanitself, an independent ecosyste.ms cross-check that only partially worked) are in the README linked above.A credential leak was found and fixed while building this dataset (nohook runs had no protection against a model reading OpenCode's credential store) — documented in the README, redacted before ever reaching this branch's pushed history, all affected credentials rotated.
Test plan
benchmark/analyze_top_opencode.py) run against the full dataset viayul scango build ./...andgo test ./...pass after mergingmain