Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
30 commits
Select commit Hold shift + click to select a range
84edfd4
feat: add Qwen Code hook support and an OpenCode benchmark harness
frankreyesgarcia Sep 13, 2026
f0cddd7
fix: block bash-based bypass of the OpenCode yul hook
frankreyesgarcia Sep 13, 2026
c6738b7
Merge main into opencode-bash-block (PR #46 landed)
frankreyesgarcia Sep 15, 2026
d1a7db5
benchmark: record full 60-case OpenCode+Qwen sweep with summary
frankreyesgarcia Sep 16, 2026
58eddf3
fix: keep OpenCode benchmark logs out of the model's working directory
frankreyesgarcia Sep 16, 2026
e3d8682
chore: move OpenCode+Qwen benchmark results to their own branch
frankreyesgarcia Sep 16, 2026
7ca8985
fix: block bash-based manifest rewrites at the yul hook level
frankreyesgarcia Sep 16, 2026
a6e9b17
chore: remove qwen-extension.json
frankreyesgarcia Sep 16, 2026
3b61c9a
benchmark: emit a per-run usage.json (tokens + cost)
frankreyesgarcia Sep 21, 2026
5805a24
benchmark: capture proof of the actual provider/model per DeepSeek run
frankreyesgarcia Sep 22, 2026
431f946
benchmark: add README for the DeepSeek-V4.1-Flash pilot
frankreyesgarcia Sep 22, 2026
fe5e76e
benchmark: match the original benchmark table format, add per-run detail
frankreyesgarcia Sep 22, 2026
2276a69
docs: explain the n/a value in the per-run detail table
frankreyesgarcia Sep 22, 2026
493626e
refactor: delegate bash-write detection to the yul binary, not the JS…
frankreyesgarcia Sep 22, 2026
d1875d6
benchmark: record the DeepSeek-V4.1-Flash pilot data (200 runs)
frankreyesgarcia Sep 22, 2026
96c4f6a
security: stop leaking DEEPSEEK_API_KEY into run transcripts
frankreyesgarcia Sep 22, 2026
cd8df01
security: stop handling DEEPSEEK_API_KEY at all
frankreyesgarcia Sep 22, 2026
c2e71a0
chore: gitignore .env
frankreyesgarcia Sep 22, 2026
421a3ab
docs: add ecosystem/per-case tables matching the paper's Table 2 style
frankreyesgarcia Sep 22, 2026
b22a263
docs: note that transcript.jsonl doesn't capture reasoning content
frankreyesgarcia Sep 24, 2026
b6eab87
benchmark: add DeepSeek V4.1 Flash pilot data via Docker runtime (60 …
frankreyesgarcia Sep 27, 2026
9b8c487
benchmark: add the remaining 50 cases to the DeepSeek Docker pilot data
frankreyesgarcia Sep 27, 2026
32e1e8e
benchmark: full ground-truth analysis of the 60-case DeepSeek Docker …
frankreyesgarcia Sep 29, 2026
4aa5184
benchmark: top up the other 50 cases to 3 reps (200 more runs)
frankreyesgarcia Sep 29, 2026
915e59f
benchmark: final ground-truth table over all 360 runs, matching the p…
frankreyesgarcia Sep 29, 2026
ef49fbd
Merge remote-tracking branch 'upstream/main' into opencode-deepseek-b…
frankreyesgarcia Sep 29, 2026
b26c41a
benchmark: drop the old Apptainer-based DeepSeek pilot data
frankreyesgarcia Sep 29, 2026
4e6b56c
benchmark: rename runs-*-docker-pilot to runs-*-docker, correct Tasks…
frankreyesgarcia Sep 29, 2026
4001233
benchmark: add a flat latest-version snapshot for all 60 target packages
frankreyesgarcia Sep 30, 2026
62ad56d
benchmark: document two ecosyste.ms staleness cases and npm's low Tas…
frankreyesgarcia Sep 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
The diff you're trying to view is too large. We only load the first 3000 changed files.
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -6,3 +6,4 @@ node_modules/
target/
opencode-yul/dist/
.claude/settings.local.json
.env
54 changes: 46 additions & 8 deletions benchmark/run_case_opencode.sh
Original file line number Diff line number Diff line change
Expand Up @@ -84,12 +84,30 @@ if [ "$CONDITION" = "hook" ]; then
# throw on exit 2 so OpenCode surfaces the stderr reason back to the
# model - the same self-correction loop Claude Code's exit-2/stderr gives.
cat > "$WORKDIR/.opencode/plugins/yul.js" <<'EOF'
// All the manifest/bash-bypass detection logic lives in main.go's runHook
// now (it handles Write, Edit, and Bash tool_names), so this plugin is just
// a thin translation layer: build yul's PreToolUse JSON shape from
// whichever OpenCode tool fired, spawn the binary, and throw on exit 2 so
// OpenCode surfaces the stderr reason back to the model - the same
// self-correction loop Claude Code's exit-2/stderr gives.
export const YulPlugin = async () => {
const YUL_BIN = process.env.YUL_BIN || "yul"
const check = async (payload) => {
const proc = Bun.spawn([YUL_BIN], { stdin: "pipe", stdout: "pipe", stderr: "pipe" })
proc.stdin.write(JSON.stringify(payload))
proc.stdin.end()
const [code, stderr] = await Promise.all([proc.exited, new Response(proc.stderr).text()])
if (code === 2) throw new Error(stderr.trim() || "yul: blocked outdated dependency")
}

return {
"tool.execute.before": async (input, output) => {
if (input.tool !== "write" && input.tool !== "edit") return
const a = output.args
if (input.tool === "bash") {
await check({ tool_name: "Bash", tool_input: { command: a.command || "" } })
return
}
if (input.tool !== "write" && input.tool !== "edit") return
const payload = input.tool === "write"
? { tool_name: "Write", tool_input: { file_path: a.filePath, content: a.content } }
: {
Expand All @@ -101,12 +119,7 @@ export const YulPlugin = async () => {
replace_all: !!a.replaceAll,
},
}

const proc = Bun.spawn([YUL_BIN], { stdin: "pipe", stdout: "pipe", stderr: "pipe" })
proc.stdin.write(JSON.stringify(payload))
proc.stdin.end()
const [code, stderr] = await Promise.all([proc.exited, new Response(proc.stderr).text()])
if (code === 2) throw new Error(stderr.trim() || "yul: blocked outdated dependency")
await check(payload)
},
}
}
Expand All @@ -122,11 +135,36 @@ git init -q
git config user.email "benchmark@example.com"
git config user.name "benchmark"

# Captured to a tempfile outside WORKDIR, not directly to transcript.jsonl/
# stderr.log - the model's cwd is WORKDIR itself, so writing the live log
# there means the model can see (and, observed in practice, delete) its own
# run's log mid-session. Moved into place only after the run finishes.
TRANSCRIPT_TMP=$(mktemp)
STDERR_TMP=$(mktemp)
YUL_BIN="$YUL_BIN" "$OPENCODE_BIN" run "$PROMPT" \
--model "$MODEL_ID" \
--auto \
--format json \
> transcript.jsonl 2> stderr.log || true
> "$TRANSCRIPT_TMP" 2> "$STDERR_TMP" || true
mv "$TRANSCRIPT_TMP" transcript.jsonl
mv "$STDERR_TMP" stderr.log

# Per-run token/cost summary, aggregated from each step's usage. cost is
# whatever OpenCode's provider pricing table reports (0 for the "local"
# self-hosted provider; real dollars for a priced provider like DeepSeek).
jq -s '
[.[] | select(.type=="step_finish")] as $steps
| {
llm_calls: ($steps | length),
tokens: {
input: ($steps | map(.part.tokens.input // 0) | add // 0),
output: ($steps | map(.part.tokens.output // 0) | add // 0),
cache_read: ($steps | map(.part.tokens.cache.read // 0) | add // 0),
cache_write: ($steps | map(.part.tokens.cache.write // 0) | add // 0)
},
cost_usd: ($steps | map(.part.cost // 0) | add // 0)
}
' transcript.jsonl > usage.json

# When .manifest listed several candidate paths, use whichever one the
# model actually wrote.
Expand Down
251 changes: 251 additions & 0 deletions benchmark/run_case_opencode_deepseek.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,251 @@
#!/usr/bin/env bash
# Runs one benchmark case under one condition (hook|nohook), same as
# run_case_opencode.sh, but against DeepSeek's API via OpenCode's built-in
# "deepseek" provider (models.dev catalog) instead of a self-hosted,
# OpenAI-compatible endpoint - no custom provider config needed.
#
# Credentials come from OpenCode's own store (`opencode auth login`, saved
# to ~/.local/share/opencode/auth.json), not an env var - this script never
# reads or holds the API key, so it can't leak it. (Earlier versions passed
# DEEPSEEK_API_KEY through the environment; a run's bash tool dumped its own
# env and the real key ended up in a committed transcript. See the
# opencode-deepseek branch history.)
#
# Usage:
# run_case_opencode_deepseek.sh <cases.json> <case_id> <hook|nohook> <output_dir> [model_id] [repeat_index]
#
# model_id opencode model spec provider/model (default:
# deepseek/deepseek-v4-flash).
# repeat_index if set, run output goes to <condition>/run-<repeat_index>/
# instead of directly under <condition>/ - use for repeated
# runs of the same case/condition.
#
# Env vars:
# YUL_BIN path to the yul binary the hook condition execs
# (default: "yul" on PATH; build one with `go build -o yul .`)
# OPENCODE_BIN path to the opencode binary (default: "opencode" on PATH)
set -euo pipefail

CASES_JSON="$1"
CASE_ID="$2"
CONDITION="$3" # hook | nohook
OUT_DIR="$4"
MODEL_ID="${5:-deepseek/deepseek-v4-flash}"
REPEAT_INDEX="${6:-}"

YUL_BIN="${YUL_BIN:-yul}"
OPENCODE_BIN="${OPENCODE_BIN:-opencode}"
command -v "$OPENCODE_BIN" >/dev/null 2>&1 || { echo "opencode not found (set OPENCODE_BIN or put it on PATH)" >&2; exit 1; }
if [ "$CONDITION" = "hook" ]; then
command -v "$YUL_BIN" >/dev/null 2>&1 || { echo "yul not found (set YUL_BIN or put it on PATH)" >&2; exit 1; }
fi

case_json() {
jq -c --arg id "$CASE_ID" '.[] | select(.id == $id)' "$CASES_JSON"
}

C="$(case_json)"
if [ -z "$C" ]; then
echo "case $CASE_ID not found" >&2
exit 1
fi

# .manifest is usually a single path, but a case may instead give an array
# of candidate paths - only meaningful for "fresh" cases, since an
# "existing" case needs one fixed path to seed.
MANIFEST=$(echo "$C" | jq -r 'if (.manifest|type)=="array" then .manifest[0] else .manifest end')
TYPE=$(echo "$C" | jq -r '.type')
PROMPT=$(echo "$C" | jq -r '.prompt')
SEED=$(echo "$C" | jq -r '.seed // empty')

WORKDIR="$OUT_DIR/$CASE_ID/$CONDITION"
if [ -n "$REPEAT_INDEX" ]; then
WORKDIR="$WORKDIR/run-$REPEAT_INDEX"
fi
rm -rf "$WORKDIR"
mkdir -p "$WORKDIR/.opencode/plugins"

if [ "$TYPE" = "existing" ]; then
mkdir -p "$(dirname "$WORKDIR/$MANIFEST")"
printf '%s' "$SEED" > "$WORKDIR/$MANIFEST"
fi

# No provider block needed - "deepseek" is a built-in OpenCode provider
# (models.dev catalog); it reads DEEPSEEK_API_KEY from the environment.
cat > "$WORKDIR/opencode.json" <<EOF
{
"\$schema": "https://opencode.ai/config.json"
}
EOF

if [ "$CONDITION" = "hook" ]; then
# Mirrors main.go's runHook: translate OpenCode's tool.execute.before
# payload (write: filePath/content; edit: filePath/oldString/newString/
# replaceAll) into yul's PreToolUse JSON shape, exec the binary, and
# throw on exit 2 so OpenCode surfaces the stderr reason back to the
# model - the same self-correction loop Claude Code's exit-2/stderr gives.
cat > "$WORKDIR/.opencode/plugins/yul.js" <<'EOF'
// All the manifest/bash-bypass detection logic lives in main.go's runHook
// now (it handles Write, Edit, and Bash tool_names), so this plugin is
// mostly a thin translation layer: build yul's PreToolUse JSON shape from
// whichever OpenCode tool fired, spawn the binary, and throw on exit 2 so
// OpenCode surfaces the stderr reason back to the model - the same
// self-correction loop Claude Code's exit-2/stderr gives.
//
// It also blocks any attempt to read OpenCode's own credential store
// directly - the run's DeepSeek API key lives there (auth login, not an
// env var, see run_case_opencode_deepseek.sh), so this is the one
// exfiltration path left for a model that goes looking for it.
const AUTH_STORE_RE = /\.local[/\\]share[/\\]opencode[/\\]auth\.json|opencode[/\\]auth\.json/i

// Catches any tool call - read, bash, grep, glob, whatever - that names
// the credential store anywhere in its arguments, without having to know
// each tool's specific field names.
function mentionsAuthStore(args) {
if (typeof args === "string") return AUTH_STORE_RE.test(args)
if (Array.isArray(args)) return args.some(mentionsAuthStore)
if (args && typeof args === "object") return Object.values(args).some(mentionsAuthStore)
return false
}

export const YulPlugin = async () => {
const YUL_BIN = process.env.YUL_BIN || "yul"
const check = async (payload) => {
const proc = Bun.spawn([YUL_BIN], { stdin: "pipe", stdout: "pipe", stderr: "pipe" })
proc.stdin.write(JSON.stringify(payload))
proc.stdin.end()
const [code, stderr] = await Promise.all([proc.exited, new Response(proc.stderr).text()])
if (code === 2) throw new Error(stderr.trim() || "yul: blocked outdated dependency")
}

return {
"tool.execute.before": async (input, output) => {
const a = output.args
if (mentionsAuthStore(a)) {
throw new Error("yul: reading OpenCode's credential store is not permitted")
}
if (input.tool === "bash") {
await check({ tool_name: "Bash", tool_input: { command: a.command || "" } })
return
}
if (input.tool !== "write" && input.tool !== "edit") return
const payload = input.tool === "write"
? { tool_name: "Write", tool_input: { file_path: a.filePath, content: a.content } }
: {
tool_name: "Edit",
tool_input: {
file_path: a.filePath,
old_string: a.oldString,
new_string: a.newString,
replace_all: !!a.replaceAll,
},
}
await check(payload)
},
}
}
EOF
fi

cd "$WORKDIR"

# Cap `git rev-parse --show-toplevel` at WORKDIR so the model can't wander
# up into the real yul checkout and find real files that make it think the
# task's already done.
git init -q
git config user.email "benchmark@example.com"
git config user.name "benchmark"

# Captured to a tempfile outside WORKDIR, not directly to transcript.jsonl/
# stderr.log - the model's cwd is WORKDIR itself, so writing the live log
# there means the model can see (and, observed in practice, delete) its own
# run's log mid-session. Moved into place only after the run finishes.
TRANSCRIPT_TMP=$(mktemp)
STDERR_TMP=$(mktemp)
# Minimal env for the opencode process - it spawns the model's bash tool
# calls as children of itself, which inherit whatever's in its environment,
# so anything beyond what opencode/yul actually need (a stray GITHUB_TOKEN,
# SLURM credentials, etc. sitting in the launching shell) would otherwise
# be exposed to a model command like `env`. No DEEPSEEK_API_KEY here at
# all - opencode reads its DeepSeek credential from its own auth store
# (`opencode auth login`), not from the environment, so there's nothing
# for a bash env dump to expose in the first place.
env -i \
PATH="$PATH" \
HOME="$HOME" \
TERM="${TERM:-xterm}" \
TMPDIR="${TMPDIR:-/tmp}" \
YUL_BIN="$YUL_BIN" \
"$OPENCODE_BIN" run "$PROMPT" \
--model "$MODEL_ID" \
--auto \
--format json \
> "$TRANSCRIPT_TMP" 2> "$STDERR_TMP" || true

# Belt-and-suspenders: the yul.js plugin above blocks reads of the auth
# store, but redact anything DeepSeek-key-shaped that slips through
# anyway (format is public: "sk-" + 32 hex chars) - this doesn't require
# knowing the actual configured key, so it still works after rotation.
sed -i -E 's/sk-[a-f0-9]{32}/***REDACTED-DEEPSEEK-API-KEY***/g' "$TRANSCRIPT_TMP" "$STDERR_TMP"

mv "$TRANSCRIPT_TMP" transcript.jsonl
mv "$STDERR_TMP" stderr.log

# Per-run token/cost summary, aggregated from each step's usage. cost_usd
# is DeepSeek's real dollar cost as OpenCode's pricing table reports it.
jq -s '
[.[] | select(.type=="step_finish")] as $steps
| {
llm_calls: ($steps | length),
tokens: {
input: ($steps | map(.part.tokens.input // 0) | add // 0),
output: ($steps | map(.part.tokens.output // 0) | add // 0),
cache_read: ($steps | map(.part.tokens.cache.read // 0) | add // 0),
cache_write: ($steps | map(.part.tokens.cache.write // 0) | add // 0)
},
cost_usd: ($steps | map(.part.cost // 0) | add // 0)
}
' transcript.jsonl > usage.json

# Direct proof of which provider/model actually served this run: OpenCode's
# transcript never records it, but its own runtime log does, per session.
# Filtered by this run's sessionID so concurrent runs sharing the same
# global log don't cross-contaminate.
OPENCODE_LOG="${OPENCODE_LOG_PATH:-$HOME/.local/share/opencode/log/opencode.log}"
SESSION_ID=$(jq -r 'select(.sessionID != null) | .sessionID' transcript.jsonl 2>/dev/null | head -1)
if [ -n "$SESSION_ID" ] && [ -f "$OPENCODE_LOG" ]; then
grep -F "session.id=$SESSION_ID" "$OPENCODE_LOG" | grep -E "providerID=|llm\.provider=" > model_used.log || true
else
: > model_used.log
fi

# When .manifest listed several candidate paths, use whichever one the
# model actually wrote.
FOUND_MANIFEST=""
shopt -s nullglob
while IFS= read -r candidate; do
if [[ "$candidate" == *"*"* ]]; then
matches=( $candidate )
if [ ${#matches[@]} -gt 0 ]; then
FOUND_MANIFEST="${matches[0]}"
break
fi
elif [ -f "$candidate" ]; then
FOUND_MANIFEST="$candidate"
break
fi
done < <(echo "$C" | jq -r 'if (.manifest|type)=="array" then .manifest[] else .manifest end')
shopt -u nullglob

if [ -n "$FOUND_MANIFEST" ]; then
cp "$FOUND_MANIFEST" "final_manifest"
echo "$FOUND_MANIFEST" > "final_manifest_path"
else
echo "MANIFEST_NOT_WRITTEN" > final_manifest
fi

# Removed below so the run output doesn't end up with a nested-repo
# gitlink when committed.
rm -rf .git

echo "done: $CASE_ID [$CONDITION] -> $WORKDIR"
27 changes: 27 additions & 0 deletions benchmark/runs-opencode-deepseek-docker/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
# Regenerable build/scratch artifacts, not experiment data - excluded from
# every run dir in this tree (see the .rsync exclude list used to produce
# this directory, kept here so a future re-sync stays consistent).
node_modules/
.venv/
.venv*/
venv/
env/
target/
.m2/
__pycache__/
*.py[cod]
.mypy_cache/
.pytest_cache/
.ruff_cache/
.tox/
.eggs/
*.egg-info/
dist/
build/
bin/
.idea/
.vscode/
.DS_Store
*.class
*.jar
*.war
Loading
Loading