From 9ae43a0fa7b16970ab582f208c88411bea85c404 Mon Sep 17 00:00:00 2001 From: Poytr1 Date: Tue, 7 Jul 2026 17:27:23 +0800 Subject: [PATCH 1/2] =?UTF-8?q?test(bench):=20verbose=20aksops=20output=20?= =?UTF-8?q?=E2=80=94=20tests=20the=20mock-confound=20hypothesis?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit @pchsu asked if the ~neutral ops result was a mock artifact (real az/kubectl emit large output that round-trips through the model per command; near-silent mocks understate what a compiled runbook skill could save). Made the mocks emit realistic sizable output (~17KB/runbook: az kubeconfig JSON, multi-line rollout progress, a 120-row `kubectl get pods -o wide` table). Result (rollouts=6, both arms 100% → iso-accuracy): still ~neutral — baseline mean 321,820 (med 330,304) vs jit mean 330,034 (med 329,740) -> -2.6% by mean, +0.2% by median. No JIT win. So the mock was NOT the main issue. Deeper finding: ~17KB (~4-5k tokens) of output is a rounding error against the ~330k baseline, which is dominated by fixed system-prompt/tool-def/cache-reread load, not task-specific work. A skill that shortcuts the task-specific portion can only ever save a small fraction — the tokens aren't where JIT optimizes. Co-Authored-By: Claude Opus 4.8 --- internal/bench/aksops_fixture.go | 38 ++++++++++++++++++++++++++++---- 1 file changed, 34 insertions(+), 4 deletions(-) diff --git a/internal/bench/aksops_fixture.go b/internal/bench/aksops_fixture.go index e16d091..550abf7 100644 --- a/internal/bench/aksops_fixture.go +++ b/internal/bench/aksops_fixture.go @@ -81,11 +81,28 @@ func (f AKSOpsFixture) SeedSessions(logsDir string) error { // output shaped like the real CLI. The header comment states plainly that these // are the runbook's intended local tools, so the agent runs them rather than // second-guessing whether "fake" commands are worth executing. +// +// The output is intentionally SIZABLE (JSON blobs, a many-row pod table, rollout +// progress) to mirror real az/kubectl: each command's output round-trips into the +// model's context in the baseline arm, which is the per-command cost a compiled +// runbook skill (run in one shot) might avoid. Near-silent mocks understate that. const mockAz = `#!/usr/bin/env bash # az — local test double for this ops runbook environment. Records to ops.log. echo "az $*" >> ops.log if [ "$1 $2" = "aks get-credentials" ]; then echo "Merged \"prod\" as current context in ~/.kube/config" + cat <<'JSON' +{ + "apiVersion": "v1", + "clusters": [ + {"cluster": {"certificate-authority-data": "LS0tLS1CRUdJTiBDRVJUSUZJQ0FURS0tLS0t...", "server": "https://prod-dns-a1b2c3d4.hcp.eastus.azmk8s.io:443"}, "name": "prod"} + ], + "contexts": [{"context": {"cluster": "prod", "user": "clusterUser_rg_prod"}, "name": "prod"}], + "current-context": "prod", + "kind": "Config", + "users": [{"name": "clusterUser_rg_prod", "user": {"token": "eyJhbGciOiJSUzI1NiIsImtpZCI6..."}}] +} +JSON else echo "{\"status\": \"ok\"}" fi @@ -96,10 +113,23 @@ const mockKubectl = `#!/usr/bin/env bash # kubectl — local test double for this ops runbook environment. Records to ops.log. echo "kubectl $*" >> ops.log case "$1" in - scale) echo "deployment.apps/web scaled" ;; - rollout) echo "deployment \"web\" successfully rolled out" ;; - get) printf 'NAME READY STATUS RESTARTS AGE\nweb-7d9f8c6b5-abcde 1/1 Running 0 5m\n' ;; - *) echo "ok" ;; + scale) + echo "deployment.apps/web scaled" ;; + rollout) + echo "Waiting for deployment \"web\" rollout to finish: 0 of 5 updated replicas are available..." + echo "Waiting for deployment \"web\" rollout to finish: 1 of 5 updated replicas are available..." + echo "Waiting for deployment \"web\" rollout to finish: 2 of 5 updated replicas are available..." + echo "Waiting for deployment \"web\" rollout to finish: 3 of 5 updated replicas are available..." + echo "Waiting for deployment \"web\" rollout to finish: 4 of 5 updated replicas are available..." + echo "deployment \"web\" successfully rolled out" ;; + get) + echo "NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES" + for i in $(seq 1 120); do + printf 'web-7d9f8c6b5-%s 1/1 Running 0 %dm 10.244.%d.%d aks-nodepool1-24680123-vmss0000%02d \n' \ + "$(printf '%05x' $((RANDOM % 1000000)))" "$i" $((i % 3)) $((i + 10)) $((i % 6)) + done ;; + *) + echo "ok" ;; esac exit 0 ` From f320fe685ce4f220a915824be7d2f061a5087834 Mon Sep 17 00:00:00 2001 From: Poytr1 Date: Tue, 7 Jul 2026 17:29:33 +0800 Subject: [PATCH 2/2] =?UTF-8?q?docs:=20the=20structural=20finding=20?= =?UTF-8?q?=E2=80=94=20token=20cost=20is=20fixed=20agent=20overhead,=20not?= =?UTF-8?q?=20the=20work?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds the deepest conclusion from the verbose-mock test: every episode is ~250-330k tokens dominated by fixed system-prompt/tool-def/cache-reread overhead; the task-specific work (edit, few commands, even 17KB of output) is a small slice, so JIT can only shortcut that slice and is structurally limited at saving tokens for Claude-Code-shaped tasks. Reframes AgentJIT's likely value as latency/determinism (not measured here) and recommends a benchmark v2 with wall-clock + reliability. Co-Authored-By: Claude Opus 4.8 --- .../specs/2026-07-06-benchmark-findings.md | 14 ++++++++++++++ 1 file changed, 14 insertions(+) diff --git a/docs/superpowers/specs/2026-07-06-benchmark-findings.md b/docs/superpowers/specs/2026-07-06-benchmark-findings.md index ec4e8d9..7d8aa8c 100644 --- a/docs/superpowers/specs/2026-07-06-benchmark-findings.md +++ b/docs/superpowers/specs/2026-07-06-benchmark-findings.md @@ -37,6 +37,20 @@ study finds **no case where it clearly wins** — code-edit shapes are neutral-t ops shape came out ~neutral once measured reliably. A win, if it exists, likely needs workflows with *much* larger per-episode work that a compiled action can wholesale replace (not four cheap commands). +**The structural reason (the deepest finding): the token cost is not where JIT optimizes.** Every episode +costs ~250–330k tokens, and that is dominated by **fixed agent overhead** — the system prompt, tool +definitions, and the context re-read every turn via cache (~44k cache-creation just to *boot* a session, +then cache-reads compounding across turns). The *task-specific* work — a code edit, four commands, even +17KB of verbose command output (tested: `az`/`kubectl` mocks emitting a 120-row pod table + JSON + rollout +logs still moved the result <3%) — is a small slice of the total. A skill can only shortcut that slice, so +**JIT compilation is structurally limited at saving tokens for Claude-Code-shaped tasks.** To make the +task-specific portion dominate, the per-episode work would have to be enormous (hundreds of KB of output, +or very long tool sequences) — not typical. + +**Implication:** AgentJIT's real value is more likely **latency and determinism** (a compiled script runs +in milliseconds and can't hallucinate or skip a step) than raw token savings — a dimension this token +benchmark does not yet measure. A benchmark v2 should add wall-clock and reliability. + **Lesson baked in:** always compare at iso-accuracy and check the *verified* sample size on both arms — a mean over a handful of lucky successes is not a result. (`aj bench --compare` prints an iso-accuracy WARN when success rates differ; trust it.)