Skip to content

feat(bench): aksops SRE fixture — where JIT actually pays off - #15

Merged
Poytr1 merged 2 commits into
mainfrom
feat/aj-bench-aksops
Jul 7, 2026
Merged

Poytr1 merged 2 commits into
mainfrom
feat/aj-bench-aksops

Conversation

@Poytr1

@Poytr1 Poytr1 commented Jul 7, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds aksops — a repetitive SRE/operations runbook (az aks get-credentials + kubectl scale/rollout/get), the tool-use shape AgentJIT was actually designed for. az/kubectl are mocked (bin/ scripts printing plausible CLI output + recording to ops.log), so it's hermetic — no real cluster. Bash-shaped, so the deterministic compiler produces a real runnable runbook skill; the JIT arm invokes it to do the work instead of reasoning through each command.

Result (n=3, rollouts=3, live claude) — the first positive one

aksops-3:  baseline 100% / jit 100%   (iso-accuracy, both ran the runbook)
   T2S baseline 332,620 (med 332,565)  vs  jit 285,391 (med 332,996)  →  saving +47,228/use (~14%)

This vindicates the ops-vs-code distinction from the earlier findings: a compiled runnable skill replaces work (running the command sequence), whereas a code-edit skill only annotates (the agent still reads + edits). So JIT compilation pays off for operations/tool-use workflows, not code-edit ones.

Caveat: jit mean < median → variance across rollouts; the +14% direction is clear but a tight number needs more rollouts.

Tests

CompilableFixture (skill name verified empirically). Verifier fails-before / passes-after; a real aj compile E2E installs the runbook script (zero-token deterministic). golangci-lint clean.

Methodology note (worth reading)

An early version failed because the agent inspected the mock tools, realized they were stubs, and refused to run them. Fix: mocks print realistic output + the prompt is mechanical ("execute exactly"). A real lesson that hermetic mocks can change agent behavior.

Findings-doc update to follow.

🤖 Generated with Claude Code

Poytr1 and others added 2 commits July 7, 2026 16:21
Adds `aksops`: a repetitive AKS operations runbook (az aks get-credentials
+ kubectl scale/rollout/get), the SRE/tool-use shape AgentJIT was designed
for. az/kubectl are MOCKED (bin/ scripts printing plausible CLI output +
recording to ops.log), so it's hermetic — no real cluster. Bash-shaped, so
the deterministic compiler produces a REAL runnable runbook skill; the JIT
arm invokes it to DO the work instead of reasoning through each command.

CompilableFixture (SkillName verified empirically = bin-az-aks-get-credentials).
Tests: verifier fails-before/passes-after, and a real `aj compile` E2E installs
the runbook script (zero-token deterministic).

RESULT (n=3, rollouts=3, live claude) — first positive one:
  baseline 332,620 (med 332,565)  vs  jit 285,391 (med 332,996)  →  saving +47,228/use (~14%)
Both 100% iso-accuracy. This is the vindication of the ops-vs-code distinction:
a compiled *runnable* skill replaces work (running the commands), unlike a
code-edit skill that only annotates. Caveat: jit mean<median => variance;
the +14% direction is clear but a tighter number needs more rollouts.

Note: the mock tools had to print realistic output + the prompt made mechanical
("execute exactly") — an early version failed because the agent noticed the
stubs were fake and REFUSED to run them (a real lesson about hermetic mocks
changing agent behavior).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rewrites the TL;DR around the sharp code-vs-ops split: nullcheck/shellseq
~0%, migrate −25%, aksops +14%. Adds the aksops results section, the
traced "why a code-edit skill doesn't save" mechanism (agent explores
before opening the skill, reads files anyway, skill adds context), and the
hermetic-mock-changes-behavior lesson. Guidance: compile repetitive
Bash/tool-use (SRE runbooks), be skeptical of code-edit patterns — value is
replacing execution, not annotating edits.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@Poytr1
Poytr1 merged commit d699aad into main Jul 7, 2026
2 checks passed
@Poytr1
Poytr1 deleted the feat/aj-bench-aksops branch July 7, 2026 08:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant