feat(bench): aksops SRE fixture — where JIT actually pays off - #15
Merged
Merged
Conversation
Adds `aksops`: a repetitive AKS operations runbook (az aks get-credentials
+ kubectl scale/rollout/get), the SRE/tool-use shape AgentJIT was designed
for. az/kubectl are MOCKED (bin/ scripts printing plausible CLI output +
recording to ops.log), so it's hermetic — no real cluster. Bash-shaped, so
the deterministic compiler produces a REAL runnable runbook skill; the JIT
arm invokes it to DO the work instead of reasoning through each command.
CompilableFixture (SkillName verified empirically = bin-az-aks-get-credentials).
Tests: verifier fails-before/passes-after, and a real `aj compile` E2E installs
the runbook script (zero-token deterministic).
RESULT (n=3, rollouts=3, live claude) — first positive one:
baseline 332,620 (med 332,565) vs jit 285,391 (med 332,996) → saving +47,228/use (~14%)
Both 100% iso-accuracy. This is the vindication of the ops-vs-code distinction:
a compiled *runnable* skill replaces work (running the commands), unlike a
code-edit skill that only annotates. Caveat: jit mean<median => variance;
the +14% direction is clear but a tighter number needs more rollouts.
Note: the mock tools had to print realistic output + the prompt made mechanical
("execute exactly") — an early version failed because the agent noticed the
stubs were fake and REFUSED to run them (a real lesson about hermetic mocks
changing agent behavior).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rewrites the TL;DR around the sharp code-vs-ops split: nullcheck/shellseq ~0%, migrate −25%, aksops +14%. Adds the aksops results section, the traced "why a code-edit skill doesn't save" mechanism (agent explores before opening the skill, reads files anyway, skill adds context), and the hermetic-mock-changes-behavior lesson. Guidance: compile repetitive Bash/tool-use (SRE runbooks), be skeptical of code-edit patterns — value is replacing execution, not annotating edits. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
aksops— a repetitive SRE/operations runbook (az aks get-credentials+kubectl scale/rollout/get), the tool-use shape AgentJIT was actually designed for.az/kubectlare mocked (bin/ scripts printing plausible CLI output + recording toops.log), so it's hermetic — no real cluster. Bash-shaped, so the deterministic compiler produces a real runnable runbook skill; the JIT arm invokes it to do the work instead of reasoning through each command.Result (n=3, rollouts=3, live claude) — the first positive one
This vindicates the ops-vs-code distinction from the earlier findings: a compiled runnable skill replaces work (running the command sequence), whereas a code-edit skill only annotates (the agent still reads + edits). So JIT compilation pays off for operations/tool-use workflows, not code-edit ones.
Caveat: jit mean < median → variance across rollouts; the +14% direction is clear but a tight number needs more rollouts.
Tests
CompilableFixture(skill name verified empirically). Verifier fails-before / passes-after; a realaj compileE2E installs the runbook script (zero-token deterministic).golangci-lintclean.Methodology note (worth reading)
An early version failed because the agent inspected the mock tools, realized they were stubs, and refused to run them. Fix: mocks print realistic output + the prompt is mechanical ("execute exactly"). A real lesson that hermetic mocks can change agent behavior.
Findings-doc update to follow.
🤖 Generated with Claude Code