Behavior Diff is a local agent plugin that compares behavior before and after an instruction-file change. It connects a user request, controlled agent trials, evidence analysis, and a report. The user decides whether the observed change meets their intent.
The installable source lives in plugin/, with manifests for
Claude Code and Codex. Skills define the agent workflow. Shared Bash and Python
scripts handle execution, evidence, and reporting.
+---------------------------------------------------------+
| User in Claude Code or Codex |
| Direct request or accepted edit-hook suggestion |
+----------------------------+----------------------------+
v
+---------------------------------------------------------+
| Skill orchestration |
| Select instruction change, scenario, agent, and model |
+----------------------------+----------------------------+
v
+---------------------------------------------------------+
| Trial execution |
| behavior-diff.sh builds independent Before/After copies |
| run-trial.sh runs the same task/model in each copy |
+----------------------------+----------------------------+
v
+---------------------------------------------------------+
| Saved evidence |
| Snapshots, traces, answers, and completion grades |
+----------------------------+----------------------------+
v
+---------------------------------------------------------+
| Evidence analysis |
| decisions.py: model-based decision extraction |
| reporting/: deterministic comparison and report data |
+----------------------------+----------------------------+
v
+---------------------------------------------------------+
| Presentation |
| render.py -> HTML, Markdown, and structured JSON |
| Skill explains evidence; user reviews the change |
+---------------------------------------------------------+
Trial execution and decision extraction call models through agent CLIs. Report assembly and rendering run locally without model calls.
The user can request the skill directly. The plugin can also suggest it after
an agent edits CLAUDE.md, AGENTS.md, or SKILL.md.
The hook configuration connects host events to
scripts in plugin/scripts/:
rules-edit-backup.shsaves Before content for untracked instruction files before supported edits. Tracked files normally use Git HEAD instead.rules-edit-detect.shrecords edited paths and asks the host agent to offer Behavior Diff after the current task.rules-edit-remind.shprovides a fallback reminder without a duplicate offer.
Hooks supply baselines and discovery, not trial execution. A hook suggestion requires explicit user acceptance. Hook failures do not block the host session.
SKILL.md guides the host agent to
select the changed file and prepare a scenario at the relevant decision point.
The task must not reveal the expected behavior. The instruction change must
remain the only source of that guidance.
The skill selects the trial host and model, calls the runner, and explains the result. It owns scenario quality and interpretation, not filesystem copying or report formatting.
behavior-diff.sh
coordinates one comparison. It resolves the Before version from an explicit
file, Git HEAD, or a hook baseline. It creates one project copy per trial and
installs the appropriate instruction version.
For Git projects, copies start from git archive HEAD. Only the selected
instruction change enters the After copies. Other uncommitted edits stay out.
Without a usable Git HEAD, the runner copies the folder instead.
run-trial.sh runs one
fresh agent process inside one copy. The default comparison runs three trials
per side concurrently, with the same task and model. The adapter supports
Claude, Codex, Pi, and OMP. It suppresses recursive plugin hooks and converts
host events into a common trace.jsonl format.
The trace contains tool actions and a final answer. Raw host logs and stderr
remain available where captured. After all trials finish, the coordinator
writes grades.tsv: REVIEW for a trace with a final answer, otherwise
BLOCKED. This checks completeness, not correctness.
Trial and extraction models have separate roles. Claude defaults to Opus for
trials and Sonnet for extraction. Codex uses Sol and Luna, respectively.
codex_model.py resolves Codex selectors to concrete IDs. Both trial variants
use the same resolved ID. Model selection
describes overrides and Pi/OMP behavior.
A scenario combines a task, project state, and an instruction version. The skill prepares the task. The runner constructs the project state and instruction versions without asking a model to rewrite either version.
- Choose the Before instruction. The runner uses
--before-filefirst. Otherwise, an existing target in Git HEAD supplies the Before version. The remaining case uses the newest hook baseline. A baseline markedABSENTmeans the Before copy must not contain the file. Without a baseline, the runner stops rather than inventing an original. - Build a fresh base for every trial. In a Git project,
git archive HEADexports committed files intobefore-N/project/orafter-N/project/. These are full copies, not branches or worktrees. Outside Git, a folder copy excludes.gitand Behavior Diff storage. - Install the instruction version. Before retains the HEAD version,
receives the explicit/baseline file, or removes the file for
ABSENT. After always receives the selected file from the working directory. All other initial project files remain the same between sides. - Run the same task independently. Each copy gets a new Git repository
and an initial snapshot commit.
run-trial.shreceives that project directory, the sharedtask.md, and the selected model. The runner starts all trials as background processes and waits before analysis.
Each side contains N independent copies, not one shared directory:
+-----------------------------+
| Common project base |
| Git HEAD or folder copy |
+--------------+--------------+
|
+-----------------+---------------+
v v
+-----------------------------+ +-----------------------------+
| Before: before-1..N/project | | After: after-1..N/project |
| Resolved Before instruction | | Working-tree instruction |
| or no file for ABSENT | | Selected file only |
+-----------------------------+ +-----------------------------+
For example, a tracked instruction edit produces this initial state:
| Input | Before | After |
|---|---|---|
| Project files | Git HEAD | Git HEAD |
| Selected instruction | HEAD version | Working-directory version |
| Task | Shared task.md |
Shared task.md |
| Trial model | Selected model | Same selected model |
| Agent context and files | Fresh for each trial | Fresh for each trial |
The run keeps task.md and config.json at its root. Each before-N/ and
after-N/ directory contains project/, trace.jsonl, and stderr.log.
The configuration records the target path, scenario, and Before/After labels.
Uncommitted setup files do not enter Git-based trials. Required task state must
already exist in the base project or in a deliberately prepared fixture.
One trial is shown below. The coordinator repeats this sequence for all
2 x N trials, with concurrent execution.
behavior-diff.sh run-trial.sh Agent CLI
| | |
|-- start ----------->| |
| |-- task + model ---->|
| |< host events -------|
| | |
| | capture/normalize trace.jsonl
|< process ends ------|
|
| wait for all 2 x N trials
| write grades.tsv
| call decisions.py, then render.py
v
The modules below live in the skill's scripts/ directory.
- Model-based interpretation:
decisions.pyreads the task, trial actions, final answers, and numbered instruction-diff hunks. A separate model extracts choices, counts, a primary result, supported implications, links to edits, plain-language Summary text, and the edit's likely aim in the same call. The inferred aim cites instruction hunks independently of trial outcomes. The script validates references and exact choice coverage before writingdecisions.json. Invalid interpretation does not discard valid observations. - Deterministic assembly:
reporting/load.pyreads saved evidence and compares recorded command sequences.reporting/instruction.pysupplies the instruction diff.reporting/content.pyderives shared wording and evidence limits.reporting/summary.pyvalidates narrative and selects the visual lead, retaining mixed-result, incomplete-evidence, and single-trial cautions. It also derives the story's aim from supplied expected behavior or a validated inference whose stored diff matches the current instruction diff. Together these modules build schema-v7ReportDatainreporting/schema.py. Summary counts come from existing decision rows.
Command flow comes from recorded events. Decision comparisons come from model interpretation of those events and answers. The report keeps these sources distinct. Links between decisions and edits do not prove causality.
render.py passes the same ReportData to reporting/render_markdown.py and
reporting/render_html.py. It writes four artifacts:
| Artifact | Purpose |
|---|---|
report.html |
Local browser report with Summary, Decision diff, Flow diff, and Trial evidence views. |
report.md |
Markdown version for reading and sharing after review. |
report-data.json |
Structured, versioned report data. |
report-artifact.html |
Embeddable HTML body. |
The default Summary tells a numbered story: the edit's aim, the scenario's
observations, and their meaning. The goal is one sentence with a compact
source label and edit link; an info popup holds the source caveat. Supplied
expectations and inferred aims remain distinct; neither becomes proof of
author intent or goal completion. Markdown retains the caveat as plain text.
reporting/illustrations.py supplies fixed SVG shapes for the Before/After
comparison. Model output supplies text and selectors, never markup.
Plans are not presented as executions. content.additional_findings selects
at most three compact comparisons; full decisions stay in Decision diff.
Markdown shares the same story and findings without illustrations.
Saved decisions without interpretations show explicit availability notices;
rendering never requests new explanations.
The scenario disclosure explains the simulated situation, instruction versions,
trial setup, and supplied expectation (or its absence). A nested disclosure
holds the full original scenario prompt, closed by default.
The loader stores content.task separately from the scenario description,
preferring the run's task.md over a capsule copy. Missing task text remains
unavailable rather than treating a description as the prompt. Shared
content.scenario_sections supplies both HTML and Markdown; rendering adds
no inferred setup details.
Instruction edit begins with a plain-language explanation from the saved
intent, labeled as a likely aim or supplied expectation. Rendering does not
infer missing explanations or claim confirmed author intent. It contains one
complete diff, with no duplicate excerpt or nested diff toggle.
content.instruction_edit_counts counts added and removed
lines across validated hunks. Stable diff and hunk anchors open the enclosing
disclosure. Missing or unparseable diffs do not acquire invented counts.
Both output formats retain the full evidence.
The runner attempts to open the HTML report. The skill summarizes observed differences and evidence limits in the conversation. Missing extraction does not prevent a report of commands and answers. Saved evidence can be rendered again without new trials or model calls.
The runner calls the renderer after decision extraction:
python3 "$scripts/render.py" "$run" "$run" "$model" "$run/config.json"Here, $scripts is the bundled scripts directory and $run is the saved run.
The second directory argument supplies scenario assets, such as a fallback
rule.md. The standard runner uses the run directory for both arguments.
All output formats share ReportData. Only the HTML branch uses
report.css and the document wrapper:
+-----------------------------------------------------+
| Saved run: configuration, grades, traces, snapshots |
| Optional decisions.json |
+--------------------------+--------------------------+
|
v
load_report()
|
v
+------------+
| ReportData |
+------+-----+
|
+--> to_json() -> report-data.json
|
+--> render_markdown() -> report.md
|
+--> render_artifact(report, css)
|
+--> report-artifact.html
|
+--> render_document()
|
v
report.html
- Load evidence.
load_report()readsconfig.json,grades.tsv, each trial'strace.jsonl, and optionaldecisions.json. It groups trials by Before/After and extracts actions and final answers. - Recover the instruction diff.
reporting/instruction.pyreads the target file frombefore-1/project/andafter-1/project/. Python'sdifflib.unified_diffproduces the added and removed lines. The report therefore uses saved trial files, not the current source project. - Build one report object. The loader computes command comparisons and
assembles metadata, results, decisions, instruction diff, and trial evidence
into
ReportData. Its schema checks the structure before rendering. - Build the HTML body.
render_artifact(report, css)creates HTML from that object and inlinesreporting/report.css. Python functions build the Summary, decision comparisons, command flow, and expandable trial evidence. Decision diff appears only with extracted decisions. Evidence text is HTML-escaped, so recorded code and answers appear as text rather than markup. - Create the page.
render_document(artifact)adds the HTML document wrapper, character encoding, and viewport metadata.render.pywrites the body toreport-artifact.htmland the full document toreport.html. It also serializesReportDatato JSON and renders Markdown separately.
HTML is not converted from Markdown, and a model does not write the page. The page includes JavaScript for evidence links, expand/collapse controls, and printing. It opens as a local file without an application server. Styles are inline, but the page also references Google Fonts.
State lives under ${BEHAVIOR_DIFF_HOME:-~/.behavior-diff}/:
baselines/: Before content for instruction files without a Git baseline.nudge/: session state that prevents repeated hook suggestions.runs/: each comparison's task, configuration, project copies, traces, completion grades, extracted decisions, and generated reports.
The runner prepares copies without changing the source project. Separate working directories are not a universal security sandbox. Tool restrictions depend on the host, and permitted shell commands can have side effects.
Reports stay local, but trial and extraction inputs go to the selected model provider. Reports can contain private code and answers. CI uses deterministic checks and synthetic evidence, never live model runs.