Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
Expand Up @@ -462,9 +462,11 @@ def _blinded_report(source):
("p", "summary-provenance"),
("p", "summary-status"),
("div", "summary-pair"),
("nav", "evidence-nav"),
]
)
if any(node.has_class("primary-result-context") for node in _children(bodies[1])):
evidence_shape.append(("section", "primary-result-context"))
evidence_shape.append(("nav", "evidence-nav"))
evidence_body = _shape(bodies[1], evidence_shape)
if _plain(evidence_body[0]) != "What the evidence shows":
raise ValueError("Unknown evidence heading.")
Expand Down
21 changes: 15 additions & 6 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -157,9 +157,11 @@ one side. Each run creates a local HTML report with five tabs:

- **Summary** tells a numbered story: the intended change, what the evidence
shows, and what it means. Illustrated Before/After cards retain exact counts
and distinguish plans, answers, and recorded actions. Mixed results and
evidence gaps remain visible. **View this comparison** opens the supporting
comparison, or links to trial records when extraction is unavailable.
and distinguish plans, answers, and recorded actions. When the cards focus on
an explanation or another secondary comparison, the primary result and its
full Before/After distribution remain visible beside them. Mixed results and
evidence gaps remain visible. One **View behavior comparisons** button opens
Behavior diff for both the featured comparison and the primary result.
**Full scenario and expected behavior** explains the simulated situation,
instruction versions, trial setup, and supplied expectation or its absence.
**View full scenario prompt** is nested inside that disclosure.
Expand Down Expand Up @@ -214,8 +216,11 @@ come from recorded commands or self-reported actions; **Answer detail** comparis
come from the final answer. Wording differences alone do not establish an action change.

In Summary and Behavior diff, counts such as **3 of 3 trials** refer to trials,
not repeated actions within one trial. A model extracts these counts from the
evidence. Separate row counts do not show a complete sequence within one trial.
not repeated actions within one trial. For new extractions, the model assigns
named trials to behaviors; code derives counts from complete, unique membership
on each side. Those assignments remain in `decisions.json` for auditing.
This validates bookkeeping, not whether a trial was classified correctly.
Separate row counts do not show a complete sequence within one trial.
**Changed** compares extracted behavior proportions. **Unchanged** is shown for
complete, unanimous same-behavior evidence; matching proportions without that
evidence are labeled **Same proportions**. **Unavailable** means extracted
Expand All @@ -233,7 +238,7 @@ The command-category table groups each trial into one category combination.
Category order is not execution order. Matching categories can contain different
commands or files, so matching patterns do not establish unchanged behavior.

Follow **View this comparison** to Behavior diff, or inspect the trial records
Follow **View behavior comparisons** to Behavior diff, or inspect the trial records
for original commands and answers. Markdown retains the five sections, story,
counts, complete instruction diff, and grouped trial evidence without illustrations.

Expand All @@ -243,6 +248,10 @@ interprets their relationship to the supplied edit. **Related edit** links are
labeled as model interpretation, not causal proof or knowledge of author intent.
The same call supplies optional plain-language Summary text: a takeaway, short
scenario, Before/After descriptions, and supported implications or cautions.
The writing instructions require concrete actor/action/object contrasts, parallel
Before/After descriptions, and a clear distinction between changed decisions or
actions and changed explanations, citations, or presentation. The decisive
contrast belongs in the headline and cards, not only in an additional finding.
It also interprets the edit's likely aim, citing instruction-diff hunks rather
than inferring intent from trial outcomes. **Edit goal** shows one sentence
with an **Inferred** badge and an edit link; the info popup explains the source
Expand Down
32 changes: 24 additions & 8 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,7 +35,7 @@ scripts handle execution, evidence, and reporting.
v
+---------------------------------------------------------+
| Evidence analysis |
| decisions.py: comparisons and trial summaries |
| decisions.py: trial assignments -> counts and summaries |
| reporting/: deterministic comparison and report data |
+----------------------------+----------------------------+
v
Expand Down Expand Up @@ -189,13 +189,18 @@ The modules below live in the skill's [`scripts/`](../plugin/skills/behavior-dif

- **Model-based interpretation:** `decisions.py` reads the task, trial actions,
final answers, and numbered instruction-diff hunks. A separate model extracts
choices, counts, a primary result, supported implications, links to edits,
plain-language Summary text, the edit's likely aim, and short per-group trial
summaries in the same call. Each trial summary names its exact Before/After
records and separates answers and plans from recorded actions.
The inferred aim cites instruction hunks independently of trial outcomes.
The script validates references and exact choice coverage before writing
`decisions.json`. Invalid interpretation does not discard valid observations.
choices with named trial assignments, a primary result, supported implications,
links to edits, plain-language Summary text, the edit's likely aim, and short
per-group trial summaries in the same call. New raw branches supply `trials`,
not aggregate counts. The script requires each completed trial exactly once
per side of every row, derives `n`, and retains both in `decisions.json`.
Foreign, duplicate, incomplete, or missing assignments invalidate the row.
Count-only raw extraction is not accepted; saved normalized report evidence
remains renderable. Membership validation cannot prove classification truth.
Each trial summary names its exact Before/After records and separates answers
and plans from recorded actions. The inferred aim cites instruction hunks
independently of trial outcomes. The script validates references and exact
choice coverage; invalid interpretation does not discard valid observations.
- **Deterministic assembly:** `reporting/load.py` reads saved evidence and
compares recorded command sequences. `reporting/instruction.py` supplies the
instruction diff. `reporting/content.py` derives shared wording and evidence
Expand Down Expand Up @@ -238,6 +243,17 @@ expectations and inferred aims remain distinct; neither becomes proof of
author intent or goal completion. Markdown retains the caveat as plain text.
`reporting/illustrations.py` supplies fixed SVG shapes for the Before/After
comparison. Model output supplies text and selectors, never markup.
Summary and trial-summary instructions require concrete, parallel descriptions
of the same subject, with the decisive contrast and any unchanged decision
explicitly stated. Changed explanations, citations, or presentation must not
be described as changed actions. When the selected lead is not the primary
result, `content.primary_result_context` exposes the primary status and full
distribution beside the cards in both formats. Missing or incomplete evidence
cannot become an unchanged-result claim. This is derived presentation, not a
new serialized report field; schema v8 is unchanged.
The primary-result block is context, not a second navigation choice. One
**View behavior comparisons** link opens Behavior diff for the selected
comparison and the primary result; unavailable evidence links to trial records.
Plans are not presented as executions. `content.additional_findings` selects
at most three compact comparisons; full comparisons stay in Behavior diff.
Markdown shares the same story and findings without illustrations.
Expand Down
Loading
Loading