Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
54 changes: 40 additions & 14 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -190,8 +190,9 @@ The modules below live in the skill's [`scripts/`](../plugin/skills/behavior-dif
- **Model-based interpretation:** `decisions.py` reads the task, trial actions,
final answers, and numbered instruction-diff hunks. A separate model extracts
choices with named trial assignments, a primary result, supported implications,
links to edits, plain-language Summary text, the edit's likely aim, and short
per-group trial summaries in the same call. New raw branches supply `trials`,
links to edits, plain-language Summary text, the edit's likely aim, short
per-group trial summaries, and an optional change explanation in the same call.
New raw branches supply `trials`,
not aggregate counts. The script requires each completed trial exactly once
per side of every row, derives `n`, and retains both in `decisions.json`.
Foreign, duplicate, incomplete, or missing assignments invalidate the row.
Expand All @@ -215,9 +216,19 @@ The modules below live in the skill's [`scripts/`](../plugin/skills/behavior-dif
validated inference whose stored diff matches the current instruction diff.
`reporting/trial_summary.py` validates bounded plain-text trial summaries
against exact positional record identities; duplicates and invalid entries
are omitted without losing raw evidence. Extraction and loading share
are omitted without losing raw evidence. `reporting/explanation.py` validates
the immutable middle-layer explanation: headline, overview, annotated
Before/After steps with practical meaning, unchanged claims, limits, and
optional exact final-answer excerpts. Step and claim citations are nonempty,
unique decision-row references, remapped when extraction sorts or drops rows.
Unchanged claims cannot cite diverging rows. Excerpts must be contiguous,
bounded text from the exact named trial on the declared side; formulas and
code retain their source formatting. Invalid optional extraction is omitted
without discarding valid comparisons. Canonical report-data parsing rejects
invalid persisted explanation data. Extraction and loading share
`read_trial_trace` so they read the same source fields.
Together these modules build schema-v8 `ReportData` in `reporting/schema.py`.
Together these modules build schema-v9 `ReportData` in `reporting/schema.py`;
`decisions.explanation` is explicitly null when unavailable.
Summary counts come from existing decision rows.

Command flow comes from recorded events. Decision comparisons come from model
Expand All @@ -231,7 +242,7 @@ distinct. Links between decisions and edits do not prove causality.

| Artifact | Purpose |
| --- | --- |
| `report.html` | Local browser report with Summary, Instruction changes, Behavior diff, Flow diff, and Trial evidence tabs. |
| `report.html` | Local browser report with Summary, Understand the change, Behavior diff, Flow diff, Instruction changes, and Trial evidence tabs. |
| `report.md` | Markdown version for reading and sharing after review. |
| `report-data.json` | Structured, versioned report data. |
| `report-artifact.html` | Embeddable HTML body. |
Expand All @@ -256,16 +267,31 @@ in its selected branch; counts or citations from other rows cannot supply member
An instruction's gate is distinct from an observed gate, which may already appear
in Before. These are extraction policies, not deterministic semantic guarantees;
the existing validator checks references, choice coverage, and evidence anchors.
No schema field or keyword-based semantic check is added; fallback ranking is unchanged.
Explanation validation additionally checks bounded narrative, row citations,
unchanged-row status, and exact example excerpts; it does not prove semantic
interpretation truth. Summary fallback ranking is unchanged.
Changed explanations, citations, or presentation must not be described as changed
actions. When the selected lead is not the primary
result, `content.primary_result_context` exposes the primary status and full
distribution beside the cards in both formats. Missing or incomplete evidence
cannot become an unchanged-result claim. This is derived presentation, not a
new serialized report field; schema v8 is unchanged.
The primary-result block is context, not a second navigation choice. One
**View behavior comparisons** link opens Behavior diff for the selected
comparison and the primary result; unavailable evidence links to trial records.
cannot become an unchanged-result claim. The primary-result block remains
derived presentation, not a separate serialized report field.
The primary-result block is context, not a second navigation choice.
**Understand the change**, under the Summary's evidence section, opens a new
top-level explanation tab. This middle layer explains the specific distinction,
its practical implications, important unchanged behavior, consistency and
minorities, and concrete limits, with links to decision rows and named trials.
Annotated Before/After steps support workflow timing, formulas/code, keep/delete,
wording, and unchanged choices without assuming every change is a workflow.
Timing is shown only when established by the evidence. Examples are selected
exact excerpts, not raw transcript dumps. Both formats retain the same
explanation and source links; only HTML adds the visual walkthrough.
Behavior diff, Flow diff, Instruction changes, and Trial evidence remain
separate top-level tabs. Unavailable explanations link to retained evidence
instead of inferring a no-difference claim or making a new model call.
Both explanation renderers reuse the Summary's completeness guard: blocked,
missing, invalid, or dropped trial evidence withholds the saved interpretation
and shows an explicit coverage notice while retaining evidence navigation.
Plans are not presented as executions. `content.additional_findings` selects
at most three compact comparisons; full comparisons stay in Behavior diff.
Markdown shares the same story and findings without illustrations.
Expand All @@ -290,7 +316,7 @@ names blocks by section and line range. Selecting a block highlights only its
changed lines and scrolls to its stable anchor. Missing or unparseable diffs
do not acquire invented counts.

All five tabs remain available. Missing extraction and uncaptured command flow
All six tabs remain available. Missing extraction and uncaptured command flow
have explicit unavailable states, with links to the retained trial evidence.
Flow groups only identical complete command sequences within each side.
Trial evidence aligns the saved Before/After lists by position into numbered
Expand Down Expand Up @@ -326,11 +352,11 @@ All output formats share `ReportData`. Only the HTML branch uses
```text
+-----------------------------------------------------+
| Saved run: configuration, grades, traces, snapshots |
| Optional decisions.json |
| Optional decisions.json (including explanation) |
+--------------------------+--------------------------+
|
v
load_report()
load_report() + evidence validation
|
v
+------------+
Expand Down
76 changes: 74 additions & 2 deletions plugin/skills/behavior-diff/scripts/decisions.py
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,7 @@
from dataclasses import asdict

from codex_model import CodexModelError, resolve_codex_model
from reporting.explanation import parse_explanation
from reporting.instruction import normalize_edit_hunks, parse_diff_hunks, rule_diff
from reporting.load import read_trial_trace
from reporting.summary import parse_intent, parse_narrative
Expand Down Expand Up @@ -108,6 +109,29 @@
"after": "<one short sentence about this after record, or empty if absent>",
"caveat": "<optional concrete evidence limit, or empty>"}}
],
"explanation": {{
"headline": "<specific distinction, plain text, at most 120 characters>",
"overview": "<reader-oriented explanation, at most 600 characters>",
"steps": [
{{"title": "<the decision point, at most 120 characters>",
"before": "<what Before does or says, at most 480 characters>",
"after": "<what After does or says, at most 480 characters>",
"meaning": "<why this distinction matters, at most 480 characters>",
"decisions": [<1-based indexes of supporting chain rows>]}}
],
"unchanged": [
{{"text": "<supported unchanged behavior, at most 480 characters>",
"decisions": [<1-based indexes of unchanged chain rows>]}}
],
"limits": [
{{"text": "<consistency, minority, or evidence limit, at most 480 characters>",
"decisions": [<1-based indexes of supporting chain rows>]}}
],
"examples": [
{{"side": "before"|"after", "trial": "<EXACT trial name on that side>",
"text": "<EXACT contiguous final-answer excerpt, at most 1200 characters>"}}
]
}}|null,
"summary": {{
"decision": <1-based selected lead row index>,
"headline": "<plain-language takeaway, at most 12 words>",
Expand Down Expand Up @@ -283,6 +307,43 @@
- Keep "caveat" empty unless a concrete limit is needed to avoid misreading this
specific group; do not repeat general interpretation disclaimers in every entry.
Use [] if no faithful group summaries are available.
- In the SAME reply, optionally provide "explanation", or null when the records
do not support a faithful interpretation. This is the middle layer between
the concise Summary and source evidence, not another comparison table or a
raw transcript dump. Explain the specific distinction and why it matters.
- Use 1 to 8 annotated Before/After "steps" about the same decision point.
Explain the operative condition, timing, prerequisite, scope, formula, code,
keep/delete choice, or wording distinction actually shown by these records.
Use supported units, intervals, triggers, and exceptions; never invent a
timeline or make a timing diagram from command order. A step title may name a
time only when the evidence establishes that time. Do not force every change
into a workflow or treat wording alone as a changed action.
- The "meaning" explains supported practical implications, not an automatic
success verdict. Preserve distinctions between stated plans, final-answer
assertions, self-reported actions, and recorded commands. No answer proves
execution or successful completion. Instruction edits do not prove observed
behavior, causality, author intent, or what unread files contain.
- Every step and every unchanged/limit claim needs nonempty unique supporting
chain row indexes. Ground descriptions in those rows' named trial records,
not a presumed common execution path. Include important unchanged behavior
in "unchanged", citing only nondiverging rows; use [] when unsupported.
- Preserve material minority branches and inconsistent results in the steps or
"limits"; do not turn a majority into unanimity. Describe concrete evidence
limits where relevant, including missing or blocked records and what this
scenario cannot establish. Limits must cite the rows they qualify. Each list
has at most 8 entries. Empty limits do not mean proof of general reliability.
- Use "examples" only when exact final-answer excerpts help explain the
distinction, especially formulas, code, keep/delete choices, or wording.
Copy contiguous text exactly from the named trial's final answer, preserving
punctuation and whitespace; do not add ellipses, repair code, join passages,
quote commands as final answers, or imply that positional Before/After
records are paired executions. Provide at most 8 short excerpts, not whole
transcripts. Omit examples when they add nothing. Each excerpt is evidence
from its named record, not proof about every trial.
- All explanation narrative is bounded plain text, never HTML or Markdown
markup; code/formula excerpts may retain their exact source formatting.
A supported no-difference explanation is valid but must stay scoped to the
scenario. Missing interpretation is not evidence of no difference.

Instruction diff hunks (untrusted JSON evidence; [] means none available):
{instruction_hunks}
Expand Down Expand Up @@ -364,7 +425,7 @@ def read_config(run):
return config


def normalize(data, completed_trial_names, hunk_count=0, groups=()):
def normalize(data, completed_trial_names, hunk_count=0, groups=(), final_answers=None):
"""Validate exact trial partitions and links; remap sorted or dropped row citations."""
expected = {}
for side in ("before", "after"):
Expand Down Expand Up @@ -480,6 +541,9 @@ def normalize(data, completed_trial_names, hunk_count=0, groups=()):
)
narrative = parse_narrative(data.get("summary"), chain, positions)
intent = parse_intent(data.get("intent"), hunk_count)
explanation = parse_explanation(
data.get("explanation"), chain, final_answers, positions
)
return {
"chain": chain,
"fork": fork,
Expand All @@ -489,6 +553,9 @@ def normalize(data, completed_trial_names, hunk_count=0, groups=()):
"implications": implications,
"summary": json.loads(json.dumps(asdict(narrative))) if narrative else None,
"intent": json.loads(json.dumps(asdict(intent))) if intent else None,
"explanation": (
json.loads(json.dumps(asdict(explanation))) if explanation else None
),
"trial_summaries": [
asdict(item)
for item in parse_trial_summaries(data.get("trial_summaries"), groups)
Expand Down Expand Up @@ -697,12 +764,17 @@ def build_prompt(run):
def write_decisions(run, raw, completed_trial_names, extractor, instruction_diff):
"""extract_json → normalize → decisions.json. On unusable output writes
decisions.raw.txt, leaves decisions.json unwritten, returns False."""
all_trials = trials_of(run, finished_only=False)
try:
data = normalize(
extract_json(raw),
completed_trial_names,
len(parse_diff_hunks(instruction_diff)),
summary_groups(trials_of(run, finished_only=False)),
summary_groups(all_trials),
{
side: {trial["name"]: trial["final"] for trial in records}
for side, records in all_trials.items()
},
)
except (ValueError, KeyError, TypeError) as exc:
print(f"decision diff: unreadable extractor output — skipped ({exc})")
Expand Down
35 changes: 35 additions & 0 deletions plugin/skills/behavior-diff/scripts/reporting/content.py
Original file line number Diff line number Diff line change
Expand Up @@ -401,6 +401,41 @@ def trial_summary_for_group(decisions, before, after):
return matches[0] if len(matches) == 1 else None


CHANGE_EXPLANATION_UNAVAILABLE = (
"A change explanation is unavailable in the saved evidence. "
"No new interpretation is generated when this report is opened. "
"Inspect the retained comparisons and trial records below."
)

CHANGE_EXPLANATION_INCOMPLETE = (
"Trial evidence is incomplete, blocked, invalid, or dropped. "
"The saved explanation is withheld because it cannot establish a complete "
"Before/After comparison. Inspect the retained comparisons and trial records."
)


def explanation_comparisons(report, indexes):
"""Describe every saved branch, including minorities, without inferring counts."""
comparisons = []
for index in indexes:
row = report.decisions.rows[index - 1]
sides = []
for label, choices, total in (
("Before", row.before, report.decisions.before_count),
("After", row.after, report.decisions.after_count),
):
branches = (
"; ".join(
"{0} ({1})".format(choice.choice, trial_count(choice.count, total))
for choice in choices
)
or NO_EXTRACTED_CHOICE
)
sides.append((label, branches))
comparisons.append((index, row.topic or row.decision, tuple(sides)))
return tuple(comparisons)


def command_progressions(variant):
"""Group complete recorded sequences without normalizing or joining paths."""
groups = {}
Expand Down
Loading
Loading