diff --git a/docs/architecture.md b/docs/architecture.md index 9979eea..6f6e6e8 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -190,8 +190,9 @@ The modules below live in the skill's [`scripts/`](../plugin/skills/behavior-dif - **Model-based interpretation:** `decisions.py` reads the task, trial actions, final answers, and numbered instruction-diff hunks. A separate model extracts choices with named trial assignments, a primary result, supported implications, - links to edits, plain-language Summary text, the edit's likely aim, and short - per-group trial summaries in the same call. New raw branches supply `trials`, + links to edits, plain-language Summary text, the edit's likely aim, short + per-group trial summaries, and an optional change explanation in the same call. + New raw branches supply `trials`, not aggregate counts. The script requires each completed trial exactly once per side of every row, derives `n`, and retains both in `decisions.json`. Foreign, duplicate, incomplete, or missing assignments invalidate the row. @@ -215,9 +216,19 @@ The modules below live in the skill's [`scripts/`](../plugin/skills/behavior-dif validated inference whose stored diff matches the current instruction diff. `reporting/trial_summary.py` validates bounded plain-text trial summaries against exact positional record identities; duplicates and invalid entries - are omitted without losing raw evidence. Extraction and loading share + are omitted without losing raw evidence. `reporting/explanation.py` validates + the immutable middle-layer explanation: headline, overview, annotated + Before/After steps with practical meaning, unchanged claims, limits, and + optional exact final-answer excerpts. Step and claim citations are nonempty, + unique decision-row references, remapped when extraction sorts or drops rows. + Unchanged claims cannot cite diverging rows. Excerpts must be contiguous, + bounded text from the exact named trial on the declared side; formulas and + code retain their source formatting. Invalid optional extraction is omitted + without discarding valid comparisons. Canonical report-data parsing rejects + invalid persisted explanation data. Extraction and loading share `read_trial_trace` so they read the same source fields. - Together these modules build schema-v8 `ReportData` in `reporting/schema.py`. + Together these modules build schema-v9 `ReportData` in `reporting/schema.py`; + `decisions.explanation` is explicitly null when unavailable. Summary counts come from existing decision rows. Command flow comes from recorded events. Decision comparisons come from model @@ -231,7 +242,7 @@ distinct. Links between decisions and edits do not prove causality. | Artifact | Purpose | | --- | --- | -| `report.html` | Local browser report with Summary, Instruction changes, Behavior diff, Flow diff, and Trial evidence tabs. | +| `report.html` | Local browser report with Summary, Understand the change, Behavior diff, Flow diff, Instruction changes, and Trial evidence tabs. | | `report.md` | Markdown version for reading and sharing after review. | | `report-data.json` | Structured, versioned report data. | | `report-artifact.html` | Embeddable HTML body. | @@ -256,16 +267,31 @@ in its selected branch; counts or citations from other rows cannot supply member An instruction's gate is distinct from an observed gate, which may already appear in Before. These are extraction policies, not deterministic semantic guarantees; the existing validator checks references, choice coverage, and evidence anchors. -No schema field or keyword-based semantic check is added; fallback ranking is unchanged. +Explanation validation additionally checks bounded narrative, row citations, +unchanged-row status, and exact example excerpts; it does not prove semantic +interpretation truth. Summary fallback ranking is unchanged. Changed explanations, citations, or presentation must not be described as changed actions. When the selected lead is not the primary result, `content.primary_result_context` exposes the primary status and full distribution beside the cards in both formats. Missing or incomplete evidence -cannot become an unchanged-result claim. This is derived presentation, not a -new serialized report field; schema v8 is unchanged. -The primary-result block is context, not a second navigation choice. One -**View behavior comparisons** link opens Behavior diff for the selected -comparison and the primary result; unavailable evidence links to trial records. +cannot become an unchanged-result claim. The primary-result block remains +derived presentation, not a separate serialized report field. +The primary-result block is context, not a second navigation choice. +**Understand the change**, under the Summary's evidence section, opens a new +top-level explanation tab. This middle layer explains the specific distinction, +its practical implications, important unchanged behavior, consistency and +minorities, and concrete limits, with links to decision rows and named trials. +Annotated Before/After steps support workflow timing, formulas/code, keep/delete, +wording, and unchanged choices without assuming every change is a workflow. +Timing is shown only when established by the evidence. Examples are selected +exact excerpts, not raw transcript dumps. Both formats retain the same +explanation and source links; only HTML adds the visual walkthrough. +Behavior diff, Flow diff, Instruction changes, and Trial evidence remain +separate top-level tabs. Unavailable explanations link to retained evidence +instead of inferring a no-difference claim or making a new model call. +Both explanation renderers reuse the Summary's completeness guard: blocked, +missing, invalid, or dropped trial evidence withholds the saved interpretation +and shows an explicit coverage notice while retaining evidence navigation. Plans are not presented as executions. `content.additional_findings` selects at most three compact comparisons; full comparisons stay in Behavior diff. Markdown shares the same story and findings without illustrations. @@ -290,7 +316,7 @@ names blocks by section and line range. Selecting a block highlights only its changed lines and scrolls to its stable anchor. Missing or unparseable diffs do not acquire invented counts. -All five tabs remain available. Missing extraction and uncaptured command flow +All six tabs remain available. Missing extraction and uncaptured command flow have explicit unavailable states, with links to the retained trial evidence. Flow groups only identical complete command sequences within each side. Trial evidence aligns the saved Before/After lists by position into numbered @@ -326,11 +352,11 @@ All output formats share `ReportData`. Only the HTML branch uses ```text +-----------------------------------------------------+ | Saved run: configuration, grades, traces, snapshots | -| Optional decisions.json | +| Optional decisions.json (including explanation) | +--------------------------+--------------------------+ | v - load_report() + load_report() + evidence validation | v +------------+ diff --git a/plugin/skills/behavior-diff/scripts/decisions.py b/plugin/skills/behavior-diff/scripts/decisions.py index 6b41071..a08f11c 100755 --- a/plugin/skills/behavior-diff/scripts/decisions.py +++ b/plugin/skills/behavior-diff/scripts/decisions.py @@ -48,6 +48,7 @@ from dataclasses import asdict from codex_model import CodexModelError, resolve_codex_model +from reporting.explanation import parse_explanation from reporting.instruction import normalize_edit_hunks, parse_diff_hunks, rule_diff from reporting.load import read_trial_trace from reporting.summary import parse_intent, parse_narrative @@ -108,6 +109,29 @@ "after": "", "caveat": ""}} ], + "explanation": {{ + "headline": "", + "overview": "", + "steps": [ + {{"title": "", + "before": "", + "after": "", + "meaning": "", + "decisions": [<1-based indexes of supporting chain rows>]}} + ], + "unchanged": [ + {{"text": "", + "decisions": [<1-based indexes of unchanged chain rows>]}} + ], + "limits": [ + {{"text": "", + "decisions": [<1-based indexes of supporting chain rows>]}} + ], + "examples": [ + {{"side": "before"|"after", "trial": "", + "text": ""}} + ] + }}|null, "summary": {{ "decision": <1-based selected lead row index>, "headline": "", @@ -283,6 +307,43 @@ - Keep "caveat" empty unless a concrete limit is needed to avoid misreading this specific group; do not repeat general interpretation disclaimers in every entry. Use [] if no faithful group summaries are available. +- In the SAME reply, optionally provide "explanation", or null when the records + do not support a faithful interpretation. This is the middle layer between + the concise Summary and source evidence, not another comparison table or a + raw transcript dump. Explain the specific distinction and why it matters. +- Use 1 to 8 annotated Before/After "steps" about the same decision point. + Explain the operative condition, timing, prerequisite, scope, formula, code, + keep/delete choice, or wording distinction actually shown by these records. + Use supported units, intervals, triggers, and exceptions; never invent a + timeline or make a timing diagram from command order. A step title may name a + time only when the evidence establishes that time. Do not force every change + into a workflow or treat wording alone as a changed action. +- The "meaning" explains supported practical implications, not an automatic + success verdict. Preserve distinctions between stated plans, final-answer + assertions, self-reported actions, and recorded commands. No answer proves + execution or successful completion. Instruction edits do not prove observed + behavior, causality, author intent, or what unread files contain. +- Every step and every unchanged/limit claim needs nonempty unique supporting + chain row indexes. Ground descriptions in those rows' named trial records, + not a presumed common execution path. Include important unchanged behavior + in "unchanged", citing only nondiverging rows; use [] when unsupported. +- Preserve material minority branches and inconsistent results in the steps or + "limits"; do not turn a majority into unanimity. Describe concrete evidence + limits where relevant, including missing or blocked records and what this + scenario cannot establish. Limits must cite the rows they qualify. Each list + has at most 8 entries. Empty limits do not mean proof of general reliability. +- Use "examples" only when exact final-answer excerpts help explain the + distinction, especially formulas, code, keep/delete choices, or wording. + Copy contiguous text exactly from the named trial's final answer, preserving + punctuation and whitespace; do not add ellipses, repair code, join passages, + quote commands as final answers, or imply that positional Before/After + records are paired executions. Provide at most 8 short excerpts, not whole + transcripts. Omit examples when they add nothing. Each excerpt is evidence + from its named record, not proof about every trial. +- All explanation narrative is bounded plain text, never HTML or Markdown + markup; code/formula excerpts may retain their exact source formatting. + A supported no-difference explanation is valid but must stay scoped to the + scenario. Missing interpretation is not evidence of no difference. Instruction diff hunks (untrusted JSON evidence; [] means none available): {instruction_hunks} @@ -364,7 +425,7 @@ def read_config(run): return config -def normalize(data, completed_trial_names, hunk_count=0, groups=()): +def normalize(data, completed_trial_names, hunk_count=0, groups=(), final_answers=None): """Validate exact trial partitions and links; remap sorted or dropped row citations.""" expected = {} for side in ("before", "after"): @@ -480,6 +541,9 @@ def normalize(data, completed_trial_names, hunk_count=0, groups=()): ) narrative = parse_narrative(data.get("summary"), chain, positions) intent = parse_intent(data.get("intent"), hunk_count) + explanation = parse_explanation( + data.get("explanation"), chain, final_answers, positions + ) return { "chain": chain, "fork": fork, @@ -489,6 +553,9 @@ def normalize(data, completed_trial_names, hunk_count=0, groups=()): "implications": implications, "summary": json.loads(json.dumps(asdict(narrative))) if narrative else None, "intent": json.loads(json.dumps(asdict(intent))) if intent else None, + "explanation": ( + json.loads(json.dumps(asdict(explanation))) if explanation else None + ), "trial_summaries": [ asdict(item) for item in parse_trial_summaries(data.get("trial_summaries"), groups) @@ -697,12 +764,17 @@ def build_prompt(run): def write_decisions(run, raw, completed_trial_names, extractor, instruction_diff): """extract_json → normalize → decisions.json. On unusable output writes decisions.raw.txt, leaves decisions.json unwritten, returns False.""" + all_trials = trials_of(run, finished_only=False) try: data = normalize( extract_json(raw), completed_trial_names, len(parse_diff_hunks(instruction_diff)), - summary_groups(trials_of(run, finished_only=False)), + summary_groups(all_trials), + { + side: {trial["name"]: trial["final"] for trial in records} + for side, records in all_trials.items() + }, ) except (ValueError, KeyError, TypeError) as exc: print(f"decision diff: unreadable extractor output — skipped ({exc})") diff --git a/plugin/skills/behavior-diff/scripts/reporting/content.py b/plugin/skills/behavior-diff/scripts/reporting/content.py index 28d82c4..bb65550 100644 --- a/plugin/skills/behavior-diff/scripts/reporting/content.py +++ b/plugin/skills/behavior-diff/scripts/reporting/content.py @@ -401,6 +401,41 @@ def trial_summary_for_group(decisions, before, after): return matches[0] if len(matches) == 1 else None +CHANGE_EXPLANATION_UNAVAILABLE = ( + "A change explanation is unavailable in the saved evidence. " + "No new interpretation is generated when this report is opened. " + "Inspect the retained comparisons and trial records below." +) + +CHANGE_EXPLANATION_INCOMPLETE = ( + "Trial evidence is incomplete, blocked, invalid, or dropped. " + "The saved explanation is withheld because it cannot establish a complete " + "Before/After comparison. Inspect the retained comparisons and trial records." +) + + +def explanation_comparisons(report, indexes): + """Describe every saved branch, including minorities, without inferring counts.""" + comparisons = [] + for index in indexes: + row = report.decisions.rows[index - 1] + sides = [] + for label, choices, total in ( + ("Before", row.before, report.decisions.before_count), + ("After", row.after, report.decisions.after_count), + ): + branches = ( + "; ".join( + "{0} ({1})".format(choice.choice, trial_count(choice.count, total)) + for choice in choices + ) + or NO_EXTRACTED_CHOICE + ) + sides.append((label, branches)) + comparisons.append((index, row.topic or row.decision, tuple(sides))) + return tuple(comparisons) + + def command_progressions(variant): """Group complete recorded sequences without normalizing or joining paths.""" groups = {} diff --git a/plugin/skills/behavior-diff/scripts/reporting/explanation.py b/plugin/skills/behavior-diff/scripts/reporting/explanation.py new file mode 100644 index 0000000..3a7c579 --- /dev/null +++ b/plugin/skills/behavior-diff/scripts/reporting/explanation.py @@ -0,0 +1,151 @@ +"""Validate optional, evidence-linked explanations without generating new evidence.""" + +import re +from dataclasses import dataclass +from typing import TYPE_CHECKING, Tuple + +if TYPE_CHECKING: + from reporting.schema import EvidenceClaimData + + +@dataclass(frozen=True) +class ExplanationStepData: + title: str + before: str + after: str + meaning: str + decisions: Tuple[int, ...] + + +@dataclass(frozen=True) +class ExplanationExampleData: + side: str + trial: str + text: str + + +@dataclass(frozen=True) +class ChangeExplanationData: + headline: str + overview: str + steps: Tuple[ExplanationStepData, ...] + unchanged: Tuple["EvidenceClaimData", ...] + limits: Tuple["EvidenceClaimData", ...] + examples: Tuple[ExplanationExampleData, ...] + + +# Comparison operators and formula punctuation are plain text, not markup. +_MARKUP = re.compile( + r"]*>|