Repository navigation
feat: guided report story and human-readable trial summaries - #23
Conversation
Signed-off-by: Kent Huang <kent@infuseai.io>
Independent review resultsTwo independent read-only subagents reviewed the implementation now included in this PR (commit Standards and correctness — APPROVE0 material findings. Reviewed the extraction → validation → loading → schema → HTML/Markdown pipeline. Exact trial identities, malformed and duplicate summaries, unequal sides, missing-summary states, and schema-v8 serialization preserve the underlying evidence. Escaping was checked in its actual HTML and Markdown output contexts. No new dependencies or additional model-call path were introduced. Spec and acceptance — APPROVE0 material findings. The implementation provides one short takeaway, Before/After descriptions, and an optional concrete caveat for each trial group. Summaries remain specific to that group rather than borrowing aggregate conclusions. The served mixed-case Markdown correctly distinguishes changed trials 1–2 from unchanged trial 3. Planned-action examples do not claim execution. Full answers and supporting evidence remain available, and documentation matches the implemented contract. Review limitsThese reviews inspected source, tests, synthetic fixtures, and served static content. The reviewers did not run tests, interactive browser checks, formatters, or live models, and made no edits. The implementation author’s prior deterministic and browser verification is documented separately in the PR description. Live-generated summary prose quality remains unvalidated; existing verification used synthetic evidence. Summary: 0 standards findings, 0 spec findings; both reviewers approved. |
Summary
Data contract
Report data advances to schema v8 for typed
decisions.trial_summaries. Regenerate reports from original run artifacts; older serialized report-data files are rejected. Existing decisions.json without summaries remains renderable with explicit unavailability. Rendering never calls a model. No new dependency or additional model-call stage.Verification
Passed locally before commit:
bash tests/hooks-test.shpython3 plugin/skills/behavior-diff/scripts/decisions.py --checkpython3 tests/report-schema-test.pybash tests/live-report-contract.shbash tests/release-workflow-test.shgit diff --checkGenerated all synthetic gallery scenarios through the real ingestion/rendering pipeline. Browser smoke covered desktop/mobile trial alignment and summaries, grouped navigation, instruction highlighting, hover/keyboard/touch legends, mixed/blocked/missing/self-reported/planned-action evidence, and print disclosure-state restoration.
Two independent read-only reviews returned APPROVE with no material findings on both standards/correctness and spec/acceptance. Live-generated summary prose quality remains untested: verification used synthetic evidence, with no live model calls.
Tracking
DRC-4793 — Implement the approved guided-story Behavior Diff report
DCO-signed commit. No version bump or release included.