diff --git a/.agents/skills/preview-behavior-diff-summary/SKILL.md b/.agents/skills/preview-behavior-diff-summary/SKILL.md new file mode 100644 index 0000000..c01ee37 --- /dev/null +++ b/.agents/skills/preview-behavior-diff-summary/SKILL.md @@ -0,0 +1,85 @@ +--- +name: preview-behavior-diff-summary +description: Builds temporary, synthetic Behavior Diff summary comparisons using the existing renderer. Use when previewing report hierarchy or navigation changes, comparing outcome-first and rule-first cards, or requesting a before/after UI prototype without live model calls. +--- + +# Preview Behavior Diff summaries + +Compare presentations of the same evidence before changing production behavior. +This is a maintainer workflow, not plugin payload or a human-evaluation run. + +## Quick start + +From the repository root: + +```bash +python3 tests/report-demo.py --serve --port 0 +``` + +Open the printed loopback URL. The existing gallery uses synthetic fixtures, +shipped ingestion, and the shipped renderer; it makes no model calls. Keep the +server running for review. Ctrl+C stops it and removes its temporary directory. + +## Prepare the comparison + +1. State the reader's question and the presentation decision being tested. + Example: “Can the reader see the changed retry interval without expanding + Other findings?” Distinguish report-version labels from instruction + Before/After labels inside each report. +2. Read `tests/report-demo.py`, `tests/report_fixtures.py`, and the affected + renderer/summary contract. Choose representative synthetic evidence. Include + mixed branches or an unchanged result when they matter to the decision. + Never copy private evaluation reports or source excerpts into the checkout. +3. Build paired reports in a fresh temporary directory outside the checkout. + Use the existing fixture builder, ingest command, and renderer invocation + from `build_reports` in `tests/report_fixtures.py`; inspect current signatures + rather than copying an old invocation. For a throwaway driver, put `tests` + on Python's import path and restrict the fixture catalog to the selected + scenario before calling `build_reports` on an empty directory. +4. Copy that generated scenario into two sibling variant directories. Preserve + task, instruction versions, trial names, traces, chain choices/memberships, + and primary outcome. For a narrative comparison, change only the selected + summary row and its supported narrative in each `extraction.json`, then run + the same shipped `decisions.py --ingest` and `render.py` commands in each copy. + Do not alter counts or hide minority branches to improve the illustration. +5. If comparing renderer implementations instead, feed identical saved inputs + to each renderer revision. Do not change renderer and authored narrative at + once unless explicitly comparing both; name every changed variable. + Reject invalid narratives rather than bypassing ingestion validation. + +## Present and verify + +6. Create a temporary HTML wrapper around the two rendered reports, not a new + renderer. Offer side-by-side and full-width views with shareable + `?variant=compare`, `?variant=before`, and `?variant=after` states. Keep a + visible switcher and clearly label the two presentation versions. +7. State on the page: “Authored synthetic illustration; no model was called.” + Do not call an authored baseline a historical report. Show unchanged evidence + and remaining uncertainty, including plans versus executed actions. +8. Serve only the temporary root on loopback, for example: + + ```bash + python3 -m http.server 8767 --bind 127.0.0.1 --directory "$PREVIEW_ROOT" + ``` + + Set `PREVIEW_ROOT` to the fresh directory from step 3; select another local + port if occupied. Use a managed persistent service while awaiting feedback. +9. Open the actual page in a browser. Exercise view switching, URL reload, + full-width links, and expanded findings. Check desktop and narrow layouts, + including iframe overflow and the visibility of operative details. Capture + a screenshot when available; if capture fails, report the limit separately + from successful DOM and interaction checks. Do not submit a real quiz. +10. Share the URL and explain the single decision to review. An authored preview + proves presentation of supplied content, not extractor adherence or improved + comprehension. Live extraction and human evaluation require their own + approved workflows; do not launch either as part of this skill. + +## Finish + +Record the accepted presentation and rationale in the tracking issue or commit +message. Stop the preview service and remove only its owned temporary directory +when the user has finished reviewing; do not stop unrelated evaluation servers. +Implement the accepted change in canonical production files and add meaningful +synthetic regression coverage where needed. Delete the comparison wrapper and +losing variants rather than shipping a permanent prototype route. No new test +is needed merely to assert this skill's wording. diff --git a/CODING_GUIDELINES.md b/CODING_GUIDELINES.md index 1135e07..eb1223c 100644 --- a/CODING_GUIDELINES.md +++ b/CODING_GUIDELINES.md @@ -22,6 +22,18 @@ Engram-only Reflection and Tricorder rules are not part of this repository. - Use the least expensive check that covers the demonstrated risk. Keep fast, deterministic checks in CI. Keep model-backed checks manual. +## Report presentation changes + +Before implementing changes to report hierarchy, navigation, or the main +comparison, show a representative synthetic preview and agree on the reader's +question and information hierarchy. Small copy corrections do not require a +prototype. Use the local `preview-behavior-diff-summary` skill for the workflow. + +Reuse the shipped renderer and hold trial evidence constant across presentation +variants. Label authored previews separately from model-generated results. +Keep prototypes temporary; retain only the accepted decision and appropriate +synthetic regression coverage, not a second rendering system. + ## Agent-skill Markdown ### Structure and discovery @@ -136,6 +148,20 @@ output. Run the checks relevant to the changed files. Do not claim checks that did not run. +Name the claim being verified and the evidence that supports it: + +- Schema and contract checks establish structure and invariants. +- Authored synthetic previews establish how supplied content is presented. +- Raw-trial audits establish fidelity for the inspected generated reports. +- Approved live extraction establishes observed prompt behavior on those inputs. +- Human evaluations establish reader performance subject to question validity. + +Do not substitute one evidence level for another. A renderer check does not +prove that a model follows a revised prompt; a stated plan does not prove +execution. Report remaining gaps explicitly. Live checks still require consent +and must remain outside CI. Keep evaluation-specific validity procedures in the +human-evaluation skill rather than duplicating them here. + ```bash docker run --rm -v "$PWD:/mnt" -w /mnt \ mvdan/shfmt:v3.14.0 -d -i 2 -ci . diff --git a/RETRO_NOTES.md b/RETRO_NOTES.md index ba60423..2a160b7 100644 --- a/RETRO_NOTES.md +++ b/RETRO_NOTES.md @@ -107,3 +107,17 @@ Process lessons, both of which cost real runs here: as "the diagram appeared but the plain-English half barely moved"; at 3+3 the drawing is what is unanimous. Use `--fast` to show shape in a demo, never to characterise an effect. + +## 2026-10-05 — summary previews and development tooling + +- Internal artifact URIs are tool-specific, not ordinary filesystem paths. + Before passing a PR body to an external CLI, resolve its file path or feed + its content through stdin. A successful internal read does not mean `gh` + can open the same URI. +- Long-lived Python kernels retain imported modules after their source changes. + Use a fresh process for verification, or explicitly reload a changed fixture + module during interactive preview work. Re-importing alone can use old data. +- Repeated screenshot timeouts are not proof that the page failed to render. + Check DOM and interactions separately, then try another capture backend. + Attaching to the terminal browser succeeded when headless capture timed out. + Report capture limits; never label DOM inspection as visual verification.