Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
85 changes: 85 additions & 0 deletions .agents/skills/preview-behavior-diff-summary/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,85 @@
---
name: preview-behavior-diff-summary
description: Builds temporary, synthetic Behavior Diff summary comparisons using the existing renderer. Use when previewing report hierarchy or navigation changes, comparing outcome-first and rule-first cards, or requesting a before/after UI prototype without live model calls.
---

# Preview Behavior Diff summaries

Compare presentations of the same evidence before changing production behavior.
This is a maintainer workflow, not plugin payload or a human-evaluation run.

## Quick start

From the repository root:

```bash
python3 tests/report-demo.py --serve --port 0
```

Open the printed loopback URL. The existing gallery uses synthetic fixtures,
shipped ingestion, and the shipped renderer; it makes no model calls. Keep the
server running for review. Ctrl+C stops it and removes its temporary directory.

## Prepare the comparison

1. State the reader's question and the presentation decision being tested.
Example: “Can the reader see the changed retry interval without expanding
Other findings?” Distinguish report-version labels from instruction
Before/After labels inside each report.
2. Read `tests/report-demo.py`, `tests/report_fixtures.py`, and the affected
renderer/summary contract. Choose representative synthetic evidence. Include
mixed branches or an unchanged result when they matter to the decision.
Never copy private evaluation reports or source excerpts into the checkout.
3. Build paired reports in a fresh temporary directory outside the checkout.
Use the existing fixture builder, ingest command, and renderer invocation
from `build_reports` in `tests/report_fixtures.py`; inspect current signatures
rather than copying an old invocation. For a throwaway driver, put `tests`
on Python's import path and restrict the fixture catalog to the selected
scenario before calling `build_reports` on an empty directory.
4. Copy that generated scenario into two sibling variant directories. Preserve
task, instruction versions, trial names, traces, chain choices/memberships,
and primary outcome. For a narrative comparison, change only the selected
summary row and its supported narrative in each `extraction.json`, then run
the same shipped `decisions.py --ingest` and `render.py` commands in each copy.
Do not alter counts or hide minority branches to improve the illustration.
5. If comparing renderer implementations instead, feed identical saved inputs
to each renderer revision. Do not change renderer and authored narrative at
once unless explicitly comparing both; name every changed variable.
Reject invalid narratives rather than bypassing ingestion validation.

## Present and verify

6. Create a temporary HTML wrapper around the two rendered reports, not a new
renderer. Offer side-by-side and full-width views with shareable
`?variant=compare`, `?variant=before`, and `?variant=after` states. Keep a
visible switcher and clearly label the two presentation versions.
7. State on the page: “Authored synthetic illustration; no model was called.”
Do not call an authored baseline a historical report. Show unchanged evidence
and remaining uncertainty, including plans versus executed actions.
8. Serve only the temporary root on loopback, for example:

```bash
python3 -m http.server 8767 --bind 127.0.0.1 --directory "$PREVIEW_ROOT"
```

Set `PREVIEW_ROOT` to the fresh directory from step 3; select another local
port if occupied. Use a managed persistent service while awaiting feedback.
9. Open the actual page in a browser. Exercise view switching, URL reload,
full-width links, and expanded findings. Check desktop and narrow layouts,
including iframe overflow and the visibility of operative details. Capture
a screenshot when available; if capture fails, report the limit separately
from successful DOM and interaction checks. Do not submit a real quiz.
10. Share the URL and explain the single decision to review. An authored preview
proves presentation of supplied content, not extractor adherence or improved
comprehension. Live extraction and human evaluation require their own
approved workflows; do not launch either as part of this skill.

## Finish

Record the accepted presentation and rationale in the tracking issue or commit
message. Stop the preview service and remove only its owned temporary directory
when the user has finished reviewing; do not stop unrelated evaluation servers.
Implement the accepted change in canonical production files and add meaningful
synthetic regression coverage where needed. Delete the comparison wrapper and
losing variants rather than shipping a permanent prototype route. No new test
is needed merely to assert this skill's wording.
26 changes: 26 additions & 0 deletions CODING_GUIDELINES.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,18 @@ Engram-only Reflection and Tricorder rules are not part of this repository.
- Use the least expensive check that covers the demonstrated risk. Keep fast,
deterministic checks in CI. Keep model-backed checks manual.

## Report presentation changes

Before implementing changes to report hierarchy, navigation, or the main
comparison, show a representative synthetic preview and agree on the reader's
question and information hierarchy. Small copy corrections do not require a
prototype. Use the local `preview-behavior-diff-summary` skill for the workflow.

Reuse the shipped renderer and hold trial evidence constant across presentation
variants. Label authored previews separately from model-generated results.
Keep prototypes temporary; retain only the accepted decision and appropriate
synthetic regression coverage, not a second rendering system.

## Agent-skill Markdown

### Structure and discovery
Expand Down Expand Up @@ -136,6 +148,20 @@ output.
Run the checks relevant to the changed files. Do not claim checks that did not
run.

Name the claim being verified and the evidence that supports it:

- Schema and contract checks establish structure and invariants.
- Authored synthetic previews establish how supplied content is presented.
- Raw-trial audits establish fidelity for the inspected generated reports.
- Approved live extraction establishes observed prompt behavior on those inputs.
- Human evaluations establish reader performance subject to question validity.

Do not substitute one evidence level for another. A renderer check does not
prove that a model follows a revised prompt; a stated plan does not prove
execution. Report remaining gaps explicitly. Live checks still require consent
and must remain outside CI. Keep evaluation-specific validity procedures in the
human-evaluation skill rather than duplicating them here.

```bash
docker run --rm -v "$PWD:/mnt" -w /mnt \
mvdan/shfmt:v3.14.0 -d -i 2 -ci .
Expand Down
14 changes: 14 additions & 0 deletions RETRO_NOTES.md
Original file line number Diff line number Diff line change
Expand Up @@ -107,3 +107,17 @@ Process lessons, both of which cost real runs here:
as "the diagram appeared but the plain-English half barely moved"; at 3+3
the drawing is what is unanimous. Use `--fast` to show shape in a demo,
never to characterise an effect.

## 2026-10-05 — summary previews and development tooling

- Internal artifact URIs are tool-specific, not ordinary filesystem paths.
Before passing a PR body to an external CLI, resolve its file path or feed
its content through stdin. A successful internal read does not mean `gh`
can open the same URI.
- Long-lived Python kernels retain imported modules after their source changes.
Use a fresh process for verification, or explicitly reload a changed fixture
module during interactive preview work. Re-importing alone can use old data.
- Repeated screenshot timeouts are not proof that the page failed to render.
Check DOM and interactions separately, then try another capture backend.
Attaching to the terminal browser succeeded when headless capture timed out.
Report capture limits; never label DOM inspection as visual verification.
Loading