Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 15 additions & 7 deletions .agents/skills/run-behavior-diff-human-evaluation/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -44,9 +44,12 @@ plugin payload. Never trigger it from hooks, CI, or a scheduled job.
For every case, prepare a neutral synthetic decision-point fixture and
`scenario.json`, then four patch-grounded options in `question.json`.
Prefer a fresh scenario worker that cannot see the options/answer key.
The correct statement must concern behavior this scenario can expose;
do not bundle unrelated changes. Keep unchanged or weak results.
6. Freeze all five fixtures and questions **before** any live trial:
The keyed statement must distinguish the anticipated After behavior from
Before at that decision point, not merely describe behavior true of both.
Complete the workflow's prefreeze contrast audit in the keyed rationale;
source support alone does not establish an observable contrast. Keep
unchanged or weak results.
6. Freeze all five fixtures and audited questions **before** any live trial:

```bash
python3 "$EVAL" freeze "$SESSION"
Expand Down Expand Up @@ -76,10 +79,15 @@ python3 "$EVAL" results "$SESSION"
```

If no submission exists, say so; never invent a score. Report correct out of
five, confidence, and insufficient-evidence count. Compare misses with the
saved scenarios, patches, and trial answers—not just the answer key. Separate
readability, factual fidelity, scenario coverage, and quiz ambiguity. An
explanation-only change is not necessarily a changed action or outcome.
five, confidence, and insufficient-evidence count. Assess question validity for
**all five cases** against saved Before/After answers, source sides, scenarios,
and blinded summaries using [the workflow](references/workflow.md#question-validity-assessment).
Distinguish an observed delta from an unchanged or both-sides match. Retain the
original score and explicitly qualify non-discriminating questions, whether
answered correctly or incorrectly; do not reinterpret them as comprehension
successes or failures. Separate readability, factual fidelity, scenario coverage,
and quiz ambiguity. An explanation-only change is not necessarily a changed
action or outcome.

Chance averages 1.25/5. Five cases may share a skill; this score is not an
estimate of overall product accuracy. Record methodology and aggregate
Expand Down
114 changes: 101 additions & 13 deletions .agents/skills/run-behavior-diff-human-evaluation/references/workflow.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,10 +67,10 @@ pinned parent revision. Do not inspect unrelated private company material.

```json
{
"stem": "Which statement describes the instruction change?",
"stem": "Which statement describes the Before/After contrast in this scenario?",
"scope": "The decision this scenario can expose; identify the relevant patch section.",
"options": [
{"statement": "First candidate statement.", "correct": true, "rationale": "Patch support and why the scenario covers it."},
{"statement": "First candidate contrast.", "correct": true, "rationale": "Before: source support. After: changed rule. Scenario: triggering facts. Contrast: anticipated observable difference and limits."},
{"statement": "Second candidate statement.", "correct": false, "rationale": "Why the patch contradicts or does not introduce it."},
{"statement": "Third candidate statement.", "correct": false, "rationale": "Why this is false for the selected change."},
{"statement": "Fourth candidate statement.", "correct": false, "rationale": "Why this is false for the selected change."}
Expand All @@ -79,15 +79,47 @@ pinned parent revision. Do not inspect unrelated private company material.
```

This is a schema example, not ready-to-use question content. Author four
concrete, similarly specific statements with exactly one supported by the
patch. Avoid bundled claims outside the scenario, trivia, obvious nonsense,
concrete, similarly specific statements with exactly one patch-grounded
contrast that the scenario can expose. The keyed statement must distinguish
anticipated After behavior from Before; a statement true of both sides is
not a discriminating answer, even if it accurately describes After.
Avoid bundled claims outside the scenario, trivia, obvious nonsense,
uniquely repeated keywords, or an answer longer than all the distractors.
A clarification-only change is valid; do not claim it changes the outcome.
Check each rationale against the actual patch, independently of the report.
5. **Freeze all cases.** `freeze` validates inputs, shuffles options, writes
the private answer key and public questions, and records input hashes.
A clarification-only change can target a stated reason, condition, or timing;
do not claim it changes the action or outcome. Check each rationale against
both complete source sides and the patch, independently of the report.
5. **Complete the prefreeze contrast audit.** Before authorizing `freeze`,
record these four checks in the keyed option's existing `rationale`:

- **Before:** cite the relevant source section and say whether the keyed
behavior is already required, permitted, or illustrated there. Check
surrounding rules and relevant references, not only removed patch lines.
- **After:** cite the changed rule and identify the precise added or changed
condition, timing, next step, or explanation. Separate explicit wording
from an inference about how a trial might respond.
- **Scenario:** identify the local facts and decision point that activate
that rule. State what the read-only task can and cannot demonstrate.
- **Contrast:** state what observable Before/After difference would support
the keyed statement and what would instead make it an unchanged or
both-sides match. A difference must concern the whole statement, not merely
a keyword appearing in After.

Have the maintainer review all four checks before freezing all five cases.
Revise an unsupported or both-sides statement, or its scenario, only during
preparation, without seeing live results. Do not replace a sampled commit
because its rule is redundant or its expected contrast is weak. If no
defensible contrast can be authored, report the design limitation before
live execution rather than inventing a difference.
6. **Freeze all cases.** `freeze` validates input structure, shuffles options,
writes the private answer key and public questions, and records input hashes.
The existing schema is unchanged: `scope` bounds the claim and the keyed
`rationale` holds the audit. The deterministic helper enforces structure and
immutability, **not semantic contrast**; a successful freeze does not prove
that the question discriminates. The audit is a required maintainer gate.
After freezing, do not edit fixtures, questions, or the key. Start a new
evaluation if the design must change; retain the abandoned session and why.
Once results exist, a weak or absent contrast is evidence, never grounds to
reject, resample, retry, or change the keyed answer.

Scenario design is model-assisted maintainer work, not a claim that the installed
Behavior Diff plugin independently drafted the scenario. When delegating, keep
Expand Down Expand Up @@ -147,11 +179,67 @@ source commit links, and the full reports become available afterward. Keep the
service running until the human finishes; restart `serve` to resume later.

`results SESSION` reads the saved submission without a model call. Report the
score out of five, confidence, insufficient-evidence count, and comments.
Investigate misses against the original reports, scenarios, patches, and traces.
Do not infer comprehension from keyword matching or score alone. Inspect factual
consistency even in correctly answered cases. Preserve the first score; do not
rescore after changing a question or revealing the answers.
original score out of five, confidence, insufficient-evidence count, and comments.
Then complete the question-validity assessment below for all cases, not just
misses. Do not infer comprehension from keyword matching or score alone. Preserve
the first submission and score; do not rescore after changing a question,
excluding weak cases, or revealing the answers. These instructions apply when
analyzing older sessions too; never retrofit their questions or frozen rationales.

## Question validity assessment

After submission, use only saved evidence. For every case, compare the frozen
keyed statement and rationale with both complete source sides, the patch,
scenario, original Before/After trial answers, and the blinded summaries that
the human saw. Record the assessment in private analysis notes alongside the
session, without editing hashed inputs, reports, the answer key, or submission.
No new model run is needed or authorized by this assessment.

Separate two judgments:

1. **Observed contrast:** does the whole keyed statement distinguish After from
Before in the saved answers?
- **Observed delta:** the answers support the stated Before/After distinction.
Name the changed condition, timing, next step, or explanation and any trial
variability; do not turn a partial pattern into a universal claim.
- **Unchanged / both-sides match:** the keyed behavior is present on both
sides, or the scoped behavior is unchanged. Explicitly label the question
**non-discriminating in this run**, even if the key is source-supported or
the human selected it correctly.
- **Not observed / contradicted:** usable answers do not show the forecast
contrast or instead show a different direction. Preserve that result.
- **Inconclusive:** missing, blocked, or inconsistent evidence prevents a
supported distinction. Name the missing evidence; do not infer a delta.
2. **Summary exposure:** do the blinded summaries faithfully expose the supported
distinction? Separate omitted operative rules or timing from factual errors,
scenario undercoverage, and question ambiguity. A delta visible only in full
traces does not prove the human could infer it from the quiz excerpt.

Cite private evidence locations and describe what each side actually says.
Assess correctly answered cases as well as misses. Retain the raw keyed score
out of five and qualify non-discriminating, unobserved, or inconclusive cases in
the analysis. Do not count an unchanged match as evidence of successful delta
comprehension, or its missed key as evidence of poor comprehension. Report
validity counts separately from the score; do not publish a revised denominator
or retroactively choose a different correct option.

### Synthetic table-and-wait example

Suppose Before already documents a blast-radius table, and After adds a rule to
disclose that table before applying a change. In saved trials, both responses
show the table, but only After explicitly says it will wait before proceeding.
The keyed statement “After shows a blast-radius table” is a both-sides match,
not an observed delta. Preserve its original score and mark it non-discriminating.
The supported observed contrast is the stated wait/next-step difference, not
the presence of the table. A future question can target that timing distinction
only if its source-and-scenario audit supports it, and must be frozen before
its own trials; do not rewrite this session's key using the example.

Keep source intent separate from observed behavior: a disclosure-before-action
rule does **not** by itself require literal user approval. If After says it
will wait for approval, report that as observed response wording, not as an
explicit source requirement unless the source actually contains that requirement.
A stated wait is not evidence that any command executed or approval was obtained.

A random guess averages 1.25/5. Five possibly correlated cases cannot establish
product-wide accuracy. Separate summary readability/fidelity from scenario
Expand Down
13 changes: 11 additions & 2 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -245,8 +245,17 @@ author intent or goal completion. Markdown retains the caveat as plain text.
comparison. Model output supplies text and selectors, never markup.
Summary and trial-summary instructions require concrete, parallel descriptions
of the same subject, with the decisive contrast and any unchanged decision
explicitly stated. Changed explanations, citations, or presentation must not
be described as changed actions. When the selected lead is not the primary
explicitly stated. Summary selection prefers the observed changed operative rule
(conditions, timing, scope, or prerequisites) over its downstream outcome, including
mixed comparisons. Material intervals, deadlines, units, and gates belong in the
main cards, not only Other findings. Each card detail must hold for every named trial
in its selected branch; counts or citations from other rows cannot supply membership.
An instruction's gate is distinct from an observed gate, which may already appear
in Before. These are extraction policies, not deterministic semantic guarantees;
the existing validator checks references, choice coverage, and evidence anchors.
No schema field or keyword-based semantic check is added; fallback ranking is unchanged.
Changed explanations, citations, or presentation must not be described as changed
actions. When the selected lead is not the primary
result, `content.primary_result_context` exposes the primary status and full
distribution beside the cards in both formats. Missing or incomplete evidence
cannot become an unchanged-result claim. This is derived presentation, not a
Expand Down
31 changes: 24 additions & 7 deletions plugin/skills/behavior-diff/scripts/decisions.py
Original file line number Diff line number Diff line change
Expand Up @@ -212,23 +212,40 @@
decision observations, or presumed author motivation. Do not claim the aim was met.
This edit-only model interpretation is independent of "summary" and chain indexes.
- In the SAME reply, optionally give a concise plain-language "summary" grounded
in a meaningful selected chain row, or null when unsupported. Prefer a changed
primary result, then a unanimous changed action, then a mixed primary result,
then a changed mixed action or changed edit-linked comparison. Do not elevate
answer wording over a supported process difference. For no observed difference,
limit the headline to this scenario, never claim the edit has no effect.
in a meaningful selected chain row, or null when unsupported. Prefer the observed
changed operative rule: the condition, timing, scope, or prerequisite that changes
what the agent does or plans. Select that row rather than only its downstream
result, even when the primary result also changes or either row has mixed choices.
The application keeps the primary result and its full distribution beside the
cards. An edit link alone does not establish an observed rule change.
If no operative-rule contrast is supported, prefer a changed primary result,
then a unanimous changed action, then a mixed primary result, then a changed mixed
action or changed edit-linked comparison. Do not elevate answer wording over a
supported process difference. For no observed difference, limit the headline to
this scenario, never claim the edit has no effect.
- Include EVERY choice on both sides of that row, copying its canonical "choice"
exactly. Do not provide summary counts: the application uses validated row counts.
Mixed primary results must not be described as unanimous even when the lead is
another action. Same-result/different-process is a valid finding.
- Ground each side's label and detail in the records of EVERY named trial assigned
to that selected branch. Do not borrow a condition, interval, deadline, or other
fact from another branch or row's memberships, even when their counts match.
If a detail is not shared by those trials, split the observed choices faithfully,
select the row that records the rule, or omit that detail; do not combine rows
into a counted execution path. Why/caution citations do not support card details.
- Write concrete actor + verb + object sentences in plain language. Explain an
internal workflow name only when the reader needs it to understand the finding.
Use parallel before/after sentences about the same subject; say what changed
and what stayed the same. Distinguish changed choices, actions, or stated plans
from changed explanations, citations, or presentation alone. Do not infer
actions from answers. Put the decisive contrast in the headline and main side
details, not only why/caution; avoid abstract correlation or the author's stance.
Do not duplicate count claims in prose: the application derives counts.
details, not only why/caution or Other findings. Include material intervals,
deadlines and their units, triggers, prerequisites, and exceptions when supported;
a vague "waits longer" or "uses a stricter gate" is not enough. Describe the observed
gate separately from what the instruction requires: a newly written gate may
already appear in Before, and a rule in the diff is not evidence that After used it.
Avoid abstract correlation or the author's stance. Do not duplicate count claims
in prose: the application derives counts.
- Use "plans" only for plans stated in final answers, "answers" for other final
answer choices, and "actions" only for a numbered action/command anchor. Plans
are not executed actions. Final answers do not prove tool execution. Self-reported
Expand Down
Loading
Loading