Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
87 changes: 87 additions & 0 deletions .agents/skills/run-behavior-diff-human-evaluation/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,87 @@
---
name: run-behavior-diff-human-evaluation
description: Run a private, five-question human evaluation of the current local Behavior Diff summaries using randomly sampled skill commits from DataRecce/recce-team. Use when a maintainer asks for a human evaluation, a blind report quiz, or to evaluate current summary ability; also use to resume or analyze an existing evaluation.
---

# Human evaluation of Behavior Diff

Manually run five real comparisons, then ask a human to identify each change
from four shuffled statements. This is a repository-maintainer skill, not
plugin payload. Never trigger it from hooks, CI, or a scheduled job.

## Start or resume

1. For **analyze/resume**, locate the requested private session under
`~/.behavior-diff/human-evaluations/`; use its saved evidence. Never create
fresh trials to answer a request about an existing result.
2. For **new evaluation**, require this repository's checkout, Python 3.9+,
Git, authenticated `gh` access to `DataRecce/recce-team`, Bash, `jq`, and
the Claude Code CLI.
A Claude or Codex maintainer session can orchestrate it; the trial stack
is explicitly Claude Code with read-only tools, not the orchestrator's
default host. Read [the workflow](references/workflow.md) before starting.
3. Explain the spend and privacy boundary. Require fresh approval for this
evaluation: five reports, three Before and three After trials each,
plus extraction; private skill text is sent to the configured Claude
provider. Default selectors are `opus` for trials and `sonnet` for
extraction. Previous sessions' approvals do not authorize this one.
4. Initialize a fresh sample using the bundled helper:

```bash
EVAL=.agents/skills/run-behavior-diff-human-evaluation/scripts/evaluate.py
SESSION=$(python3 "$EVAL" init)
```

Always use the fixed upstream `DataRecce/recce-team`. The helper pins its
current `main`, generates a new random seed, and samples five eligible
commits without replacement. Never handpick commits or reuse a saved
sample. Independent random samples may overlap; disclose known familiarity.
Keep SHAs, patches, subjects, and correct options out of human-facing chat.

## Prepare and run

5. Follow [case preparation](references/workflow.md#case-preparation).
For every case, prepare a neutral synthetic decision-point fixture and
`scenario.json`, then four patch-grounded options in `question.json`.
Prefer a fresh scenario worker that cannot see the options/answer key.
The correct statement must concern behavior this scenario can expose;
do not bundle unrelated changes. Keep unchanged or weak results.
6. Freeze all five fixtures and questions **before** any live trial:

```bash
python3 "$EVAL" freeze "$SESSION"
python3 "$EVAL" run "$SESSION" --approve-live
python3 "$EVAL" build "$SESSION"
python3 "$EVAL" serve "$SESSION" --port 0
```

Pass `--approve-live` only after step 3's approval. Use a managed long-lived
service for `serve`; give the human the printed loopback URL. The helper
uses this checkout's code, including uncommitted changes. Never switch to
the installed plugin, silently retry a case, or change code mid-evaluation.
7. Check actual trial completions, intended revision reads, and generated
report artifacts. A blocked or unchanged result is evidence, not grounds
for replacement. Verify the quiz in a browser without submitting answers.
Do not use the real session for a scoring smoke test.
8. Deliver the URL and instructions: read the summaries, choose one option
per case, rate confidence, flag insufficient evidence, and submit once.
Do not show ground truth until submission. Original reports and commit
links unlock afterward. Artifacts stay outside the checkout and are
never attached to a public issue or committed.

## Analyze the human's answers

```bash
python3 "$EVAL" results "$SESSION"
```

If no submission exists, say so; never invent a score. Report correct out of
five, confidence, and insufficient-evidence count. Compare misses with the
saved scenarios, patches, and trial answers—not just the answer key. Separate
readability, factual fidelity, scenario coverage, and quiz ambiguity. An
explanation-only change is not necessarily a changed action or outcome.

Chance averages 1.25/5. Five cases may share a skill; this score is not an
estimate of overall product accuracy. Record methodology and aggregate
findings in the evaluation's tracking issue without private excerpts. Do not
rewrite summaries, edit answers, rerun models, or implement fixes unless asked.
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Behavior Diff — human evaluation</title>
<link rel="stylesheet" href="/quiz.css">
<script src="/quiz.js" defer></script>
</head>
<body>
<header>
<p class="eyebrow">Behavior Diff · Human evaluation</p>
<h1>Read the summary. Choose the supported statement.</h1>
<p>For each case, read the generated Before/After summary and select one of four statements.
Record your confidence. If the summary does not provide enough evidence to choose, flag that too;
still select your best answer. An optional note can explain ambiguity.</p>
<p class="notice">Your first complete submission is final. Answers, rationales, source commits and full reports
become available afterward. Draft answers are saved only in this browser for this session.</p>
</header>
<main>
<p id="status" role="status" aria-live="polite">Loading this evaluation…</p>
<form id="quiz" hidden>
<div id="cases"></div>
<div class="submit-row"><button id="submit" type="submit">Submit all five answers</button>
<span id="progress"></span></div>
</form>
<section id="results" hidden aria-labelledby="results-title">
<h2 id="results-title">Your session results</h2>
<div id="score"></div>
<div id="reviews"></div>
</section>
</main>
<footer>These questions measure comprehension of the displayed reports in this session, not overall model accuracy.</footer>
</body>
</html>
31 changes: 31 additions & 0 deletions .agents/skills/run-behavior-diff-human-evaluation/assets/quiz.css
Original file line number Diff line number Diff line change
@@ -0,0 +1,31 @@
:root { color-scheme: light; --ink: #1c2733; --muted: #526171; --border: #d9e1e7; --accent: #0b6e75; }
* { box-sizing: border-box; }
body { margin: 0 auto; max-width: 1140px; padding: 2rem 1.25rem; color: var(--ink); background: #f6f8f9; font: 16px/1.55 -apple-system, BlinkMacSystemFont, "Segoe UI", sans-serif; }
h1 { font-size: clamp(1.6rem, 5vw, 2.3rem); line-height: 1.25; margin-top: .25rem; }
h2, h3 { line-height: 1.35; }
.eyebrow { color: var(--accent); text-transform: uppercase; letter-spacing: .08em; font-size: .8rem; font-weight: 700; }
.notice { padding: 1rem; background: #e3f0f1; border-left: 3px solid var(--accent); }
.case { background: #fff; border: 1px solid var(--border); border-radius: .65rem; padding: 1.5rem; margin: 1.5rem 0; overflow-wrap: anywhere; }
.case > h2, .case > h3 { margin-top: 0; }
iframe { display: block; width: 100%; height: 600px; border: 1px solid var(--border); border-radius: .4rem; margin: 1rem 0 1.5rem; background: #f6f8f9; }
fieldset { border: 0; margin: 0 0 1rem; padding: 0; min-width: 0; }
legend { font-weight: 650; margin-bottom: .8rem; padding: 0; font-size: 1.1rem; }
.option { display: flex; align-items: flex-start; gap: .75rem; border: 1px solid var(--border); border-radius: .4rem; padding: .8rem; margin: .5rem 0; cursor: pointer; }
.option:has(input:checked) { border-color: var(--accent); background: #e3f0f1; }
input[type="radio"], input[type="checkbox"] { flex: none; margin: .35rem 0 0; width: 1.1rem; height: 1.1rem; accent-color: var(--accent); }
select, textarea, button { font: inherit; }
select, textarea { border: 1px solid #9baab8; border-radius: .3rem; padding: .5rem; color: var(--ink); background: #fff; max-width: 100%; }
textarea { display: block; width: 100%; resize: vertical; margin-top: .4rem; }
.sufficiency { display: flex; gap: .7rem; margin: 1rem 0; align-items: flex-start; }
.submit-row { display: flex; align-items: center; flex-wrap: wrap; gap: 1rem; padding: 1rem 0; }
button { padding: .75rem 1.2rem; border: 0; border-radius: .4rem; color: #fff; background: var(--accent); font-weight: 650; cursor: pointer; }
button:disabled { opacity: .6; cursor: default; }
a { color: var(--accent); text-underline-offset: .18em; }
a:focus-visible, button:focus-visible, input:focus-visible, select:focus-visible, textarea:focus-visible, summary:focus-visible { outline: 3px solid var(--accent); outline-offset: 3px; }
.links { display: flex; flex-wrap: wrap; gap: .7rem 1.2rem; }
summary { cursor: pointer; font-weight: 650; padding: .6rem 0; }
.score { font-size: 1.5rem; font-weight: 700; }
#status, #progress, footer { color: var(--muted); }
footer { margin-top: 2rem; border-top: 1px solid var(--border); padding-top: 1rem; }
[hidden] { display: none !important; }
@media (max-width: 600px) { body { padding: 1rem .6rem; } .case { padding: .8rem; } iframe { margin: .7rem 0 1rem; } .option { padding: .7rem; } }
196 changes: 196 additions & 0 deletions .agents/skills/run-behavior-diff-human-evaluation/assets/quiz.js
Original file line number Diff line number Diff line change
@@ -0,0 +1,196 @@
"use strict";

const byId = (id) => document.getElementById(id);
let sessionId;
let questions;
let storageKey;
let submitted = false;
const frames = new Map();

function element(tag, text, className) {
const node = document.createElement(tag);
if (text !== undefined) node.textContent = text;
if (className) node.className = className;
return node;
}

async function request(path, options) {
const response = await fetch(path, { credentials: "same-origin", ...options });
const data = await response.json();
if (!response.ok) throw new Error(data.error || "The local server rejected this request.");
return data;
}

function fitFrame(frame) {
try {
const doc = frame.contentDocument;
if (!doc || !doc.body) return;
// Measure content without changing the viewport inside ResizeObserver.
const style = frame.contentWindow.getComputedStyle(doc.body);
const margins = (parseFloat(style.marginTop) || 0) + (parseFloat(style.marginBottom) || 0);
const height = Math.ceil(doc.body.getBoundingClientRect().height + margins) + 24;
if (Math.abs((parseFloat(frame.style.height) || 0) - height) > 2) frame.style.height = `${height}px`;
} catch (_) {
// Full reports remain accessible through their separate post-submit link.
}
}

function reportFrame(path, title, full = false) {
const frame = element("iframe");
frame.title = title;
frame.src = path;
frame.setAttribute("sandbox", full ? "allow-scripts allow-same-origin" : "allow-same-origin");
frame.addEventListener("load", () => {
fitFrame(frame);
const doc = frame.contentDocument;
if (!doc) return;
doc.addEventListener("toggle", () => requestAnimationFrame(() => fitFrame(frame)), true);
doc.addEventListener("click", () => setTimeout(() => fitFrame(frame), 40));
if (window.ResizeObserver) {
const observer = new ResizeObserver(() => fitFrame(frame));
observer.observe(doc.body);
frames.set(frame, observer);
}
});
return frame;
}

function answerFor(question) {
const id = question.id;
const selected = document.querySelector(`input[name="choice-${id}"]:checked`);
const confidence = byId(`confidence-${id}`).value;
return { id, letter: selected ? selected.value : "", confidence,
insufficient: byId(`insufficient-${id}`).checked, note: byId(`note-${id}`).value };
}

function saveDraft() {
if (submitted) return;
const answers = questions.map(answerFor);
try { localStorage.setItem(storageKey, JSON.stringify({ session_id: sessionId, answers })); } catch (_) { /* Storage is optional. */ }
const complete = answers.filter((answer) => answer.letter && answer.confidence).length;
byId("progress").textContent = `${complete} of ${questions.length} complete`;
}

function restoreDraft() {
try {
const draft = JSON.parse(localStorage.getItem(storageKey));
if (!draft || draft.session_id !== sessionId || !Array.isArray(draft.answers)) return;
for (const answer of draft.answers) {
if (!questions.some((question) => question.id === answer.id)) continue;
const choice = document.querySelector(`input[name="choice-${answer.id}"][value="${["A", "B", "C", "D"].includes(answer.letter) ? answer.letter : ""}"]`);
if (choice) choice.checked = true;
if (["low", "medium", "high"].includes(answer.confidence)) byId(`confidence-${answer.id}`).value = answer.confidence;
byId(`insufficient-${answer.id}`).checked = answer.insufficient === true;
if (typeof answer.note === "string") byId(`note-${answer.id}`).value = answer.note.slice(0, 2000);
}
} catch (_) { /* Ignore malformed or unavailable browser drafts. */ }
}

function renderQuestions() {
for (const question of questions) {
const section = element("section", undefined, "case");
section.append(element("h2", `Case ${question.id}`));
section.append(reportFrame(`/summary-${question.id}.html`, `Blinded summary for case ${question.id}`));
const fieldset = element("fieldset");
fieldset.append(element("legend", question.stem));
for (const option of question.options) {
const label = element("label", undefined, "option");
const input = element("input");
input.type = "radio"; input.name = `choice-${question.id}`; input.value = option.letter; input.required = true;
label.append(input, element("span", `${option.letter}. ${option.statement}`));
fieldset.append(label);
}
section.append(fieldset);
const confidenceLabel = element("label", "Confidence ");
const confidence = element("select");
confidence.id = `confidence-${question.id}`; confidence.required = true;
for (const [value, label] of [["", "Choose confidence"], ["low", "Low"], ["medium", "Medium"], ["high", "High"]]) {
const option = element("option", label); option.value = value; confidence.append(option);
}
confidenceLabel.append(confidence);
const insufficientLabel = element("label", undefined, "sufficiency");
const insufficient = element("input");
insufficient.type = "checkbox"; insufficient.id = `insufficient-${question.id}`;
insufficientLabel.append(insufficient, element("span", "The summary provides insufficient evidence to choose confidently."));
const noteLabel = element("label", "Optional note");
const note = element("textarea");
note.id = `note-${question.id}`; note.maxLength = 2000; note.rows = 3;
noteLabel.append(note);
section.append(confidenceLabel, insufficientLabel, noteLabel);
byId("cases").append(section);
}
restoreDraft();
byId("quiz").addEventListener("input", saveDraft);
byId("quiz").addEventListener("change", saveDraft);
saveDraft();
byId("quiz").hidden = false;
}

function renderResults(data) {
submitted = true;
byId("quiz").hidden = true;
byId("results").hidden = false;
byId("status").textContent = "First complete submission saved. You can now inspect the answer key and full reports.";
byId("score").replaceChildren(element("p", `${data.score.correct} of ${data.score.total} answers correct.`, "score"), element("p", data.interpretation));
byId("reviews").replaceChildren();
for (const answer of data.answers) {
const review = element("article", undefined, "case");
review.append(element("h3", `Case ${answer.id}: ${answer.correct ? "correct" : "incorrect"}`));
review.append(element("p", `Your answer: ${answer.letter}. Correct answer: ${answer.correct_letter}. Confidence: ${answer.confidence}. Insufficient evidence flagged: ${answer.insufficient ? "yes" : "no"}.`));
review.append(element("p", `Question scope: ${answer.scope}`));
const reasons = element("ul");
for (const option of answer.options) {
const reason = element("li");
reason.append(element("strong", `${option.letter}. ${option.statement}${option.correct ? " (correct)" : ""}`), element("p", option.rationale));
reasons.append(reason);
}
review.append(reasons);
if (answer.note) review.append(element("p", `Your note: ${answer.note}`));
const links = element("p", undefined, "links");
for (const [url, label] of [[answer.source_url, "Source commit"], [answer.before_source_url, "Before commit"], [answer.review_url, "Open full report"]]) {
const link = element("a", label); link.href = url; link.target = "_blank"; link.rel = "noopener noreferrer"; links.append(link);
}
review.append(links);
const details = element("details");
details.append(element("summary", "Full original Behavior Diff report"));
details.addEventListener("toggle", () => {
if (details.open && !details.querySelector("iframe")) details.append(reportFrame(answer.review_url, `Full report for case ${answer.id}`, true));
else if (details.open) fitFrame(details.querySelector("iframe"));
});
review.append(details);
byId("reviews").append(review);
}
}

byId("quiz").addEventListener("submit", async (event) => {
event.preventDefault();
if (submitted || !byId("quiz").reportValidity()) return;
byId("submit").disabled = true;
byId("status").textContent = "Saving your first complete submission…";
try {
renderResults(await request("/submit", { method: "POST", headers: { "Content-Type": "application/json" },
body: JSON.stringify({ session_id: sessionId, answers: questions.map(answerFor) }) }));
} catch (error) {
byId("status").textContent = error.message;
byId("submit").disabled = false;
}
});

window.addEventListener("resize", () => { for (const frame of document.querySelectorAll("iframe")) fitFrame(frame); });

(async () => {
try {
const session = await request("/session-public.json");
sessionId = session.session_id;
storageKey = `behavior-diff-human-evaluation:${sessionId}`;
// Server persistence, not browser storage, is authoritative about submission.
const response = await fetch("/results", { credentials: "same-origin" });
if (response.ok) { renderResults(await response.json()); return; }
if (response.status !== 403) throw new Error("Unable to read this session's submission state.");
questions = await request("/questions.json");
renderQuestions();
byId("status").textContent = "Read all five summaries before submitting. Your draft is private to this session.";
} catch (error) {
byId("status").textContent = error.message;
}
})();
Loading
Loading