feat(report): surface everything eval-harness 1.6 records - #20
Conversation
eval-harness 1.6 added repeated sampling with real statistics, content hashes and per-row aggregates, agent trajectories, cost and budgets. All of it arrives in the report artifact this screen already fetches in full, so the work was reading it rather than exposing it: no new endpoints, no server-side change. Three tabs on report detail: - Rows: per-row pass rate with its Wilson interval, the row_hash the regression gate joins on, and an expandable worst execution carrying what the pipeline produced, the judge's own reason and the tool calls behind it. Defaults to failing rows worst-first -- a run-level macro-F1 says whether to worry, this says which row to open. A run that stopped on a pending approval says so, because "I have submitted that" while an approval is queued reads as success and is not. - Sampling & precision: repetitions, the smallest difference the run could actually detect, whether the difference being gated on is detectable at this sample size and how many repetitions it would take, and the unstable rows -- the ones that disagree with themselves and fail builds nobody broke. - Cost: total split into billed and derived, per-model breakdown, and "this total is a floor, not a figure" in words when calls ran on a model with no declared rate. A cost total that omits what it cannot price is an unknown bill, and the small number is the one that gets quoted. A run halted on its budget raises an alert in the page header, visible from every tab: every figure in that report describes a partial run, and the rows that never executed are unknowns rather than passes. Absent data is never rendered as zero. A report from before repeated sampling did not measure a resolution of 0 -- it measured nothing -- so each panel degrades to a sentence and a pointer at the flag that would produce the data. Tests 8 -> 26. tsc --noEmit, vitest and the production build are green. The PHP gate was not run locally (this session's proxy cannot authenticate to github.com for the Pest dist downloads); the diff touches no PHP. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EZ4u82zdoa8naYm75kppMC
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 13459b1108
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| const resolution = precision.resolution; | ||
| const resolvable = precision.target_resolvable === true; |
There was a problem hiding this comment.
Read precision values from the nested run block
When a valid precision payload places run-level statistics in precision.run—a shape explicitly represented by ReportPrecision.run—this panel ignores them and reads only the optional top-level fields. Such reports display — for the detectable difference and claim the target is not resolvable even when run.resolution and run.target_resolvable contain the recorded result; use the nested values as fallbacks and avoid asserting non-resolvability when neither form is present.
Useful? React with 👍 / 👎.
| .map((entry) => entry?.score) | ||
| .filter((score): score is number => typeof score === 'number'); | ||
|
|
||
| return scores.length === 0 ? Number.POSITIVE_INFINITY : scores.reduce((a, b) => a + b, 0) / scores.length; |
There was a problem hiding this comment.
Rank unscored executions as failures when selecting the worst
If a row has an errored or otherwise unscored execution alongside a scored execution, assigning the unscored run positive infinity makes it rank as the best candidate. Expanding that failing row can consequently show a successful scored response while hiding the execution that produced no score, contradicting the panel's worst-execution contract; unscored/error runs need to sort before scored runs or be classified explicitly.
Useful? React with 👍 / 👎.
padosoft/eval-harness1.6 added repeated sampling with real statistics, content hashes and per-row aggregates, agent trajectories, cost and budgets.All of it already arrives in the report artifact this screen fetches in full, so the work here was reading it — no new endpoints, no server-side change, nothing to release in lockstep.
Three new tabs on report detail
Rows — per-row pass rate with its Wilson confidence interval, the
row_hashthe regression gate joins on, and an expandable worst execution carrying what the pipeline produced, the judge's own reason, and the tool calls behind it.Defaults to failing rows, worst first. A run-level macro-F1 says whether to worry; this says which row to open, which is the only number that leads to a fix. Showing the first execution instead of the worst would hide the failure the reader opened the row to see.
A run that stopped on a pending approval says so in bold — "I have submitted that refund" while an approval is queued reads as success and is not.
Sampling & precision — repetitions, the smallest difference the run could actually detect, whether the difference you are gating on is even detectable at this sample size (and how many repetitions it would take), and the unstable rows: the ones that disagree with themselves and fail builds nobody broke.
Cost — total split into what the provider billed and what was derived from configured token rates, a per-model table, and — in words rather than in a footnote — "this total is a floor, not a figure" when calls ran on a model with no declared rate. A cost total that omits what it cannot price is an unknown bill, and the small number is the one that gets quoted in a meeting.
Two decisions worth reviewing
A budget halt is in the page header, not only on the cost tab. Every figure in a halted report describes a partial run, and the rows that never executed are unknowns rather than passes — so a reader looking at the metrics tab has to know before they read the metrics.
Absent data is never rendered as zero. A report from before repeated sampling did not measure a resolution of
0— it measured nothing. Each panel degrades to a sentence plus a pointer at the flag that would produce the data:0.0%would be a number somebody could act on that nobody produced. A test asserts it never appears.Verification
tsc --noEmit,vitest, and the production build are green.enandit.🤖 Generated with Claude Code
https://claude.ai/code/session_01EZ4u82zdoa8naYm75kppMC
Generated by Claude Code