Skip to content

feat(report): surface everything eval-harness 1.6 records - #20

Merged
lopadova merged 1 commit into
mainfrom
claude/laravel-iam-rebel-toolkit-drlhbe
Aug 24, 2026
Merged

lopadova merged 1 commit into
mainfrom
claude/laravel-iam-rebel-toolkit-drlhbe

Conversation

@lopadova

Copy link
Copy Markdown
Contributor

padosoft/eval-harness 1.6 added repeated sampling with real statistics, content hashes and per-row aggregates, agent trajectories, cost and budgets.

All of it already arrives in the report artifact this screen fetches in full, so the work here was reading it — no new endpoints, no server-side change, nothing to release in lockstep.

Three new tabs on report detail

Rows — per-row pass rate with its Wilson confidence interval, the row_hash the regression gate joins on, and an expandable worst execution carrying what the pipeline produced, the judge's own reason, and the tool calls behind it.

Defaults to failing rows, worst first. A run-level macro-F1 says whether to worry; this says which row to open, which is the only number that leads to a fix. Showing the first execution instead of the worst would hide the failure the reader opened the row to see.

A run that stopped on a pending approval says so in bold — "I have submitted that refund" while an approval is queued reads as success and is not.

Sampling & precision — repetitions, the smallest difference the run could actually detect, whether the difference you are gating on is even detectable at this sample size (and how many repetitions it would take), and the unstable rows: the ones that disagree with themselves and fail builds nobody broke.

Cost — total split into what the provider billed and what was derived from configured token rates, a per-model table, and — in words rather than in a footnote — "this total is a floor, not a figure" when calls ran on a model with no declared rate. A cost total that omits what it cannot price is an unknown bill, and the small number is the one that gets quoted in a meeting.

Two decisions worth reviewing

A budget halt is in the page header, not only on the cost tab. Every figure in a halted report describes a partial run, and the rows that never executed are unknowns rather than passes — so a reader looking at the metrics tab has to know before they read the metrics.

Absent data is never rendered as zero. A report from before repeated sampling did not measure a resolution of 0 — it measured nothing. Each panel degrades to a sentence plus a pointer at the flag that would produce the data:

This run recorded no sampling statistics.
Run with --repetitions=N to measure the smallest difference the run can actually detect.

0.0% would be a number somebody could act on that nobody produced. A test asserts it never appears.

Verification

  • 26 tests (was 8) — 18 new across the three panels and the readers.
  • tsc --noEmit, vitest, and the production build are green.
  • 40 new i18n keys in both en and it.
  • No screen was removed and no existing tab changed; older reports render exactly as before.

The PHP gate was not run locally: this session's proxy cannot authenticate to github.com for the Pest dist downloads. The diff touches no PHP — only resources/js/**, README.md and docs/ — so CI covers it.

🤖 Generated with Claude Code

https://claude.ai/code/session_01EZ4u82zdoa8naYm75kppMC


Generated by Claude Code

eval-harness 1.6 added repeated sampling with real statistics, content hashes
and per-row aggregates, agent trajectories, cost and budgets. All of it arrives
in the report artifact this screen already fetches in full, so the work was
reading it rather than exposing it: no new endpoints, no server-side change.

Three tabs on report detail:

- Rows: per-row pass rate with its Wilson interval, the row_hash the regression
  gate joins on, and an expandable worst execution carrying what the pipeline
  produced, the judge's own reason and the tool calls behind it. Defaults to
  failing rows worst-first -- a run-level macro-F1 says whether to worry, this
  says which row to open. A run that stopped on a pending approval says so,
  because "I have submitted that" while an approval is queued reads as success
  and is not.
- Sampling & precision: repetitions, the smallest difference the run could
  actually detect, whether the difference being gated on is detectable at this
  sample size and how many repetitions it would take, and the unstable rows --
  the ones that disagree with themselves and fail builds nobody broke.
- Cost: total split into billed and derived, per-model breakdown, and "this
  total is a floor, not a figure" in words when calls ran on a model with no
  declared rate. A cost total that omits what it cannot price is an unknown
  bill, and the small number is the one that gets quoted.

A run halted on its budget raises an alert in the page header, visible from
every tab: every figure in that report describes a partial run, and the rows
that never executed are unknowns rather than passes.

Absent data is never rendered as zero. A report from before repeated sampling
did not measure a resolution of 0 -- it measured nothing -- so each panel
degrades to a sentence and a pointer at the flag that would produce the data.

Tests 8 -> 26. tsc --noEmit, vitest and the production build are green. The PHP
gate was not run locally (this session's proxy cannot authenticate to github.com
for the Pest dist downloads); the diff touches no PHP.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EZ4u82zdoa8naYm75kppMC

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 13459b1108

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +30 to +31
const resolution = precision.resolution;
const resolvable = precision.target_resolvable === true;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Read precision values from the nested run block

When a valid precision payload places run-level statistics in precision.run—a shape explicitly represented by ReportPrecision.run—this panel ignores them and reads only the optional top-level fields. Such reports display — for the detectable difference and claim the target is not resolvable even when run.resolution and run.target_resolvable contain the recorded result; use the nested values as fallbacks and avoid asserting non-resolvability when neither form is present.

Useful? React with 👍 / 👎.

.map((entry) => entry?.score)
.filter((score): score is number => typeof score === 'number');

return scores.length === 0 ? Number.POSITIVE_INFINITY : scores.reduce((a, b) => a + b, 0) / scores.length;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Rank unscored executions as failures when selecting the worst

If a row has an errored or otherwise unscored execution alongside a scored execution, assigning the unscored run positive infinity makes it rank as the best candidate. Expanding that failing row can consequently show a successful scored response while hiding the execution that produced no score, contradicting the panel's worst-execution contract; unscored/error runs need to sort before scored runs or be classified explicitly.

Useful? React with 👍 / 👎.

@lopadova
lopadova merged commit ec79649 into main Aug 24, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants