You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
P1 measurement prerequisite for #868. Commit the development pins, runner and scoring protocol first. The specified holdout contains >=30 PRs created after 2026-10-01, with pins frozen before #874/#909–#913 merge. As of this review, that future sampling window cannot supply a completed holdout; no freeze is claimed here.
Record the accountable owner, sampling window and actual freeze commit before reader merges. If the window cannot provide the required cases in time, record the shortfall and an explicit methodology/scheduling decision; do not silently use exposed development cases, backdate pins, lower the bar or call a development rerun holdout evidence. Oct 14 remains a checkpoint, not a guarantee that the sample or ten-case outcome exists.
This is independent of #830's released-build outreach gate. The existing counts and acceptance below are unchanged.
Evidence base (2026-09-30). Main 6ced6f70 (#864, #865, #875, #876 and #872 merged; #874 open) was run as diff --application --json with a derived scope (no --scope), which is the default a new user gets. The corpus is the #868 set of 49 real PRs that modify an existing OpenAI Agents SDK or Google ADK agent's tools= / sub_agents= / handoffs= in non-test code (46 merged, 3 open, all from the last 90 days). Each run is pinned from merge-base to head. Every row was scored by hand against the source using #868's Q1 and Q2 definitions.
main a430e81a (2026-09-25)
main 6ced6f70 (2026-09-30)
PRs producing any row
3 / 42
9 / 49
Q2 (a useful, correct and complete answer on an existing agent)
40 PRs produce no row. Every one of them is honestly partial or not_established, and names what it could not read. The remaining work is therefore coverage and usefulness, not honesty.
Outcome
Every advisory release reports, on a corpus we did not tune against, how many real agent-changing PRs get a correct, complete and useful answer. "Is it useful yet?" then becomes a number we can check, not an impression.
The 49-PR corpus lived in a session scratchpad under /tmp. On 2026-09-30 the OS's daily /tmp cleanup deleted its pin list and the clones' Git metadata, and it had to be re-pinned from PR numbers. A measurement we cannot reproduce is not a measurement.
We have fixed readers against these same 49 PRs. A count on them alone overstates how the next PR a stranger opens will fare. We need a holdout set that no reader change has seen.
Scope
Commit the development corpus in a benchmark directory:
one row per PR: repository, number, state, merge-base SHA, head SHA and the selection criteria it met;
the runner (full-object clones; diff --application --json with the derived scope; no network during the diff);
a per-release ledger with, for each PR: status, row counts by change kind, sides naming a reach or a named unresolved hop, the Q0/Q1/Q2 score and a one-line rationale.
Freeze a holdout corpus of at least 30 PRs:
chosen with the same criteria from a later window (PRs created after 2026-10-01);
pins committed before any of the reader issues below land;
scored only at release time and never used to debug a reader.
Release step:docs/release-runbook.md § Cutting the release gains one step. Each advisory release's record (docs/changelog/<version>.md) states Q2: n/49 development, m/≥30 holdout and the change since the previous release.
Acceptance criteria
The committed runner reproduces the 2026-09-30 numbers from the committed pins:
PM sequencing note — 2026-10-01
P1 measurement prerequisite for #868. Commit the development pins, runner and scoring protocol first. The specified holdout contains >=30 PRs created after 2026-10-01, with pins frozen before #874/#909–#913 merge. As of this review, that future sampling window cannot supply a completed holdout; no freeze is claimed here.
Record the accountable owner, sampling window and actual freeze commit before reader merges. If the window cannot provide the required cases in time, record the shortfall and an explicit methodology/scheduling decision; do not silently use exposed development cases, backdate pins, lower the bar or call a development rerun holdout evidence. Oct 14 remains a checkpoint, not a guarantee that the sample or ten-case outcome exists.
This is independent of #830's released-build outreach gate. The existing counts and acceptance below are unchanged.
Evidence base (2026-09-30). Main
6ced6f70(#864, #865, #875, #876 and #872 merged; #874 open) was run asdiff --application --jsonwith a derived scope (no--scope), which is the default a new user gets. The corpus is the #868 set of 49 real PRs that modify an existing OpenAI Agents SDK or Google ADK agent'stools=/sub_agents=/handoffs=in non-test code (46 merged, 3 open, all from the last 90 days). Each run is pinned from merge-base to head. Every row was scored by hand against the source using #868's Q1 and Q2 definitions.a430e81a(2026-09-25)6ced6f70(2026-09-30)comparedwhile a binding changed unseen)40 PRs produce no row. Every one of them is honestly
partialornot_established, and names what it could not read. The remaining work is therefore coverage and usefulness, not honesty.Outcome
Every advisory release reports, on a corpus we did not tune against, how many real agent-changing PRs get a correct, complete and useful answer. "Is it useful yet?" then becomes a number we can check, not an impression.
Why now
/tmp. On 2026-09-30 the OS's daily/tmpcleanup deleted its pin list and the clones' Git metadata, and it had to be re-pinned from PR numbers. A measurement we cannot reproduce is not a measurement.Scope
diff --application --jsonwith the derived scope; no network during the diff);docs/release-runbook.md§ Cutting the release gains one step. Each advisory release's record (docs/changelog/<version>.md) statesQ2: n/49 development, m/≥30 holdoutand the change since the previous release.Acceptance criteria
compared, 42partial, 6not_established;Non-goals
Metric this moves
It defines the metric: Q1 and Q2 counts on development and holdout. Related: #868, #830.