Skip to content

Q2 measurement: commit the acceptance corpus, freeze a holdout, and report Q1/Q2 on every advisory release #908

Description

@pengfei-threemoonslab

PM sequencing note — 2026-10-01

P1 measurement prerequisite for #868. Commit the development pins, runner and scoring protocol first. The specified holdout contains >=30 PRs created after 2026-10-01, with pins frozen before #874/#909–#913 merge. As of this review, that future sampling window cannot supply a completed holdout; no freeze is claimed here.

Record the accountable owner, sampling window and actual freeze commit before reader merges. If the window cannot provide the required cases in time, record the shortfall and an explicit methodology/scheduling decision; do not silently use exposed development cases, backdate pins, lower the bar or call a development rerun holdout evidence. Oct 14 remains a checkpoint, not a guarantee that the sample or ten-case outcome exists.

This is independent of #830's released-build outreach gate. The existing counts and acceptance below are unchanged.


Evidence base (2026-09-30). Main 6ced6f70 (#864, #865, #875, #876 and #872 merged; #874 open) was run as diff --application --json with a derived scope (no --scope), which is the default a new user gets. The corpus is the #868 set of 49 real PRs that modify an existing OpenAI Agents SDK or Google ADK agent's tools= / sub_agents= / handoffs= in non-test code (46 merged, 3 open, all from the last 90 days). Each run is pinned from merge-base to head. Every row was scored by hand against the source using #868's Q1 and Q2 definitions.

main a430e81a (2026-09-25) main 6ced6f70 (2026-09-30)
PRs producing any row 3 / 42 9 / 49
Q2 (a useful, correct and complete answer on an existing agent) 0 1 (jpka/attest#3)
False complete answers (compared while a binding changed unseen) 1 (kkmiecik-coder/CRM#5) 0

40 PRs produce no row. Every one of them is honestly partial or not_established, and names what it could not read. The remaining work is therefore coverage and usefulness, not honesty.


Outcome

Every advisory release reports, on a corpus we did not tune against, how many real agent-changing PRs get a correct, complete and useful answer. "Is it useful yet?" then becomes a number we can check, not an impression.

Why now

  • Application-agent PR review without prior setup: exact-ref comparison, tool bindings, and 10 verified cases #868 closes at 10 Q2 cases. Today we have 1 (see below). Several reader issues are about to move that number, and we need to see which ones actually do.
  • The 49-PR corpus lived in a session scratchpad under /tmp. On 2026-09-30 the OS's daily /tmp cleanup deleted its pin list and the clones' Git metadata, and it had to be re-pinned from PR numbers. A measurement we cannot reproduce is not a measurement.
  • We have fixed readers against these same 49 PRs. A count on them alone overstates how the next PR a stranger opens will fare. We need a holdout set that no reader change has seen.

Scope

  1. Commit the development corpus in a benchmark directory:
    • one row per PR: repository, number, state, merge-base SHA, head SHA and the selection criteria it met;
    • the runner (full-object clones; diff --application --json with the derived scope; no network during the diff);
    • a per-release ledger with, for each PR: status, row counts by change kind, sides naming a reach or a named unresolved hop, the Q0/Q1/Q2 score and a one-line rationale.
  2. Freeze a holdout corpus of at least 30 PRs:
    • chosen with the same criteria from a later window (PRs created after 2026-10-01);
    • pins committed before any of the reader issues below land;
    • scored only at release time and never used to debug a reader.
  3. Scoring protocol in the benchmark README:
  4. Release step: docs/release-runbook.md § Cutting the release gains one step. Each advisory release's record (docs/changelog/<version>.md) states Q2: n/49 development, m/≥30 holdout and the change since the previous release.

Acceptance criteria

Non-goals

  • Automating the hand scoring.
  • Growing the corpus beyond what scoring can sustain per release.

Metric this moves

It defines the metric: Q1 and Q2 counts on development and holdout. Related: #868, #830.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1Next after P0; blocks other work or ships a misleading resultarea:benchmarkLabeled corpus, accuracy measurementenhancementNew feature or requestquestionFurther information is requestedworkstream:applicationManifest-free application-agent comparisons, binding coverage and Q1/Q2 proof.

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions