Commit the application review benchmark and freeze the holdout rule (#908) - #926
Conversation
…908) The 49-PR development corpus behind #868's Q1/Q2 counts lived in a session scratchpad, and the OS cleanup deleted its pins on 2026-09-30. This commits it as benchmark/application-q2/: - the pinned corpus; - a runner (full-object clones; `diff --application --json` with a derived scope; nothing fetched during a diff) and a summarizer whose counts are mechanical; - the hand-scoring protocol, with #868's levels verbatim; - the 2026-09-30 ledger. Given the committed pins and the same build (6ced6f7), the committed runner reproduced every status, row and comparison id: 1 compared, 42 partial, 6 not_established; 9 PRs with rows; Q2 1/49. Re-scoring found that 8 members are LiveKit, ZeroRuntime, rustic-ai or hand-written agents rather than SDK/ADK changes, so Q0 is 41/49. They are kept, so the counts stay the published ones. The holdout window (PRs created from 2026-10-02) holds no pull request yet, so its pins cannot exist before the reader issues merge. By the owner's decision (2026-10-01), the selection rule is frozen here instead: pool queries, window, criteria, order and n=30. Membership is then generated, never chosen. The release runbook's step 1 now records `Q2: n/49 development, m/>=30 holdout` in each release's record. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Reviewed commit fba503421b30800116dce84c8ac7f54c46c1b1d2.
The committed pins, explicit hand scoring, retained non-Q0 development cases, and separation of development evidence from a future holdout make the measurement easier to audit. I found two runner defects that can misrepresent a failed measurement, described inline; I recommend fixing them before this runner becomes the release-runbook path.
Validation: the benchmark, documentation-link, privacy and public-surface suites passed (607 tests), and Ruff passed on the new Python files. I exercised the runner's timeout, missing-commit and nonzero-engine paths with a synthetic one-member corpus and controlled subprocess results. A timeout retained the prior compared JSON and returned exit 0; the summarizer counted that old row as current. An engine exit 2 also left the runner's exit at 0. These cases are not covered by the new tests.
The head's GitHub CI and both Agents Shipgate workflows are successful. Local advisory verification of the pinned base/head returned control_state=complete, decision=passed. I did not independently rerun all 49 external repositories or rescore their source; the published Q2 count remains the committed ledger's claim. The holdout has no pins yet, as explicitly recorded in this PR.
Posted by Codex at the repository owner's request.
Invalidate corpus outputs before clone preparation, record unavailable and timed-out members with current diagnostics, and return a failing runner status when any member is unavailable, times out, or exits nonzero. Add eight regression cases covering reruns and mixed successful/failed members, and document the runner's failure behavior.
Review fixes for #926. Runner: one full-object clone per repository with pins kept under refs/q2/<slug>, lazy fetching disabled, completeness checked over every object, explicit refspecs when fetching, timeouts and no credential prompts. Each member's outcome (ok, refused, timeout, unavailable) and reason is recorded in runs.json; earlier outputs and runs.json are removed first, and the exit code is non-zero unless every member answered. Scores: each names the answer_id it judged, the digest of the answer without the engine's own version, platform and build, so the same answer read on another machine or release keeps its score and a moved answer is listed as stale. Re-read against source: twelve members are not SDK/ADK wiring changes (Q0 37/49) and MIS_TALENT#6 is Q1 (Q1 2/49); Q2 stays 1/49. Holdout: the definitions the rule uses are frozen with it, and the digest of that text and of the recorded pool (pool.json, 8,126 repositories from 160 size-sharded code-search requests) is held by the tests. Relevance is judged from source at the pins, before the engine runs; framework repositories are excluded by their own package files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…enchmark # Conflicts: # CHANGELOG.md
Why
#868 closes at 10 Q2 cases, and the reader issues (#874, #909–#913) are about to move that count. The 49-PR development corpus lived in a session scratchpad, and the OS's
/tmpcleanup deleted its pins on 2026-09-30, so they had to be re-pinned from PR numbers. A count that cannot be reproduced is not a measurement. A count taken only on the PRs the readers were fixed against also overstates how the next stranger's PR will fare, which is why #908 asks for a holdout.What this commits:
benchmark/application-q2/development.jsonframeworksthe engine read for each (not what the repository imports: two members import a repository-localagentspackage) and the search that admitted them. 46 merged, 3 open.run.pyrefs/q2/<slug>/{base,head}so garbage collection cannot drop them. Completeness is checked over every object of both pins, with lazy fetching disabled.--fetchis the only network step, with explicit refspecs, timeouts and no credential prompts. Thendiff --application --base <merge_base> --head <head> --jsonruns with no--scope, as for a new user. Each member's outcome (ok,refused,timeout,unavailable) and reason go toruns.json. Earlier outputs andruns.jsonare removed before anything else, so a failed rerun never reports an earlier build's answer. It exits non-zero unless every member answered.summarize.pyanswer_idit judged: the digest of the answer without the engine's own version, Python, platform and build. A byte-identical answer from another machine or release keeps its score; a moved answer is listed as stale and not counted.results/2026-09-30-6ced6f70.{md,scores.json}pool.json,pool.pyholdout.jsonpool.json. No pins yet; see below.README.mddocs/release-runbook.mdrecordsQ2: n/49 development, m/≥30 holdoutin step 1 of § Cutting the release, and in the advisory list. Every release so far went through the advisory list, which joins the main path at step 4, after step 1. A test fails anydocs/changelog/<version>.mdafter 1.2.0 that lacks the line.Acceptance (#908)
6ced6f70, from fresh clones made with--fetch, all 49 answeredok. Statuses, rows andcomparison_ids are identical to the 2026-09-30 outputs: 1compared, 42partial, 6not_established; 9 PRs with rows; 60 changed, 8 added, 230 not established; Q2 = 1 (Memory Bank: Vertex AI working memory, and fix the import recursion it exposed jpka/attest#3).What the review changed
Agentclass (2) and a hand-written ReAct loop (1). Four change no application agent's wiring: a model-string update, library docstrings, and two agent-framework internals.Finance_Agent's construction and every tool it binds are unchanged, so the spread list the reader cannot resolve hides no row.1a0acb7on the original runner.aa11aeareplaces that runner with the per-repository design above, after a second review found three more defects: blobless clones were never fully hydrated, a bare console-script--enginefailed, and git calls had no timeout. The eight regression cases from1a0acb7are ported to it, with two more for an interrupted run and a git timeout.src/agents/function_schema.py,src/google/adk/runners.py) rather than by judgement. And it freezes the definitions it uses together with the steps.Choices for you to check
pool.pybefore it ran, and no window PR was looked at.contributory/starchatter-telegram→starfall-orb/starchatter-telegram). It is recorded under its canonical name, withrenamed_from.Validation
tests/test_application_q2_benchmark.py(20 tests):summarize.pyover the reproduced run with the committed scores: 49 scored, 0 stale, Q0 37, Q1 2, Q2 1.-m "not perf", plus the separately run CI files) passed on the review-fix tree; the benchmark, privacy and docs-link tests pass on the final one.ruff check .is clean.Refs #868, #830. The holdout and release boxes stay open, so this does not close #908.
🤖 Generated with Claude Code