Skip to content

Application-agent PR review without prior setup: exact-ref comparison, tool bindings, and 10 verified cases #868

Description

@pengfei-threemoonslab

Current execution and evidence — PM review 2026-10-01

Primary track under #778; milestone 13 retains the Oct 14 evidence checkpoint. Latest measured baseline is the September 30 ledger in this discussion: 1 Q2 / 49 development cases, 9 PRs with rows, zero false-complete answers on main 6ced6f70. This is source-build development evidence, not a new released-build run or observed adoption.

  1. Q2 measurement: commit the acceptance corpus, freeze a holdout, and report Q1/Q2 on every advisory release #908: preserve pins/runner/scoring and freeze the untouched holdout before the gated reader merges. Record its future-window scheduling constraint explicitly.
  2. Application diff: resolve tools lists built by spread, concatenation, conditionals and comprehensions #909 list expressions + Application diff: resolve tool lists passed into agent-building functions and constructors #874 parameter/instance flow; keep their ownership separate. OpenAI Agents SDK: resolve bounded literal tool-list concatenation in static binding evidence #584's concatenation acceptance stays with Application diff: resolve tools lists built by spread, concatenation, conditionals and comprehensions #909.
  3. Application diff: read tools bound as objects — MCP toolsets, agent-as-tool and hosted tools #910 object tools and built-ins, including Google ADK: preserve built-in load_memory identity and binding in capability diffs #866's load_memory regression.
  4. Tool reach: name effects beyond HTTP — databases, subprocesses, files, cloud SDKs and messaging #913 observed effects + Application diff output: one finding per changed agent, the first unresolved hop, grouped uncertainty #914 concise reviewer findings.
  5. Application diff: read agents derived with clone() as agents with their own bindings #911 clones and Application diff: identify agents whose name is not a literal (base-class template methods, computed names) #912 nonliteral identities, after the higher-yield shapes; their bounded identity work can proceed independently where documented.
  6. GitHub Action: an application mode that posts the comparison as a PR comment #915 Action delivery can ship with current truthful text and adopt Application diff output: one finding per changed agent, the first unresolved hop, grouped uncertainty #914 later. Measure which agent frameworks real agent-changing PRs use, and decide the next reader (ADR) #916 measures framework mix without implementing a reader.

#580/#864/#865/#872/#875/#876 are closed delivered prerequisites. #655/#867 remain open for acceptance reconciliation: establish which residual clauses are unmet before rebuilding already-present comparison behavior. Checked prerequisite boxes below record issue completion only; they do not establish this epic's ten-case outcome.

Exit remains ten qualifying Q2 cases with separate development/holdout results on a released build. Any material false row or omission disqualifies its case. Preserve named partial coverage and static-only operation. No automatic semantic declarations, qualification claim or outreach authorization follows from this plan.


Selected execution — 2026-09-24

Days 1–5 rebaseline and design: PR #869. Fresh PyPI 1.1.0/contract 40 and main d9a6d0e/contract 41 produce the same limitations on the selected corpus: 27 PRs × two builds = 54 head scans, plus eight fresh exact-ref verifies on four representative PRs. Attest, Capstone and ScopeIQ have missing_manifest/zero top changes on both builds; Visulate refuses the mode-160000 Gitlink. The ledger includes pinned refs, scope/config identity, release wheel digest and dependency versions. This is zero qualifying cases, not a success-rate estimate, and does not close this epic. Selected milestone 13; #580/#655 design first, then #867 and bounded readers. See PR for the reviewed comparison/absence semantics.


The original engine baseline below is historical; PR #869 records the fresh released/main rebaseline. The end-to-end acceptance remains unchanged.

Outcome

Deliver useful, reproducible application-agent capability comparisons for real open-source PRs that have never committed an Agents Shipgate manifest. The report should tell a reviewer which agent's source-observed callable surface changed, show the evidence and coverage limits, and support a concrete review decision before project adoption.

This tracks the product direction requested on 2026-09-23: unconfigured-project comparison plus tool-binding coverage. Existing #655, #656, and #580 remain the owners of their implementation scope; this issue joins them to the newly reproduced reader and per-agent comparison work.

Evidence baseline

Engine: source commit 9df307ab11f8d17b9e342a8b451b29f96e62bb31, local version 1.0.0, contract 40.

The investigation screened 310 distinct public PRs and locally ran Agents Shipgate on 27 priority candidates. No candidate yet met the full ten-case selection standard. 310 is a screening count, not 310 local runs or 310 demonstrated product failures. These targeted observations are development evidence, not a representative success-rate estimate or the released-build value gate in #830.

Real application-agent change Observed limitation Work owner
attest#3 adds memory read/write tools to an existing agent Exact-ref verify reports base_status=missing_manifest, empty top_changes; head reads only the three old local tools #655, #864
capstone_project#3 adds SQL and memory tools A manually supplied real base report enables nonempty SQL tool additions, but load_memory is unresolved #655, #866
visulate-for-oracle#526 adds repository-memory tool factories Head inventory is empty; factory variables are unresolved #864, #865
scopeiq#2 introduces tools on two SDK agent objects Five direct edges are already extracted, but an unresolved deployment root leaves root-scoped inventory empty #867

The source changes above establish relevant scenarios. They are not claims that Agents Shipgate already delivered complete review value, nor claims of runtime vulnerabilities in those projects. Each reader issue records exact SHAs, commands, observations, and bounded acceptance.

Exact-ref comparison reproduction

For attest#3, after the scoped init --local-review in #864:

agents-shipgate verify --workspace subject   --config agents/attest_orchestrator/.agents-shipgate-local-review.yaml   --base 10e1f949d48b5349a180677c41c4ba30deab943d   --head cab295001ec01e7fca388acee67842c6f53b5150   --ci-mode advisory --format json

Observed base_status=missing_manifest, capability_review.top_changes=[]. Providing Git refs alone does not establish an application-tool comparison when the base has no manifest. #655 owns automatic comparison; #580 owns accurate added/relocated scope and absence semantics. The existing shipgate diff host-grant route does not cover this application source change.

Delivery sequence and dependencies

Track additional obstacles explicitly during the rerun. Visulate's unrelated gitlink currently refuses verifier materialization (see #865); #430 and host fix #688 are relevant prior work. Do not count paired extraction as successful whole-PR verification while this remains unresolved. #584's SDK list-concatenation extension is adjacent work, not the cause of ScopeIQ's root-selection result.

End-to-end acceptance

  • One documented local comparison flow accepts a repository/scope and exact base/head refs without committing a manifest or manually authoring a base report. Report both requested refs and the actual compared trees/merge base, engine identity, input scope, and reader coverage.
  • Rows carry agent/tool identity, before/after, change kind/direction, source locations, and evidence basis. A reviewer can trace each conclusion to the specific PR. Stable output on identical inputs is verified.
  • Known structural changes remain visible beside precise unresolved edges. Definitions without wiring cannot masquerade as agent capability; ambiguous deployment roots cannot erase already-observed per-agent wiring.
  • Added/removed roots, moved files, changed bindings, unchanged tools, and unsupported dynamic paths have paired positive/negative regression cases. Incomplete reads cannot become “no changes” or a fabricated empty base.
  • Preserve source evidence versus inference, declared authority versus observed wiring, and advisory comparison versus release authorization. Never auto-fill purpose, effects, authority, or reviewed agent_bindings merely to make a result pass.
  • Validate 10 real application-agent PRs locally at pinned refs. For each, retain command/version, raw output, source-to-row reconciliation, coverage assessment, and the concrete reviewer decision the output enables. Verify the changed callable surface and direction; a material omission or false claim disqualifies the case.
  • A standalone insufficient_evidence, a head-only scan, an empty diff, or coding-agent host grants do not satisfy the ten-case result. Useful structural comparison does not require a release pass, but must be accurate and complete for its stated changed scope.
  • Publish attempted and qualifying counts honestly. The existing public PRs are regression candidates; do not assume they qualify after a parser repair, or generalize the result into market validation.

Related: #328 (reader limitations versus user configuration), #437 (bound-surface count semantics), #830 (separate released-build value gate). This tracker creates no outreach authorization and does not change existing release-control contracts.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

P1Next after P0; blocks other work or ships a misleading resultepicTracking issue coordinating a group of related issuesworkstream:applicationManifest-free application-agent comparisons, binding coverage and Q1/Q2 proof.

Type

No type

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions