You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Current execution and evidence — PM review 2026-10-01
Primary track under #778; milestone 13 retains the Oct 14 evidence checkpoint. Latest measured baseline is the September 30 ledger in this discussion: 1 Q2 / 49 development cases, 9 PRs with rows, zero false-complete answers on main 6ced6f70. This is source-build development evidence, not a new released-build run or observed adoption.
#580/#864/#865/#872/#875/#876 are closed delivered prerequisites. #655/#867 remain open for acceptance reconciliation: establish which residual clauses are unmet before rebuilding already-present comparison behavior. Checked prerequisite boxes below record issue completion only; they do not establish this epic's ten-case outcome.
Exit remains ten qualifying Q2 cases with separate development/holdout results on a released build. Any material false row or omission disqualifies its case. Preserve named partial coverage and static-only operation. No automatic semantic declarations, qualification claim or outreach authorization follows from this plan.
Selected execution — 2026-09-24
Days 1–5 rebaseline and design: PR #869. Fresh PyPI 1.1.0/contract 40 and main d9a6d0e/contract 41 produce the same limitations on the selected corpus: 27 PRs × two builds = 54 head scans, plus eight fresh exact-ref verifies on four representative PRs. Attest, Capstone and ScopeIQ have missing_manifest/zero top changes on both builds; Visulate refuses the mode-160000 Gitlink. The ledger includes pinned refs, scope/config identity, release wheel digest and dependency versions. This is zero qualifying cases, not a success-rate estimate, and does not close this epic. Selected milestone 13; #580/#655 design first, then #867 and bounded readers. See PR for the reviewed comparison/absence semantics.
The original engine baseline below is historical; PR #869 records the fresh released/main rebaseline. The end-to-end acceptance remains unchanged.
Outcome
Deliver useful, reproducible application-agent capability comparisons for real open-source PRs that have never committed an Agents Shipgate manifest. The report should tell a reviewer which agent's source-observed callable surface changed, show the evidence and coverage limits, and support a concrete review decision before project adoption.
This tracks the product direction requested on 2026-09-23: unconfigured-project comparison plus tool-binding coverage. Existing #655, #656, and #580 remain the owners of their implementation scope; this issue joins them to the newly reproduced reader and per-agent comparison work.
Evidence baseline
Engine: source commit 9df307ab11f8d17b9e342a8b451b29f96e62bb31, local version 1.0.0, contract 40.
The investigation screened 310 distinct public PRs and locally ran Agents Shipgate on 27 priority candidates. No candidate yet met the full ten-case selection standard. 310 is a screening count, not 310 local runs or 310 demonstrated product failures. These targeted observations are development evidence, not a representative success-rate estimate or the released-build value gate in #830.
Real application-agent change
Observed limitation
Work owner
attest#3 adds memory read/write tools to an existing agent
Exact-ref verify reports base_status=missing_manifest, empty top_changes; head reads only the three old local tools
The source changes above establish relevant scenarios. They are not claims that Agents Shipgate already delivered complete review value, nor claims of runtime vulnerabilities in those projects. Each reader issue records exact SHAs, commands, observations, and bounded acceptance.
Exact-ref comparison reproduction
For attest#3, after the scoped init --local-review in #864:
Observed base_status=missing_manifest, capability_review.top_changes=[]. Providing Git refs alone does not establish an application-tool comparison when the base has no manifest. #655 owns automatic comparison; #580 owns accurate added/relocated scope and absence semantics. The existing shipgate diff host-grant route does not cover this application source change.
Run the end-to-end PR validation below and publish the reproducible evidence bundle.
Track additional obstacles explicitly during the rerun. Visulate's unrelated gitlink currently refuses verifier materialization (see #865); #430 and host fix #688 are relevant prior work. Do not count paired extraction as successful whole-PR verification while this remains unresolved. #584's SDK list-concatenation extension is adjacent work, not the cause of ScopeIQ's root-selection result.
End-to-end acceptance
One documented local comparison flow accepts a repository/scope and exact base/head refs without committing a manifest or manually authoring a base report. Report both requested refs and the actual compared trees/merge base, engine identity, input scope, and reader coverage.
Rows carry agent/tool identity, before/after, change kind/direction, source locations, and evidence basis. A reviewer can trace each conclusion to the specific PR. Stable output on identical inputs is verified.
Known structural changes remain visible beside precise unresolved edges. Definitions without wiring cannot masquerade as agent capability; ambiguous deployment roots cannot erase already-observed per-agent wiring.
Added/removed roots, moved files, changed bindings, unchanged tools, and unsupported dynamic paths have paired positive/negative regression cases. Incomplete reads cannot become “no changes” or a fabricated empty base.
Preserve source evidence versus inference, declared authority versus observed wiring, and advisory comparison versus release authorization. Never auto-fill purpose, effects, authority, or reviewed agent_bindings merely to make a result pass.
Validate 10 real application-agent PRs locally at pinned refs. For each, retain command/version, raw output, source-to-row reconciliation, coverage assessment, and the concrete reviewer decision the output enables. Verify the changed callable surface and direction; a material omission or false claim disqualifies the case.
A standalone insufficient_evidence, a head-only scan, an empty diff, or coding-agent host grants do not satisfy the ten-case result. Useful structural comparison does not require a release pass, but must be accurate and complete for its stated changed scope.
Publish attempted and qualifying counts honestly. The existing public PRs are regression candidates; do not assume they qualify after a parser repair, or generalize the result into market validation.
Related: #328 (reader limitations versus user configuration), #437 (bound-surface count semantics), #830 (separate released-build value gate). This tracker creates no outreach authorization and does not change existing release-control contracts.
Current execution and evidence — PM review 2026-10-01
Primary track under #778; milestone 13 retains the Oct 14 evidence checkpoint. Latest measured baseline is the September 30 ledger in this discussion: 1 Q2 / 49 development cases, 9 PRs with rows, zero false-complete answers on main
6ced6f70. This is source-build development evidence, not a new released-build run or observed adoption.#580/#864/#865/#872/#875/#876 are closed delivered prerequisites. #655/#867 remain open for acceptance reconciliation: establish which residual clauses are unmet before rebuilding already-present comparison behavior. Checked prerequisite boxes below record issue completion only; they do not establish this epic's ten-case outcome.
Exit remains ten qualifying Q2 cases with separate development/holdout results on a released build. Any material false row or omission disqualifies its case. Preserve named partial coverage and static-only operation. No automatic semantic declarations, qualification claim or outreach authorization follows from this plan.
Selected execution — 2026-09-24
Days 1–5 rebaseline and design: PR #869. Fresh PyPI 1.1.0/contract 40 and main d9a6d0e/contract 41 produce the same limitations on the selected corpus: 27 PRs × two builds = 54 head scans, plus eight fresh exact-ref verifies on four representative PRs. Attest, Capstone and ScopeIQ have missing_manifest/zero top changes on both builds; Visulate refuses the mode-160000 Gitlink. The ledger includes pinned refs, scope/config identity, release wheel digest and dependency versions. This is zero qualifying cases, not a success-rate estimate, and does not close this epic. Selected milestone 13; #580/#655 design first, then #867 and bounded readers. See PR for the reviewed comparison/absence semantics.
The original engine baseline below is historical; PR #869 records the fresh released/main rebaseline. The end-to-end acceptance remains unchanged.
Outcome
Deliver useful, reproducible application-agent capability comparisons for real open-source PRs that have never committed an Agents Shipgate manifest. The report should tell a reviewer which agent's source-observed callable surface changed, show the evidence and coverage limits, and support a concrete review decision before project adoption.
This tracks the product direction requested on 2026-09-23: unconfigured-project comparison plus tool-binding coverage. Existing #655, #656, and #580 remain the owners of their implementation scope; this issue joins them to the newly reproduced reader and per-agent comparison work.
Evidence baseline
Engine: source commit
9df307ab11f8d17b9e342a8b451b29f96e62bb31, local version 1.0.0, contract 40.The investigation screened 310 distinct public PRs and locally ran Agents Shipgate on 27 priority candidates. No candidate yet met the full ten-case selection standard. 310 is a screening count, not 310 local runs or 310 demonstrated product failures. These targeted observations are development evidence, not a representative success-rate estimate or the released-build value gate in #830.
base_status=missing_manifest, emptytop_changes; head reads only the three old local toolsload_memoryis unresolvedThe source changes above establish relevant scenarios. They are not claims that Agents Shipgate already delivered complete review value, nor claims of runtime vulnerabilities in those projects. Each reader issue records exact SHAs, commands, observations, and bounded acceptance.
Exact-ref comparison reproduction
For attest#3, after the scoped
init --local-reviewin #864:Observed
base_status=missing_manifest,capability_review.top_changes=[]. Providing Git refs alone does not establish an application-tool comparison when the base has no manifest. #655 owns automatic comparison; #580 owns accurate added/relocated scope and absence semantics. The existingshipgate diffhost-grant route does not cover this application source change.Delivery sequence and dependencies
shipgate diffv1 across Route A: manifest-free base synthesis; the capability delta is neverunavailableon first adoption; verdict only as a footer #655 — manifest-free application tool-source base comparison, with reproducible evidence identity.load_memory.insufficient_evidenceleaves the headline #656's first-value separation: useful observed changes and explicit coverage limits must be available without filling semantic declarations. A whole lock-state redesign should not become a prerequisite for the bounded comparison increment.Track additional obstacles explicitly during the rerun. Visulate's unrelated gitlink currently refuses verifier materialization (see #865); #430 and host fix #688 are relevant prior work. Do not count paired extraction as successful whole-PR verification while this remains unresolved. #584's SDK list-concatenation extension is adjacent work, not the cause of ScopeIQ's root-selection result.
End-to-end acceptance
agent_bindingsmerely to make a result pass.insufficient_evidence, a head-only scan, an empty diff, or coding-agent host grants do not satisfy the ten-case result. Useful structural comparison does not require a release pass, but must be accurate and complete for its stated changed scope.Related: #328 (reader limitations versus user configuration), #437 (bound-surface count semantics), #830 (separate released-build value gate). This tracker creates no outreach authorization and does not change existing release-control contracts.