Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 27 additions & 2 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,30 @@ The current-control envelope supplies operational permissions and the next
action; verifier, handoff, PR and report surfaces project the same decision.
Execution success is not merge authority.

## Lead wedge (focus)
## Selected execution — 2026-09-24

The owner selected the **application-agent PR proof sprint** (#868) as primary,
with bounded host reliability maintenance. Accountable owner: Pengfei Hu
(`pengfei-threemoonslab`); current technical execution: Codex.
[Days 1–5 evidence and ordered backlog](docs/research/application-days1-5/README.md)
records released/main reproductions and the #580/#655 comparison design.
Weeks 2–3 implement paired inputs, then per-agent wiring; deeper readers follow
reproduced gaps. #610 remains a reproduced contract defect; #787 is the selected
recipe repair. #795/#812/#780/#369 retain their exact residual acceptance.

This selection supersedes the scheduling and recruitment instructions in the
September 14 historical plan below. #830 and external-outreach holds remain;
Oct 14 / Nov 13 / Dec 13 are evidence checkpoints, not release promises.
No ten-case value, external adoption or qualification claim is made.

Status as of 2026-09-27: #870 closed #787; #871 merged `diff --application`
(not yet released), an advisory application comparison with its own JSON
schema; #873 closed #580, #877 extended its reach and #879 closed #864. #868,
#655, #867, #865, #866, #610, #795, #812, #780, #369 and #830 remain open.

<a id="lead-wedge-focus"></a>

## Historical lead wedge (September 14 focus)

Two surfaces share one engine: **(A)** tool-surface readiness for agent builders,
and **(B)** repository-declared host configuration, MCP bindings, permissions,
Expand All @@ -40,7 +63,9 @@ blocking CI. Organization-wide adoption must be demonstrated. New surface
follows the [non-goals](#explicit-non-goals) and
[`CONTRIBUTING.md`](CONTRIBUTING.md#surface-discipline).

## Post-1.0 adoption (plan of record, 2026-09-14)
<a id="post-10-adoption-plan-of-record-2026-09-14"></a>

## Historical post-1.0 adoption (2026-09-14; superseded above)

**Adoption > completeness.** `v1.0.0` is published on PyPI and GitHub from
`bace7c1871834e0b3eb98e6f60c0627725c53a59`, and #777 moved the pins to it.
Expand Down
119 changes: 119 additions & 0 deletions docs/research/application-days1-5/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,119 @@
# Application PR review: Days 1–5 execution record

Date: 2026-09-24. Program: #778 / #868. This is engineering evidence, not
external user validation or proof of ten useful PR reviews.

## Decision and ownership

The owner requested execution of Days 1–5 of the September 23 issue review.
Select application-agent comparison as the primary proof sprint; keep host
correctness in a bounded maintenance lane. Accountable product/acceptance owner:
Pengfei Hu (`pengfei-threemoonslab`); technical execution: Codex in this task.
No additional engineer staffing is assumed. Maintainer acceptance remains a
human decision. #830 and its outreach hold remain unchanged.

| Issue | Days 1–5 disposition | Next acceptance / dependency |
| --- | --- | --- |
| #778 | Record this selected sequence and supersede earlier recruitment-first scheduling | Preserve Oct 14 / Nov 13 / Dec 13 evidence checkpoints; do not promise releases |
| #868 | Rebaseline completed below; epic remains open | Ten individually source-checked, complete, useful application PR comparisons |
| #580 | Comparison input design recorded in `comparison-design.md` | Implement independently bound scopes and truthful absence |
| #655 | Advisory synthesis design reconciled with #580 | Implement automatic paired extraction; never count synthesized input as reviewed base |
| #867 | Selected after comparison input support | Project existing direct wiring per agent without guessing deployment root |
| #864 / #865 / #866 | Selected subsequent reader work | Local imports, bounded factories, exact ADK built-in identity; keep unresolved paths visible |
| #610 | Reproduced on release and main; remains open | Planning-only semantics, compatible schemas/projections and positive/negative consumer tests |
| #787 | Selected bounded reliability implementation | Safe-path installation in non-GitHub recipes and regression checks |
| #795 | Keep open for acceptance reconciliation | Delivered field presentation does not prove all same-version CLI/JSON/PR and reviewer criteria |
| #812 | Keep open for acceptance reconciliation | Delivered coverage slices do not prove unaided reviewer interpretation; #821 is closed, do not rebuild it |
| #780 | Keep open for exactly the hosted residuals | No-change update to existing comment; Not compared; hosted setup/execution failure |
| #369 | Deferred unless a concrete operational consumer needs argv | Existing POSIX command recovery is tested; setup argv does not satisfy this issue |

#795/#812 need a criterion-by-criterion same-version evidence join before closure,
not another broad renderer rewrite. #780's ten published 1.1.0 hosted runs remain
valid; #853 merged, removing its source-pin caveat. #854–#859 retain their own
observed defects. Hosted runs do not demonstrate external adoption (#571).
This phase selects and reconciles those residuals; it does not claim they shipped.

Status as of 2026-09-27 (the table above is the 2026-09-24 snapshot): #787 was
closed by #870, #580 by #873 and #864 by #879. #871 merged `diff --application`
(not yet released), a new flag with its own advisory JSON schema; #873, #877
and #879 extended it. #868, #655, #867, #865, #866, #610, #795, #812, #780 and
#369 remain open.

## Engines and method

- Release: PyPI `agents-shipgate==1.1.0`, contract 40, Python 3.13.11.
Wheel SHA-256:
`038bdb4650d45d9c81996f60d33781b5671bfb57f006233f80e74db8a7377d33`.
Dependencies are recorded in `release-requirements.txt`.
- Source: main `d9a6d0ea5cca51b5b83eaf7743d8f1086bed1858`, version 1.1.0,
contract 41. Same release interpreter/dependencies for the paired scans.
- `replay.csv`: 27 pinned PRs × two engines = 54 head scans. Reused identical
provisional local-review config bytes on both engines (hash per row). These
configs contain unresolved setup declarations; they are not reviewed policy.
Clean exact-head trees were required; a dirty cached AI4ES tree was replaced
by a clean clone. No application code or tools were executed.
- Per engine: 23/27 empty inventories; three inventories of two tools and one
of three tools. 26/27 `insufficient_evidence`; one `review_required` with an
empty inventory. These statuses and counts are not useful PR comparisons.
- The 27 cases are a selected historical corpus. The previous 310 count is
screening only, not local runs. No population success rate follows.
- `fresh-adoption.json`: separate fresh init and exact-ref verify for four
representative PRs, on both engines (eight init + eight verify operations).
Base/head are pinned; the actual merge base equals the supplied base in all
four cases. No manually supplied base report or semantic declaration.

| PR | Fresh verify on both engines | Source/binding limitation retained |
| --- | --- | --- |
| [attest#3](https://github.com/jpka/attest/pull/3) | `missing_manifest`, zero top changes | Three old local tools; new memory tools unresolved |
| [capstone_project#3](https://github.com/zendah21/capstone_project/pull/3) | `missing_manifest`, zero top changes | SQL tools read; `load_memory` unresolved |
| [visulate-for-oracle#526](https://github.com/visulate/visulate-for-oracle/pull/526) | Exit 2, unsupported Git tree binding | Gitlink `api-server/repos/test-proj-01`, mode 160000; separate factory gap remains |
| [scopeiq#2](https://github.com/rafliogun49/scopeiq/pull/2) | `missing_manifest`, zero top changes | Five direct edges already read; ambiguous deployment root, not a missing SDK parser |

Zero cases in these fresh replays meets the ten-case bar. `insufficient_evidence`
is retained as diagnostic evidence only. A host-grant row cannot substitute for
an application-agent capability change.

### Reproduction

Use each record's public PR repository and pinned refs from `replay.csv` /
`fresh-adoption.json`. In a disposable clone, fetch the refs before analysis,
checkout the head, and verify `git merge-base BASE HEAD`. Then run, separately
with the release console script and the source checkout's `./shipgate`:

```sh
agents-shipgate init --workspace SCOPE --local-review --json
agents-shipgate verify --workspace REPO --config SCOPE/.agents-shipgate-local-review.yaml \
--base BASE_SHA --head HEAD_SHA --ci-mode advisory --format json
```

Scopes: attest `agents/attest_orchestrator`; capstone `meal_planner_agent`;
Visulate `ai-agent/root_agent`; ScopeIQ `backend`. Set
`AGENTS_SHIPGATE_AGENT_MODE=1`, unset `PYTHONPATH`, and set the source launcher's
`AGENTS_SHIPGATE_PYTHON` to the same interpreter. Do not use a shallow/partial
clone or initialize submodules. Raw CLI reports remain local; curated data here
is a research ledger, not a receipt. The 27-case historical config hashes permit
identity checks but do not reconstruct those configs; fresh init above is the
portable four-case reproduction, not a byte-identical replay of all 27 configs.

## #610 reproduction and decision

In a disposable Git repository with a committed README and an unverified
working-tree file, feed `{"changed_files": []}` to
`agents-shipgate preflight --workspace REPO --plan PLAN.json --json`.
Both builds return `state=complete`, `completion_allowed=true`, and all six
permissions true, including merge/report_complete, without a current verifier
identity. A plan listing README instead returns `agent_action_required`, a
verify next action, and all permissions false. The defect is the empty-plan
route; the docs-only control is a useful negative control.

Select planning-only completion with no publication/merge authority. Implement
that deliberately across model, schema and consumers under #610; do not silently
change a shared `complete` union or frozen predecessor schema in this research
PR. Omitting edits cannot certify a workspace. This review does not fix #610.

## Phase exit

The baseline, owner, ordered backlog, scope/absence design, #610 diagnosis and
bounded reliability selection are recorded. Weeks 2–3 implementation remains
#580 → #655 → #867. Outreach and the ten-case value claim remain gated by their
actual evidence; neither is implied by closing the Days 1–5 planning phase.
88 changes: 88 additions & 0 deletions docs/research/application-days1-5/comparison-design.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,88 @@
# Proposed application comparison inputs (#580 / #655)

Status: selected implementation design, 2026-09-24; **not a shipped API**.
This settles the Days 1–5 input contract before Weeks 2–3 implementation.

> **Status as of 2026-09-27: superseded in part.** #871 implemented part of
> this design as a new `diff --application` flag with its own advisory JSON
> schema (merged, not yet released), not through the existing capability
> projection named below. #873 (closing #580), #877 and #879 (closing #864)
> followed. See `docs/application-comparison.md` for the implemented
> behavior; this page remains the 2026-09-24 design record.

## User outcome

From a fresh unconfigured repository and precise PR refs, show what changed in
each observed agent's callable tool surface, with source evidence and named
limits. A missing manifest must not erase readable application changes. Do not
infer effects, business authority or a deployed root to obtain a result.

## One subject, independent sides

Extend the existing comparison/receipt identities, not the decision engine.
Bind requested base/head, actual merge base and compared commit/tree IDs,
engine version/contract, reader/config options, and each side's scope and input
content digests. Record discovery bounds, unread paths, excluded dependencies
and observed agent/tool wiring. Canonical ordering and deterministic hashing
must include the origin of each input selection.

Select base and head independently. A head source list may seed candidate
search on base; it cannot establish base completeness or the old deployed root.
Manifest-backed, discovery-derived and reviewer-selected scopes remain distinct
provenance. Reuse already extracted per-agent edges (#867) even when a deployed
root is unresolved; label them observed wiring rather than root reachability.

| Input situation | Required evidence and result |
| --- | --- |
| Existing app, no manifest at base | Discover/read base independently; missing configuration is not an empty surface |
| Added directory | Complete Git tree establishes absence only at that exact path; broader agent absence needs an explicit bounded subject and reviewed identity claim |
| Renamed scope | Bind old and new locations plus Git/content evidence; a file rename alone does not prove agent identity |
| Ambiguous move / split / merge | Keep separate observations and unresolved correspondence; do not invent additions/removals |
| Unread base, truncated census, missing objects | Named incomplete comparison; never zero changes or a manufactured empty base |
| Unrelated gitlink or symlink | Retain a named materialization limitation until safe bounded reading is supported; dependencies crossing it remain unresolved |
| Readable independent agents plus partial edge | Preserve proven rows with per-agent coverage limits; never label whole application complete |

## Advisory comparison versus verification

Generated comparison inputs establish source-observed structure only. They must
carry an explicit generated origin through plan, report, receipt and consumers.
They cannot be relabeled as an independently scanned, reviewed verifier base or
satisfy historical qualification bars. Existing manifest/policy/trust-root
deltas remain visible. Existing release decisions and current-control identity
rules remain authoritative; advisory rows grant no publication or merge rights.

Thus revise #655's absolute “never unavailable” aim to “show every established
structural change and precisely identify unresolved comparisons.” Missing Git
objects, dynamic wiring and ambiguous subjects are legitimate limitations.
Reject copying head inventory into base and reject keyword inference presented
as declared effect. No purpose/effect/authority/agent-binding auto-fill.

## Minimal delivery and contract review

1. Extend existing paired input selection for exact trees and independent scopes
(#580). Keep fail-closed verifier behavior until all input joins are supported.
2. Add manifest-free advisory paired extraction (#655), consumed by existing
capability projection. Reader support is unchanged in this increment.
3. Preserve per-agent direct edges (#867), then expand only the reproduced
#864 → #865 and #866 shapes. Re-run all four real examples after each change.
4. Before adding fields, choose versioned model and current JSON projections,
enumerate frozen predecessor down-projections under STABILITY.md, and test
current and legacy consumers. Field names in this proposal are concepts,
not a premature public schema or a promised new CLI flag.

## Required engineering acceptance

- Fixed added/renamed scope cases google/adk-samples#1975/#1977 from #580,
bound to its recorded SHAs; new manifest stays a trust-root change.
- Paired add/remove/unchanged tool, relocation, ambiguous rename, unrelated
deletion, multiple roots and unread/truncated base controls.
- Declarations without wiring never count as reachable; unchanged definitions
with changed bindings do count when their per-agent wiring is established.
- Identical refs, input bytes, engine/options yield identical canonical evidence;
changing either side's input, scope or generated origin invalidates identity.
- Advisory completeness cannot satisfy reviewed-base qualification or mint a
release permission. Negative tests cover forged provenance and missing base.
- Each accepted public PR has exact refs, command/engine identity, nonempty
PR-specific rows, source-reviewed tool coverage and direction, and a concrete
reviewer decision. Ten passing examples are required by #868; incomplete
examples remain failures for that acceptance even if some rows are useful.
Loading
Loading