Skip to content

feat: compare application agent wiring without prior setup - #871

Merged
pengfei-threemoonslab merged 3 commits into
mainfrom
codex/application-comparison-no-setup
Sep 25, 2026
Merged

pengfei-threemoonslab merged 3 commits into
mainfrom
codex/application-comparison-no-setup

Conversation

@pengfei-threemoonslab

@pengfei-threemoonslab pengfei-threemoonslab commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

User outcome

Review a real application-agent PR before adopting Agents Shipgate:

agents-shipgate diff --application --workspace /path/to/repository \
  --base BASE_SHA --head HEAD_SHA --scope backend --json

This is the first implementation increment for Useful application comparison without prior setup. It does not depend on merging #869. No init, manifest, baseline, purpose/effect/authority/binding declaration, or application execution is needed. This flag is unreleased; use this source checkout or a build containing the PR.

Surface discipline

  1. Headline metric: activation rate. This removes the manifest/declaration prerequisite for a first useful application PR comparison; no measured rate improvement is claimed.
  2. Existing surface: this extends diff with an application mode. Host-grant rows cannot faithfully represent per-agent callable wiring or its scoped uncertainty, so application JSON has its own explicitly advisory schema. It does not duplicate release_decision.decision.
  3. Non-goals: existing readers and binding-graph evidence only; no second verdict, authored authority, execution, LLM calls or network access. The application route and its narrower claims are now registered separately in the distribution-surface registry and parity test.

Implementation

  • Discover the base and head scopes independently from immutable Git trees, and compare existing OpenAI Agents SDK / Google ADK reader observations per source-qualified agent.
  • Show added, removed, and changed bindings with before/after signatures, schemas, binding/function locations, implementation digests, and a reviewer question. Catalog-only definitions do not become capabilities; an unresolved deployment root does not erase observed per-agent wiring.
  • Record requested refs, actual compared commits/merge base and trees, scope, engine identity, coverage, exact-blob source correspondence, and deterministic comparison identity.
  • Retain parse bounds and unresolved inputs; unread input cannot establish an empty base or removal. Exact file moves preserve identity, including handoffs. Ambiguous nonidentical moves remain partial and can use explicit old/new scopes.
  • Publish a separate advisory JSON result, retaining the existing host diff and release/verifier contracts.

Real PR reproduction and source reconciliation

Fresh clone, no Agents Shipgate files/configuration, no writes to the subject repository:

./shipgate diff --application --workspace /path/to/scopeiq \
  --base edf6d5c0a8fbb692f1e117df77eb7ba1d63ba699 \
  --head fb0d02045b357e96ebf4cdc451e1000784bf26f7 \
  --scope backend --json

ScopeIQ #2: compared, both side limits empty, exactly five added bindings:

Agent New bindings Source evidence
social_agent tool_hn_search, tool_stackexchange_search, tool_tavily_search social.py:42
synthesizer_agent python_exec, rag_query synthesizer.py:38

The exact base files contain TODO placeholders, and the exact head contains those five direct wrappers and tool-list entries. The output lets a reviewer decide whether the synthesizer should receive the Python execution and retrieval wrappers and whether the social agent should receive the three search wrappers. It does not establish sandbox correctness or deployed/runtime permissions.

Capstone #3, d146c20e937155d29d0f9ade77af9be2001f6f20...1fcb28ca619c1d681df481f4236c26775e961e4d, scope meal_planner_agent: partial, two SQL binding additions (execute_sql, inspect_schema), with unresolved load_memory explicitly retained. This is not counted as a complete qualifying case; #866 remains open.

Engineering review follow-up

Two initial self-review passes were followed by the external review on 86f1c960. This update addresses its nine inline comments:

  • Scoped uncertainty: source/agent/binding gaps affect only related candidates. Unknown implementation evidence does not erase observed additions. Uncertain candidates have not_established rows with per-side reasons; partial zero-row text explicitly denies a no-change conclusion.
  • Definition identity: use the reader's existing Python symbol and definition location, preserving name_override and separating same-named methods. No reader contract or runtime loading was added.
  • Partial clones/errors: reuse archive_fetched_tree, report the affected base/head side and hydration command. Ordinary Git process errors become ConfigError at the existing Git boundary; no new subprocess import or execution allowlist entry. The existing static-safety line pin moves with its unchanged call.
  • Scoped materialization: each side archives its own selected directory through the verified materializer, preserving link refusal. Measured --scope examples on this repository at 5.39 s versus the review's reported 84 s (local timing, not a benchmark guarantee).
  • Missing scopes: name absence and relocation flags; refuse if neither selected directory exists.
  • Surface/release docs: add gate answers, a dedicated registry row and parity entry, and an Unreleased CHANGELOG entry. Availability links to the changelog instead of hard-coding the current published version.
  • Flag/digest nits: reject explicit --max-python-files without --application, retain empty AST fields consistently (same fixture digest verified on Python 3.12/3.13), and hash the sanitized published payload so its comparison ID is reproducible.

The separate-module tools.py import limitation is reproduced in a regression and remains explicitly partial. It is existing reader work tracked by #864/#868, not claimed fixed or counted as a qualifying PR here.

Verification

  • Repository Ruff and static-only safety checks passed.
  • Application, review regressions, existing host-diff/partial-clone and distribution-parity tests passed; includes real treeless/blobless clones with base and head failures, no implicit hydration, and scoped symlink refusal.
  • Fresh ScopeIQ reproduction still returns exactly the five additions above, with no side limits.
  • Complete local test suite passed. The final test-only follow-up also verifies CLI errors with FORCE_COLOR=1, stripping ANSI escapes only in the assertion.
  • Final head 90527433a5f034c372085aa747eeb41c809b0831: CI passed, including all three test shards, combined coverage, packaging and launcher checks. Both verifier jobs passed.
  • Exact-base/head local verifier and fresh current-control read report complete. ScopeIQ and Capstone were rerun on this final code with the same results documented above.

Remaining scope

First increment supports SDK/ADK source wiring only. Unresolved imports, dynamic factories, built-ins, other frameworks, indirect helper effects, deployment-root reachability and business authority remain outside established coverage. The verified materializer retains its checks within the selected scope; scoped application comparison is not whole-repository verification. Advisory output grants no release or merge permission and is not a reviewed verifier base/receipt.

Progresses #868, #655, #867 and scoped comparison in #580. Does not close those issues or claim ten validated PRs, release-package support, market validation, or outreach readiness.

@pengfei-threemoonslab pengfei-threemoonslab left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: diff --application (86f1c96)

Verdict: request changes. This is posted as a comment because GitHub doesn't allow requesting changes on your own PR.

The overall design holds up. Both sides are materialized from immutable trees, and the dirty worktree is never read. It reuses the existing SDK/ADK readers and resolve_agent_binding_graph rather than adding a second extractor. manifest=None is threaded narrowly through the three readers. Nothing is written to the subject repo. The advisory JSON contract is separate. The size bound runs before discovery, and truncated or unparseable input stays a limit rather than turning into an empty side. Exact-blob correspondence is limited to R100 and also covers handoff targets.

The problem is that fail-closed is applied per side rather than per agent or file. On ordinary repositories that hides the rows this feature exists to show (#1 below).

What I ran: tests/test_application_diff.py passes locally (31 tests, CPython 3.12.13). I also probed the PR head in fresh repos: unrelated unchanged gaps, name_override, same-named methods, moved and misspelled scopes, treeless and blobless clones, and timing on this repository.

# Sev Finding
1 P1 Any limit on the base hides every added row in the repository, and any limit on the head hides every removed row. The text output then says "No established … changes"
2 P2 _definition can't find name_override tools, or a tool that shares its name with any method. Body changes to those tools produce no row, and they trigger #1
3 P2 Partial clones: a treeless clone crashes with a traceback (exit 1). A blobless clone exits 2 with raw git text. Host diff in the same checkouts names the #817 repair
4 P2 Both trees are archived unscoped: 84 s on this repo even with --scope examples. Archiving one commit scoped takes 1.6 s, versus 40.1 s unscoped
5 P2 Surface rules: no answers to the Surface-discipline questions for the new schema and flags. The capability_diff registry row still says "Host route only". No CHANGELOG ## Unreleased entry
6 P3 Absent scope: a moved directory reports compared with REMOVED rows and no hint. A misspelled scope exits 0 with not_established and no mention that the path doesn't exist
7 nit Stray --max-python-files, digests that depend on the interpreter, comparison_id computed before redaction, error kind (inline)

Heads-up (not a defect in this diff, but it limits what fixing #1 can recover). A common layout is @function_tool definitions in tools.py and an agent.py that imports them. With this PR, that layout comes out partial with zero rows on every PR:

  • Each discovered file becomes its own ToolSourceConfig id.
  • resolve_agent_binding_graph matches a structural binding only against tools with the same source_id (core/agent_bindings.py:212-219).
  • So assistant → lookup "matched 0 canonical tools".

ScopeIQ works because its wrappers live in the same file as the agent. It's worth measuring how often this happens in the #866/#868 corpus before counting qualifying PRs.


🤖 Generated with Claude Code

Comment thread src/agents_shipgate/cli/application_diff.py Outdated
Comment thread src/agents_shipgate/cli/application_diff.py Outdated
Comment thread src/agents_shipgate/cli/diff.py
Comment thread src/agents_shipgate/cli/application_diff.py Outdated
Comment thread src/agents_shipgate/cli/application_diff.py
Comment thread README.md Outdated
Comment thread src/agents_shipgate/cli/diff.py Outdated
Comment thread src/agents_shipgate/cli/application_diff.py Outdated
Comment thread src/agents_shipgate/cli/application_diff.py Outdated
@pengfei-threemoonslab
pengfei-threemoonslab merged commit a430e81 into main Sep 25, 2026
12 checks passed
pengfei-threemoonslab added a commit that referenced this pull request Sep 27, 2026
The ROADMAP "Selected execution" section and the research README's
issue table describe as future work items that have since merged: #870
closed #787, #871 merged diff --application, #873 closed #580 and #879
closed #864. Add a dated "Status as of 2026-09-27" note to each naming
what merged and what remains open, leaving the plan and snapshot as
written. comparison-design.md called itself the selected design,
delivered through the existing capability projection and "not a
promised new CLI flag", but #871 added a new flag with its own advisory
JSON schema; add a status banner saying it is superseded in part.

Renaming "Lead wedge (focus)" broke the inbound link from
docs/engineering/v1-release-readiness.md to #lead-wedge-focus; keep the
old slug with an anchor above the renamed heading, and do the same for
the renamed post-1.0 adoption heading.

git diff --check failed on a trailing space in the research README and
CRLF endings on every line of replay.csv. Strip the space and convert
the CSV to LF; nothing pins its digest and no field holds a newline.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pengfei-threemoonslab added a commit that referenced this pull request Sep 27, 2026
…869)

* docs: record Days 1–5 application comparison baseline and design

* docs: date the Days 1-5 status; keep roadmap anchors; fix whitespace

The ROADMAP "Selected execution" section and the research README's
issue table describe as future work items that have since merged: #870
closed #787, #871 merged diff --application, #873 closed #580 and #879
closed #864. Add a dated "Status as of 2026-09-27" note to each naming
what merged and what remains open, leaving the plan and snapshot as
written. comparison-design.md called itself the selected design,
delivered through the existing capability projection and "not a
promised new CLI flag", but #871 added a new flag with its own advisory
JSON schema; add a status banner saying it is superseded in part.

Renaming "Lead wedge (focus)" broke the inbound link from
docs/engineering/v1-release-readiness.md to #lead-wedge-focus; keep the
old slug with an anchor above the renamed heading, and do the same for
the renamed post-1.0 adoption heading.

git diff --check failed on a trailing space in the research README and
CRLF endings on every line of replay.csv. Strip the space and convert
the CSV to LF; nothing pins its digest and no field holds a newline.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
pengfei-threemoonslab added a commit that referenced this pull request Sep 29, 2026
… in the agent's function (#865) (#903)

* Read a Google ADK tool a local factory builds, and a tools list built in the agent's function (#865)

visulate/visulate-for-oracle#526 bound `save_memory_tool = create_save_memory_tool()`
in its root agent's builder, where the factory returns `FunctionTool` around a
nested function; the reader stopped at the local assignment. A factory call
bound once and unconditionally (in the agent's function or at a module's top
level), or inline in `tools=[...]`, is now followed as syntax to the factory's
one unconditional `return` of `FunctionTool(inner)`, a plain function, a name
bound once to one of those, or another factory's call, up to four deep; the
tool is the wrapped function, with the factory recorded in `import_path`.

A factory that returns from several places or under a condition, recurses, is
decorated or a generator; a wrapped function that is decorated, a parameter,
or changed or handed on (Visulate's delegate sets `__name__`); a tool changed
where it is bound; and a factory held outside the read scope stay named with
why (`factory_return`, or the resolver's reason). A third-party factory keeps
the answer it had. On visulate#526 at `ai-agent`, the two memory tools become
not_established candidate additions beside the nine named delegates.

MuhammadVT/smart-assignment#46 built `tools = [...]` in the agent's function
with a conditional `tools.append(...)`: read as a dynamic tools expression, it
lost the unconditional tools. Such a list with append/extend/insert/`+=` is
read member by member; an addition under a condition or in a loop is named on
the agent and never read as bound; any other use keeps it dynamic.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Resolve repository-local imported tools for Google ADK and OpenAI Agents SDK (#864) (#879)

* feat: resolve repository-local imported tools for ADK and SDK readers (#864)

A tool an agent binds from another module was unresolved at the module
boundary, so a PR adding one produced no row. A shared resolver
(inputs/python_imports.py) follows a tools-list reference through static
imports, package re-exports, module-qualified access and plain aliases to
one function definition, reading only regular .py files inside the read
directory through the bounded input reader, never importing or running
them. Every stop is a named reason.

- Google ADK: imported names, module.function, alias = function, and
  FunctionTool/LongRunningFunctionTool wrappers (inline, assigned, or
  built in the imported module). One tool per definition; same-named
  functions in different modules stay distinct via binding locators.
- OpenAI Agents SDK: names and module.function reaching the SDK's
  @function_tool, with guard evidence for imported definitions.
- diff --application: rows carry import_path evidence (modules, lines,
  digests) outside compared meaning; unresolved references are scoped
  to their agent with the reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): import resolution never produces a false or complete answer (review round 1)

Adversarial review of the shared resolver found answers that read as complete
while wrong:

- One SDK agent binding two same-named definitions from different modules
  kept the last and reported a false CHANGED row. Both readers now bind
  neither, whatever the list order, and name both definitions.
- A name the function building the agent binds itself (a local import, a
  parameter) was resolved through the module's binding. It is now a named
  stop (`local_binding`); a nested function that is the only definition of
  its name is that definition.
- `import a.b` then `a.b.f` read `a/__init__`'s own `b`. It now reads the
  submodule, and several `import a.x` statements are not a rebinding.
- An ADK wrapper warning shared by two agents scoped its gap to one of them.
  Every unresolved-reference record now gets its own gap.
- `scan` counted one definition twice when an import reached a module another
  configured source also reads, making `{tool: ...}` selectors ambiguous. The
  catalog keeps one observation, and the binding graph reaches it through the
  exact definition locator the reader resolved when the edge's own source has
  none.
- The module-binding walk climbed a parent chain per node; it is now linear
  (a 744 KB nested module: 32 s -> 3.7 s).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): read a reference where it is used (review round 2)

- A builder's own `from support import lookup` is followed like a module-level
  import (`ImportResolver.resolve_local_import`) instead of stopping, for bare
  names and ADK `FunctionTool(func=...)` arguments alike.
- `nonlocal` follows the outer function's binding; a name a scope binds more
  than once is a named stop.
- `tools.lookup = ...` / `setattr(tools, "lookup", ...)` in the module that
  binds `tools.lookup` is a named stop.
- A factory's own `toolset = McpToolset(...)` / `tool = FunctionTool(...)` is
  read through the existing toolset and wrapper paths again, so the MCP
  endpoint and toolset checks return.
- A module-level ADK agent binds the module-level `def`, not a nested one the
  flat function map happened to keep last (pre-existing on main).
- An SDK list variable bound twice anywhere in the file, or changed in place,
  is dynamic.
- The `scan` dedupe runs once where sources are loaded, so guard association
  and inventory completion see the same tools; a source an inventory completes
  keeps its observation; `./` spellings match.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): read a list and a flat name where they are bound (review round 3)

- SDK tool lists are read through the scope that binds the name at the
  construction. They are read only when that scope binds the name once, to
  a literal list, and no change site anywhere in the file (a method call,
  a subscript store, `global` / `nonlocal`) changes that binding. Each
  change site is indexed once, so the reader is linear. A module-level
  list's names are resolved where the list is written, not in the
  building function's scope (R3-1, R3-4).
- A flat function, wrapper or toolset map answers a module-level
  reference only when its entry is the module's single top-level binding.
  Otherwise the resolver answers (a module import wins over a nested
  `def`). If the resolver cannot establish the binding, the same-named
  definition is named, with a tool-scoped issue, and its row is never
  established (R3-5).
- A wrapper's `func` is read where the wrapper is written.
- An enclosing package `__init__.py` that reassigns the definition is a
  named stop (R3-3). `_dotted` is iterative, so a chain thousands of
  attributes deep no longer crashes (R3-2).
- The scan dedupe drops the removed copy's guard evidence too (R3-7).
- The CHANGELOG states the inventory-completion ambiguity (R3-6).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): a guess is scoped to its tool, and a list's second handle counts (review round 4)

- A guessed ADK binding no longer marks the agent's list incomplete.
  - That had dropped every per-tool `scan` finding of the agent for one
    name (R4-2). `scan` is back to the round-3 behavior: a shadowed,
    medium-confidence definition. The PR #400 tests pass unmodified again.
  - The comparison reads `tool_issues` as a tool-scoped gap and an
    agent-scoped gap. The row is `not_established` on whichever side it
    is present, added and removed included (R4-1).
- `x = FunctionTool(func=x)` right after `def x` wraps that `def` and is
  not a guess. That was noise on byte-identical files.
- The SDK list reader counts `alias = TOOLS` and `TOOLS` passed to any
  call but a read-only builtin or logging method as a change (R4-3).
- An attribute of the resolved name reassigned in an enclosing package's
  `__init__.py`, or in a module it imports relatively, is a named stop.
  This covers an alias spelled from a package above (R4-4).
- Same-file symbol lookup is a dict, not a scan of every tool per
  reference.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): scope a guess to its candidates; read list arguments and patches precisely (review round 5)

- A guessed ADK binding's reason covers the tool it bound. It also covers
  every tool the module's bindings of that name could give the agent: a
  `def`, an import, `x = other`, or `x = FunctionTool(func=f)`. It covers
  the whole agent (`ANY_TOOL`) only when one of those cannot be named.
  - A guess that binds a name the agent already lists still leaves a gap
    (R5-1).
  - A `try: import … except ImportError: def …` fallback no longer hides
    the agent's other changes (R5-2).
- A literal list passed to a call counts as changed unless the call is:
  - a read-only builtin, `pprint`, or a logging method;
  - an SDK `Agent` or a copy (`clone`, `replace`) reading its own
    `tools=`;
  - a function the module can resolve that leaves that parameter alone.

  Only names bound to a literal list are followed into a callee, so no
  other module is read for an unrelated call (R5-3).
- The package-patch check counts only attributes rooted at an import. It
  parses the modules an `__init__` imports on a separate bounded budget,
  so a package that re-exports many modules no longer exhausts the
  resolution budget (R5-4, R5-5).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): only reads keep a list literal; follow a guess's candidates (review round 6)

- A literal tool list stays readable only while every use of it is a read.
  Allowed reads: iteration, indexing, comparison, a truth test, formatting,
  a read-only builtin or logging method, an agent's or copy's own `tools=`,
  or a function whose every use of that parameter is such a read. `+=`, a
  return, a tuple, `*args`, `**kwargs` or storing it in another container
  makes it dynamic (R6-1).
- A guess's candidate tool names are followed to their definitions: an
  import through the resolver, or by its imported name when the module is
  outside the scope; `x = other` and `FunctionTool(func=f)` through the name
  they spell. A candidate that cannot be followed, or a wildcard import,
  scopes the reason to the whole agent (R6-2, R6-3).
- A module that a package `__init__` imports relatively and that cannot be
  read (a link, a missing file) is a named stop, cached with the package
  (R6-4). A module read for patches is the same object a later resolution
  uses (R6-5).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): spreads and tests are reads; a package's routine imports never stop (review round 7)

- A literal list is still read when it is spread (`[*COMMON, x]`,
  `print(*TOOLS)`), tested (`TOOLS or …` in a condition), or read through
  a dict method (`.get`, `.keys`, `.values`, `.items`). `x or y` whose
  value is the list is judged by its own use. A `globals()` or `vars()`
  call in the module makes its lists dynamic (R7-3, R7-1 b4).
- The package-patch scan treats an import above the scope as the read's
  boundary, as every import does. It skips an optional module imported
  under `except ImportError`. A link, or a missing module that is not
  optional, still stops (R7-2).
- Docs: drop a duplicated sentence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): code that runs first is checked, and unread code keeps a binding named (review round 8)

- Every module on a resolution chain (the agent's own file included), the
  `__init__.py` of every package enclosing one, and every in-scope module
  those import are checked for a reassignment of an attribute named like a
  step of the chain. `import patches` in the agent's file, or in a module
  the chain re-exports through, is now a named stop (R8-2). The defining
  module handing its function on (`registry.lookup = lookup`) does not
  count, and every location is kept, so one never hides another.
- A relative import in any of those modules that climbs above the scope
  runs code that is not read. It no longer disappears: the tool is named
  and its row is `not_established` with that import in the reason, in both
  readers; for `scan` the ADK module stays at medium (R8-1, which round 7
  had made silent). An import under `if TYPE_CHECKING:` never runs and is
  skipped.
- `sys.modules`, however spelled, and importing a module by `__name__`
  reach its lists like `globals()` does: an SDK list there is dynamic
  (R8-3).
- A redirecting package `__getattr__` stays a documented residual (R8-4):
  a gap for every hook on a chain would make #864's own attest rows
  `not_established`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): every spelling of unread code keeps the caveat; a lazy package hook is proven (review round 9)

- A caveat survives a `FunctionTool(func=f)` wrapper the imported module
  builds (R9-1), and an absolute import spelled through a directory above
  the scope (`from svc.patches import ...` with scope `svc/app`) is a
  caveat like the relative one (R9-2).
- The defining module's hand-on exemption no longer covers a patch through
  its own import (`import tools as _me; _me.lookup = ...`, R9-3), and only
  `typing`'s `TYPE_CHECKING` skips a block (R9-4).
- Ordinary code no longer stops a resolution (R9-5): every location an
  ambiguous import could mean is read, a generated `*_pb2` module is the
  boundary and any other missing relative module a caveat, and the patch
  scan reads up to 1024 modules.
- A package `__getattr__` is established only when every return it can
  reach for the name gives that submodule (`import_module(f".{name}",
  __name__)`, `from . import name`); attest's two lazy loaders stay
  established, a redirecting hook is named (R8-4).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): the repository says what is its own code; harden the lazy-hook idiom (review round 10)

- Whether an absolute import no file in the scope provides is the
  application's own code is read from the repository: the compared
  commit's tree for `diff --application` (set per side), the checkout for
  `scan`. A module or regular package at the root or under `src/` is; a
  directory without `__init__.py` only when it holds the named submodule.
  `from common.patches import ...` with scope `svc/app` is a caveat
  (R10-1), and SDK apps under `agents/` import the SDK again (R10-3: the
  round-9 ancestor-name rule had caveated every tool there).
- The scope spelled from the repository root (`svc.app.tools` with scope
  `svc/app`) is read inside the scope, so it resolves and a patch module it
  names is checked instead of guessed.
- The lazy-hook idiom requires an undecorated hook that never rebinds its
  parameter, `importlib` / `import_module` bound only by importing them,
  and a returned local bound exactly once, counting `for`, `with`, walrus
  and `except` targets (R10-2).
- A generated `_version` module is the boundary like `*_pb2` (R10-4); the
  patch-scan budget message names what it scans.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): every directory above the scope is an import root; unread links and subverted hooks are named (review round 11)

- An absolute import is looked up at the repository root, under `src/`, and
  under every directory between the root and the scope, so `backend/app`
  importing `common` from `backend/common` is a caveat (R11-1).
- A root entry that is a symbolic link or a submodule, spelled by the
  import, is unread code: a caveat (R11-2). `scan` outside a checkout reads
  the three directories above the scope instead of nothing (R11-3).
- A store into `sys.modules` in code that runs first is a named stop, and a
  package hook is not trusted when the package rebinds `__name__`, patches
  `importlib`, or stores into `sys.modules` or `globals()` other than the
  idiom's own cache (R11-4).
- The scope spelled from an import root is read inside the scope only
  through a regular package (R11-5).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): a package is not an import root; a sys.modules store is read by its key (review round 12)

- A directory between the repository root and the scope that holds an
  `__init__.py` is imported through its parent, never from the path: SDK
  apps under a regular `app/agents/` package import the SDK again, and
  `app/types.py` no longer shadows the standard library (R12-1, a round-11
  regression).
- A store into `sys.modules` (`[...] =`, `setdefault`, `__setitem__`,
  `update`) is a named stop when its key names a module on the chain or is
  built on `__name__`, a caveat when it is computed (a plugin loader's
  `spec.name`), and nothing when it names another module (R12-2, R12-3).
- `globals().update(...)` / `setdefault` in a package disqualifies its hook
  (R12-3).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(#864): packages above the scope are read; sys.modules and globals() by allow-list (review round 13)

- Every directory between the repository root and the scope is an import
  root again, package or not: a service run from `backend/` with a stray
  `backend/__init__.py` names `common.patches` (R13-1, a round-12
  regression). A standard-library name, and the scope's own package on the
  way to it unless it holds the name imported, are not the repository's
  code, so SDK apps under `app/agents/` still import the SDK (R13-5).
- The `__init__.py` of every package above the scope, and each module it
  imports, are read through the repository layout (the commit's tree, or
  the checkout) for the same reassignments (R13-2).
- `sys.modules` and `globals()` are read by allow-list. A store whose key
  names a module on the chain, a package above one, or the framework's own
  modules is a named stop; `__name__` plus a literal is that module's own
  name; any other use is a caveat; a module rebinding its own name through
  them is a reassignment; a change to `__path__` is a caveat (R13-3,
  R13-4, R13-6). A package that rebinds `__getattr__` or `__path__` does
  not have a trusted hook.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read what the packages above the scope import, and every spelling of a module's own object (#879 review, round 14)

- Above the scope, follow each import the way the in-scope reader does:
  every package on the way to the module and each submodule named
  (`from .hooks import patches`, `import svc.lib.util`). A reassignment
  there counts only when rooted at the scope or at something that cannot
  be found; guarded and generated imports are exempt.
- A standard-library name is exempt only at a root that is a regular
  package; interpreter-preloaded modules are always exempt.
- Allow-list reads: comparisons, spreads, iteration,
  pkgutil.iter_modules(__path__), namespace keywords
  (get_type_hints(globalns=globals())), and patch.dict/setitem by key.
- The module's own object through an alias, sys.modules.get,
  import_module(__name__), __dict__/vars() stores, a computed setattr, or
  handed to a function is a reassignment or caveat, never silent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* A local named vars or globals is a variable, not the builtin namespace (#879 review, round 14 corpus)

siada-cli's DebugUtils.dump binds `vars = stack[-2][-3]` and iterates it;
the bare-name rule read it as the builtin handed on and named a real
definition change not_established.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Only a two-part patch on another module is exempt above the scope; read the scope's modules an ancestor imports and the module object anywhere (#879 review, round 15)

- R15-1: above the scope, a patch is exempt only when it sets one
  attribute of another module file outside the scope that its package
  binds nothing else under. A longer path, a name imported from a module,
  a package attribute shadowing the submodule, or a module alias
  (`_t = tools`) counts, in both readers.
- R15-2: a scope module that an ancestor __init__.py imports is read like
  the chain's own.
- R15-3: a module file wins over a same-named directory without
  __init__.py.
- R15-4: a link's blob text is never read as source.
- R15-5: the module object used anywhere but an attribute, a plain alias,
  a comparison or a reader is a caveat. So are __dict__/vars().update,
  getattr(m, "__dict__") stores, f_globals, builtins.globals under another
  name, a reader of the module's own under a reader's name, `globals =
  globals`, and __path__ changes above the scope (pkgutil.extend_path
  aside). get_type_hints(fn, globals()) is a read.
- R15-6: above-scope files are read in batched `cat-file --batch` per
  directory (1100 modules: 76 s -> 3 s). The read bound is named once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read a frame's namespace by allow-list; batch a directory only while it is small (#879 review, round 16)

- R16-1: `sys._getframe(1).f_globals.get("__name__")` in a logging helper
  is a read. A frame's namespace is read by the same allow-list as
  globals(); any other use is a caveat.
- R16-2: above-scope reads batch a directory's modules in one
  `cat-file --batch` only while the directory holds at most 16 MB of
  Python. A directory of generated or vendored modules is read file by
  file, as asked (peak RSS 305 MB -> 87 MB on the 60 MB case).
- Docs: `sys.path` / `sys.meta_path` changes that make a same-named module
  elsewhere the one imported are not followed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Bind the located definition, prove readers by their binding, check the lazy hook's import (#879 PR review)

- The resolved locator decides between same-named definitions even when
  the edge's own source holds exactly one. After source deduplication, a
  `tools.lookup` imported as `shared_lookup` no longer binds the local
  `lookup`.
- A read-only reader is proven by its binding:
  - a bare builtin only when nothing binds it and no star import may;
  - a standard-library reader only when imported from that module
    (`json.dumps`, `from pprint import pprint`);
  - a logging method only on `logging` or a `getLogger()` logger.
  A `print` imported from the application, or an `.info()` on its own
  object, makes the list dynamic.
- A lazy `__getattr__` is the submodule idiom only when its alias imports
  the submodule asked for: `from . import alternate as memory` is not.
perf: read materialized Git tree blobs through batched cat-file (#878)

_materialize_isolated_tree spawned one `git cat-file blob <oid>` per file
in both its entries and links loops. Once `diff --application --scope .`
materializes the whole tree (#877), that dominated the run: one side of
TencentCloud/CubeSandbox took ~180 s to archive (the #686 cost class).

Blobs are now read by _isolated_blobs: one `cat-file --batch-check` types
and sizes every object (a missing or non-blob object refuses with a
ConfigError before any content is read), then `cat-file --batch` runs of
at most 64 MiB each return the content, with strict framing checks. Only
full object IDs reach the batch. The link-text reads in
_scope_through_boundary_links use the same reader.

Every blob is still hashed against its tree entry's oid before it is
written; path handling, containment and escape checks, blobs-before-links
order, placeholders and the final digest comparison are unchanged. The
reads go through _run_git_dir (now accepting input=) and the existing
_run_process boundary, so no subprocess call site is added; only the two
line pins in test_adapter_static_only.py moved.

diff --application --scope . with #877: CubeSandbox 412 s -> 46 s,
dlt 203 s -> 45 s, identical rows.
fix: compare application wiring past unrelated links and submodules; read google.adk.Agent (#877)

* fix: compare application wiring past unrelated links and submodules

`diff --application` at the root scope took the unscoped archive route,
which refuses every symlink and gitlink. On 2026-09-25, 15 of the first
27 runs against open third-party SDK/ADK PRs exited 2 on a path the
reader never opens: `CLAUDE.md -> AGENTS.md`, a linked skill directory,
a `VERSION` link leaving the tree, a vendored submodule.

- Materialize every scope, the root included, through the scoped
  verified materializer, which recreates links rather than refusing them
  and packs the tree instead of the history.
- Never read a Python input through a link: a target this scope already
  reads is compared at its own path, and any other in-scope target is a
  gap over the link's path. Reading the alias made one agent two
  ambiguous ones and hid the real file's change.
- Record gitlinks behind an opt-in `archive_tree(record_gitlinks=True)`,
  materialized as the empty directory an unpopulated checkout leaves.
  An unchanged gitlink commit is named in limits; any other is a coverage
  gap over its path. Other archive callers still refuse.
- Stop turning the host-configuration census's link count into
  application coverage gaps.
- Read `google.adk.Agent`, the package-root re-export, as an ADK agent
  constructor (`from google.adk import Agent`).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: keep unread links and scope-level submodules from reading as removals

PR #877 review found two cases where unread input became conclusive
negative evidence:

- Discovery drops a path that does not resolve, so replacing `agent.py`
  with a dangling link read as a definite removal, `compared`, with no
  gap. The dropped host-census gap had been the only fallback, and it also
  covered a source directory replaced by an absolute or dangling link.
  Census every link under the scope directly, without following it:
  gap a `*.py` link that does not alias an input the scope already reads,
  a directory link holding Python outside the scope, and a link resolving
  to nothing in the tree where the other side reads source at or beneath
  it. An unchanged `agent/VERSION -> ../../VERSION` still changes nothing.
- A gitlink at the selected scope itself gave a gap with source ".",
  which covers no relative binding path, so `agent.py` read as a definite
  addition or removal. Make that gap scope-wide.
fix(diff --application): key SDK identity on imports; unobserved agents are not removals (#873)

* fix(diff --application): key SDK identity on imports; unobserved agents are not removals

On speechmatics/speechmatics-academy#142, a LiveKit voice agent moved its
tools into an Agent subclass that passes them through super().__init__.
`diff --application` printed `compared` with four false REMOVED rows.

Two defects:

- Framework identity. Discovery counted any bare `@function_tool` as the
  OpenAI Agents SDK, and the SDK reader recognized `function_tool` and
  `Agent` by spelling alone, so `livekit.agents` symbols were read as the
  SDK's. Both now key on import provenance: the absolute `agents` /
  `openai_agents` package, or an unimported name (unless a foreign
  wildcard could supply it). Relative imports are no framework's signal.

- Unobserved is not removed. An agent observed on one side whose file on
  the other side still assigns its name (or passes it as `name=`), via a
  construction no reader supports (subclass, factory, Agent[Ctx], clone),
  now records a scoped coverage gap for that agent: `partial`, rows
  `not_established`. Handoff-only references do not count as an observed
  construction. A genuinely deleted agent stays an established removal.

Regression tests fail on the unfixed tree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(diff --application): resolve SDK names per scope; imports keep an agent named

Addresses the two P2 findings on #873.

- Import provenance was flattened across lexical scopes: one file-wide
  import map, accepted if any origin was the SDK, so a LiveKit
  `Builder(...)` became an SDK agent when a sibling function imported the
  SDK under the same alias. `_SdkNames` now resolves each spelling in the
  scope that uses it (nearest binding scope, class bodies skipped for
  nested code, global/nonlocal obeyed; decorators and defaults in the
  enclosing scope). An import there must be the SDK's; a parameter or
  local assignment is not. `function_tool` is decided per decorator node,
  not as a file-wide set of spellings. The walk is iterative.

- The missing-agent check ignored import bindings, so
  `from agent_factory import agent` (or `exported as agent`) still gave a
  definite removal. Import aliases now count as the file still naming the
  agent.

Regression tests for both fail on the previous head.
feat: compare application agent wiring without prior setup (#871)

* feat: compare application agent wiring without prior setup

* fix: address application comparison review findings

* test: normalize colored CLI errors in review regressions
fix(ci): prevent checkout module shadowing in non-GitHub recipes (#870)

* fix(ci): prevent checkout module shadowing in installation recipes

* test(ci): cover quoted and qualified Python installation commands
Read agent launches in CI workflows (#823) (#850)

* Read agent launches in CI workflows (#823)

The workflow grant read triggers, token permissions, reusable-workflow
secrets (#693) and step `uses:` references (#771), and nothing about how a
coding agent is launched inside a job. Changing a claude-code-action's
`claude_args` from `--allowedTools "Read"` to `--permission-mode
bypassPermissions --allowedTools "Bash(*)"`, adding a `claude -p
--permission-mode acceptEdits` run step, or checking out the pull request
head in a `pull_request_target` job each gave "No static host-grant changes
detected", and a move to `issue_comment` with `pull-requests: write` gave a
row with no agent context.

The workflow grant now lists, as text that is never executed, fetched or
evaluated:

- `agent_launches[]`: a step whose `uses:` is anthropics/claude-code-action,
  anthropics/claude-code-base-action or openai/codex-action (any ref, any
  case) with the documented permission inputs it sets, or a `run:` that is
  one literal simple command starting with `claude -p/--print` or `codex
  exec`, with its documented permission flags under their primary spelling.
  The prompt, `--model` and undocumented flags are not compared.
  `job_secrets` names the secrets the job references, as context only.
- `checkout_refs[]`: each actions/checkout step's `with.ref`, null for the
  default.

Each job's multiset of launches and refs is compared, never the step label,
so a rename or reorder is quiet; a difference is one `changed` row on the
existing workflow row naming `job/step` and both values. Direction is
claimed only by documented rules a job's launches gain, read from literal
values: bypassed permission checks (either spelling), bypassed approvals
and sandbox, a danger-full-access sandbox, `safety-strategy: unsafe`, or a
`*` user gate. Those raise `workflow_agent_widened_<added|changed>` and
make the row widened; every other edit, including `--allowedTools
"Bash(*)"` (#824's to rate), is `changed`. A workflow row whose workflow
runs an agent ends its `why` with the untrusted-input trigger, write
scopes, secrets and pull request checkout beside each agent step; it is a
note and moves no direction.

A compound `run:`, an expansion or an expression is `unresolved`, publishes
none of its text and records a non-blocking coverage issue naming
`job/step`, as an unread secret value does (#693): coverage stays
complete, adding one is a row that claims no effect, and an edit inside one
is not reported. Scripts, composite (#701) and unknown actions, and agents
reached through npx/timeout/sudo are listed as unread surfaces. Every value,
ref and secret name goes through the #802 label redaction; a rewritten one
is null with `redacted` and is neither published nor compared.

Contract 40 and host-grants 0.6 shipped in 1.1.0, so host-grants
inventory, baseline and drift move to 0.7 (the 0.6 schema files are
untouched) and the runtime contract to 41. A 0.4-0.6 baseline holding a
workflow grant is incomparable (`baseline_workflow_agent_launches_unavailable`);
one without a workflow stays comparable. Verifier 0.20 and capability diff
0.3 do not move. No check id is added or removed and `check` decides as
before.

Docs: the support page (tables, rules, note, limits, a new unread-surfaces
bullet), a STABILITY "Migration Note: Unreleased", CHANGELOG `## Unreleased`
above 1.1.0, the agent contract page, the distribution-surfaces
`capability_diff` row and its parity comment, the Stop hook note, version
tables and pins, and a rebuilt llms-full.txt. The pilot ledger's
source-tree column was re-measured on the Route H fixture: identical cells
to the 1.1.0 engine apart from contract 41 and inventory 0.7, with
byte-identical diff rows.

Tests: new tests/test_workflow_agent_launches.py covers the four
reproduction cases, what is and is not read, every unsupported shape,
direction rules and non-rules, the acceptance's negative controls,
redaction with a CLI canary sweep across every published output, the 0.6
baseline migration, schema validation, and the same row on diff, verify,
the PR comment, check, the control envelope and the Stop hook. Existing
tests move to 0.7/41. The host-config and cold-start replays reproduce
their committed outcomes and run-of-record scores, and the sample goldens
are unchanged.

Closes #823

* Read no widening rule from an agent input that holds an expression (#823)

The support page and STABILITY say a documented rule is read from literal
values only, so a value holding `${{ }}` never meets one. The `*` user
gate did not apply that: `allowed_non_write_users: "${{ vars.USERS }}, *"`
raised workflow_agent_widened_changed. Every rule now skips a value holding
an expression, as the claude_args and codex-args rules already did, and a
test pins the gate case.

* Address review cycle 1 on agent launches in CI (#823)

F1. `claude_args` and `codex-args` were split with the `run:` shell
tokenizer, which gave up on a newline or an unquoted `(`, so a widening
written the way the actions document it gave `changed`: `claude_args: |`
on several lines, `--allowedTools Bash(git:*) --dangerously-skip-permissions`,
a `# comment` line, or a multi-line `codex-args`. Each input is now split
the way its action splits it. The Claude actions' parse-sdk-options.ts
drops full `#` lines, makes `()|&;<>` literal and reads the rest with
shell-quote (newlines are whitespace, `$NAME` is empty, an unquoted `#`
ends the input, a `--` word is always a flag); codex-action reads a JSON
array of strings or string-argv. The `run:` tokenizer is kept for `run:`
steps only. The published `claude_args` is the text the action parses, so
a full-line comment is neither published nor compared.

F2. Any URL path made a whole setting `redacted`, so a
`--dangerously-skip-permissions` beside `https://example.com/style-guide`
gave no row, and every `plugin_marketplaces` value was never compared,
under a limit that wrongly called it credential-shaped. The documented
rules are now decided from the declared text when the workflow is read,
before anything is withheld, and published on the launch as
`widening_rules` (rule and setting), which the comparator keys on, so
redaction never hides a rule. A URL publishes its scheme and host with
`<redacted-path>`, as an MCP server URL does (#723), and the rest of the
value is published and compared. Other credential-shaped text in a
setting or checkout ref (a token shape, an assignment, a bearer or header
value, URL userinfo) is published redacted and makes the workflow a
blocking limit through `_uncompared_workflow_text`, as a redacted step
reference does (#767).

F3. JSON in `settings`, `mcp_config`, `--settings`, `--mcp-config` or any
argument word was published verbatim, env values and apiKeyHelper
included. A JSON object now publishes what `.claude/settings.json` and
`.mcp.json` publish: key names, with `env`/`headers` values, apiKeyHelper
and secret-named values `<redacted>`, as canonical JSON. A codex
`--config` override under env, headers or a secret-named key publishes
`<redacted>`. Text that starts like JSON and does not parse is withheld
(`unparsed_json`, a non-blocking limit). The agent-launch canary sweep
now carries JSON-shaped canaries and their digests across diff, audit,
check, verify, the PR comment and verifier.json, and a second sweep
covers the refused credential-shaped case.

Nonblocking: a rule gained where the job launched that agent before only
in an unread form is named and not claimed (the `unknown_before` rule); a
quoted word starting with `#` no longer makes a `run:` compound;
`anthropics/claude-code-action/base-action` is read as the base action;
the STABILITY note says a prompt after a variadic flag is compared; and
`__all__ =[` is spaced.

Host-grants 0.7 is unreleased, so `widening_rules`, `unparsed_json` and
the base-action agent value extend it in place; the 0.7 schema files are
regenerated. The support page, STABILITY migration note, CHANGELOG
Unreleased entry, contract summary, integrations Stop-hook sentence and
the capability_diff distribution-surface row say the same.

* Address review cycle 1 on agent launches in CI (#823)

The second review of #850 at 25c13ce8 found two P1 and four P2 defects.
The branch is rebased onto origin/main daa4ad5f.

Attached JSON values are withheld (F1). A --settings={...} or
--mcp-config={...} word inside claude_args, and codex's attached
-c<override> / -c=<override>, published their env, header and apiKeyHelper
values verbatim, because only a word starting with "{" was withheld.
_withheld_words now splits a --name=value word and withholds the value, and
reads codex's attached -c through _withheld_config, as clap reads it. The
canary sweep carries both spellings.

Redacted prose no longer refuses the comparison (F2). The #802 label
redaction rewrites ordinary prose ("never print bearer tokens",
"Authorization: headers"), and a redacted agent setting made the workflow a
blocking limit. That hid every row beside it and made every check
incomparable while the workflow existed. The rules are already read from the
declared text, so a redacted setting is now compared by its published text
and its widening_rules. uncompared_agent_launch_texts names it as a
non-blocking limit, and the row cell shows the redacted text. A redacted
checkout ref still refuses, as a redacted step reference does (#767),
because it names the code a job runs.

A renamed job's launch moves its rules (F3). _agent_rule_gains keyed a rule
on its job, so renaming a job that launches a bypassing agent was a
widening. agent_rule_gains now pairs a rule one job gains with the same rule
another job lost, when the launch that met it left that job: the job no
longer launches that agent, or the same launch (agent_launch_key less the
job) now runs in the gaining job. The why names the move. A second job
gaining a rule, or a different launch gaining one while the first job still
launches that agent, still widens.

An expression no longer turns off every rule, and the row says what it
leaves unread (F4). Rules are read from literal text a ${{ }} expression
cannot reach:
- the words of claude_args or codex-args before the first expression, less
  the word it touches and any quoted run still open at it;
- the elements of a JSON-array codex-args before the one holding it;
- the gate entries that hold none.
A setting holding an expression is published with holds_expression (the
unreleased 0.7 schema extends in place), and a row that changes it says the
text the expression reaches is not read. A gain where the job's launch held
an expression before, in the input the rule is read from, is named and not
claimed, as unknown_before is. That also fixes a false widening at
25c13ce8: replacing --model ${{ vars.CLAUDE_MODEL }} with --model opus
beside --dangerously-skip-permissions was reported as gaining the bypass.

Two non-blocking fixes:
- _published_value reads each expression as one word, so an expression in a
  URL's userinfo is withheld with it rather than garbling the URL and
  publishing its path.
- A codex --config value that starts like a table and does not parse, as
  string-argv leaves a quoted one, is withheld as unparsed_json.

Rebase (F5). CHANGELOG keeps #853's #778 line beside #823's under
github_action row. llms.txt and ai-search-summary state the source tree as
contract 41, unreleased, ahead of the published v1.1.0 (contract 40), as
test_public_surface_contract requires while the two differ. The pilot
ledger's source-tree column was re-taken on the rebased tree, through
./shipgate beside the engine of e3c6cb0c, the commit v1.1.0 was cut from:
- the only differences are contract 40 -> 41 and inventory schema 0.6 -> 0.7;
- diff rows are byte-identical;
- check JSON differs only in the launcher path its next action names.

The support page, the STABILITY preamble and migration note, the CHANGELOG
entry, agent-contract-current, the Stop hook sentence in integrations, the
capability_diff row in distribution-surfaces with its parity comment, the
0.7 schema files and llms-full.txt are updated to match.

* Address review cycle 2 on agent launches in CI (#823)

The third review of #850 at c7f67550 found one P1 and one P2 defect and
four P3 notes. The branch is on origin/main 44b9e05d; no rebase was needed.

A JSON setting publishes its shape, not its free text (C2-F1).
_withheld_json published the whole _redact_secret_values tree. That tree is
the host readers' digest input, not what they publish: it keeps every
string outside env, headers and secret-named keys. So an mcp-remote
--header "Authorization: Bearer ..." argument in --mcp-config, and a hook's
curl command in settings, reached diff text, diff --json, audit --host
--json, the PR comment and verifier.json. _json_shape now keeps key names,
numbers, booleans and null, redacts what the host readers redact, and
replaces each other string with <withheld:...>, a 12-hex digest of
redacted_config_sha256 for that string, so an edit to it is still a
changed row. It keeps only the strings a host reader publishes:
- a permissions.allow/ask/deny rule and a documented Claude Code setting's
  value (defaultMode, the switches, enabledMcpjsonServers entries);
- an MCP server's command name and its URL's scheme and host, followed by
  the digest when they drop a command's arguments or a URL's query.
A codex --config table or array is read under its key path, so
mcp_servers.gh={command="gh", ...} keeps its command name. The canary
sweep adds the mcp-remote header and the hook command in every spelling
(action input, claude_args, CLI flag, codex -c table). Re-running the
reviewer's two repositories through ./shipgate gives 0 canaries in every
output.

The documented bypasses written through listed inputs widen (C2-F2).
- Claude Code settings written as JSON meet bypass_permissions when their
  defaultMode is bypassPermissions, read by claude_setting_values as the
  settings reader reads .claude/settings.json. This covers the action's
  settings input, a --settings value in claude_args, and the CLI's
  --settings flag. A path is not read, and a settings value holding an
  expression meets none.
- openai/codex-action's permission-profile: :danger-full-access meets
  danger_full_access. ":danger-full-access" is Codex's reserved name for
  its built-in full-access profile
  (BUILT_IN_PERMISSION_PROFILE_DANGER_FULL_ACCESS).
Mode inputs are now (input, value, rule) triples. settings and
permission-profile join _RULE_SETTINGS, so a gain after an expression in
either is named and not claimed, as for claude_args.

P3 notes:
- allowed_bots: "*" now reads "accepts runs triggered by any bot" through
  agent_rule_text.
- Which of two gaining jobs a moved rule goes to no longer depends on
  declaration order: the same launch arriving is matched before a job that
  merely stopped launching the agent.

The support page, the STABILITY migration note, the contract summary, the
schema docstrings and the CHANGELOG entry now say what a structured value
publishes instead of claiming it publishes what the host readers would, and
list the two rules. host-grants 0.7 is extended in place; it is unreleased.

* Address review cycle 3 on agent launches in CI (#823)

Rebased onto origin/main 01777037 (#852, #821), which had already moved the
unreleased runtime contract 40 -> 41, verifier 0.20 -> 0.21 and capability
diff 0.3 -> 0.4. This change now extends contract 41 in place instead of
minting it: one comment in schemas/contract.py names #821 and #823, and the
texts that said verifier 0.20 and capability diff 0.3 were unchanged now say
capability_diff registry row keeps #852's two added roots and both paragraphs;
STABILITY keeps the #821, #823 and #827 notes, with #821's "host-grants stays
0.6" and #827's version sentence corrected for a tree that also carries
host-grants 0.7; CHANGELOG keeps both Unreleased entries. llms-full.txt is
rebuilt. The pilot ledger's source-tree column was re-measured on the rebased
tree beside e3c6cb0c: identical cells except contract 40/41 and inventory
schema 0.6/0.7, identical diff text, rows, check JSON and inventory apart from
its schema version, and diff --json / verifier.json differing only in #821's
schema versions and coverage members.

A codex exec step that selects the full-access sandbox through --config was a
changed row while --sandbox danger-full-access widened, and -sdanger-full-access
was not read at all. The flag reader now reads a short flag's attached value
as clap does (-s<mode>, -s=<mode>, -c<override>, -c=<override>), and a
--config override that sets sandbox_mode to danger-full-access, or
default_permissions to :danger-full-access (what codex-action's
permission-profile input passes the CLI), meets danger_full_access, its value
read as the CLI's parse_overrides reads it. As the CLI resolves them, the last
override of a key counts, default_permissions outranks sandbox_mode, and a
--sandbox flag outranks both, so an override beside --sandbox meets none. In
codex-args such an override meets none, because the action appends its own
--sandbox or default_permissions selection after codex-args. The support page
and the STABILITY Direction bullet list the spellings.

Withholding a quoted URL inside a JSON-array codex-args element no longer drops
the element's escaped closing quote, so the published value stays the JSON
array the support page describes.

* Address review cycle 2 on agent launches in CI (#823)

An agent CLI inside a double-quoted $(...) or a backtick substitution, or
after a shell reserved word, got no launch, no row and no coverage issue:
gh pr comment --body "$(claude -p --dangerously-skip-permissions ...)" read
as "No static host-grant changes detected", while the support page said such
a run: is listed as unresolved once for each agent CLI it starts at the head
of a command. The word splitter reads a double-quoted substitution as one
word and a backtick as no operator, and the head reader took then, do, { and
! for the command.

_command_substitutions now lists the text of each $(...) and backtick
substitution outside single quotes, double-quoted ones included, and an agent
CLI heading a command inside one is an unresolved launch (shell_expansion for
a single top-level command, compound_command otherwise), publishing none of
its text. An escaped \$, a single-quoted '$(...)', a comment and $((...))
arithmetic are not substitutions, so echo "claude -p ..." stays unlisted. The
scanner keeps an explicit stack and reads each nested substitution as `_` in
the one around it, so no nesting depth recurses or re-splits text. When the
run's quoting does not balance, each line's substitutions are read as well.
A command's head is now read after shell reserved words (!, time [-p], {, if,
then, elif, else, while, until, do, function NAME), and a command after one
is not a simple command, so its launch is compound_command and never read.
The compound_command and shell_expansion limit phrases, the launch schema
description (host-grants 0.7 schemas regenerated), the support page and the
STABILITY bullet say so.

A read launch that became a form this audit does not recognise (npx, a path,
codex options before exec) was worded "a step no longer launches an agent",
though the step still starts one. The removed case now reads "a step no longer
declares an agent launch this audit reads", and a row whose workflow still
exists adds that the step may still start an agent in a way this audit does
not read, naming those forms. codex with an option before exec is listed under
Known unread surfaces and in STABILITY's "What is not read".

Tests cover each shape the review named, the negative controls, the diff and
audit --host route of the reproduction, the reworded removal for npx, a path
and codex root options, and 20000 nested substitutions.

* Address review cycle 3 on agent launches in CI (#823)

A step that drafts a PR comment in a quoted here-doc, such as
cat > comment.md <<'EOF' with "Reproduce locally with `claude -p ...`" in its
body, was listed as an unresolved agent launch: a "high changed" row saying a
step now launches an agent, and a GitHub coverage issue. The shell passes a
here-doc's body to its command as input and, with a quoted delimiter, expands
nothing in it, so the step starts no agent. The phantom also hid a real gain:
a --dangerously-skip-permissions step added beside it was "not counted as a
widening" because the job "launched the agent in a form this audit does not
read" before. The substitution scanner read backticks and $( across the whole
run:, and the word splitter read a body line starting with claude -p as a
command head.

_here_documents now removes each here-doc's body and closing line before the
run: is split or scanned. <<, <<- and a quoted, escaped or partly quoted
delimiter are read outside quotes, comments and $((...)) arithmetic, inside a
$(...) or backtick substitution too; <<< is a here-string; bodies start after
the opening line and follow in order when one line opens several. A body is
never a command. Only an unquoted <<EOF body's substitutions, which the shell
runs, are read, with quotes and # in it literal, so '$(claude -p ...)' there
is listed and was missed before. A here-doc with no closing line is left in
the text and read as ordinary lines, so a misread << hides no later command.
Closing lines are looked up by content, so many here-docs cost one pass. Every
shape was checked against bash with stub claude and codex functions.

A launch that only became a form this audit does not read no longer counts
as having left its job. The moved-between-jobs rule now takes a rule as moved
only when the job that met it no longer exists, or when the same launch now
runs in the gaining job and that job's own launches of the agent still run
there or in the job it left (a moved step, a swap). Job a going from
claude -p --dangerously-skip-permissions to npx @anthropic-ai/claude-code -p
while job b's launch is edited into a bypassing one is now widened, where it
was "moved between jobs". A step moved and edited while its job remains is
claimed too.

A checkout step on one side only, such as a default checkout in an added job,
is worded "a step now declares a checkout" or "a step no longer declares a
checkout" rather than "a checkout's declared ref changed".

The support page, STABILITY's migration note and the CHANGELOG entry say
which here-doc text is read, name the two shapes still read as lines, and
state the narrowed move rule.

* Name an unchanged instruction limit reached through an in-tree link instead of refusing (#822)

An instruction file whose limit this entry cannot resolve, such as a
SKILL.md whose metadata holds a non-string value, is named as an unchanged
limit when a change leaves it alone (#721). Reached through an in-tree link
the reader reads through (#700) - `.claude/skills -> ../.agents/skills`, a
per-skill link or a file link - the same untouched file refused the whole
comparison: diff, verify and the manifest-free PR comment printed
`base_inventory_incomplete; head_inventory_incomplete` and no row, hiding a
removed deny rule beside it. `unchanged_limits` asked blob_path_unchanged
whether the source, the path the link is read under, was one regular file
in Git, and a path through a link never is.

blob_path_unchanged now resolves the path on each side the way the reader
reaches it: from Git tree entries for the base and a commit head, and for a
working-tree head without following any link (each component's own entry,
each link's own text, the file's unfiltered hash). It holds only when both
resolutions are equal - every link at the same path with the same text,
every other component a directory - and the file they land on is the same
regular-file blob at the same in-tree path. A link is followed only under
the rules the reader and the base archive already use (#700, #711): a
relative text that lands inside the tree after normalization, directories
above where it lands, at most eight links, and links at one component of
the path. A path no link reaches still takes the one `ls-tree` it took
before, and blob object IDs are still compared, so no filter or textconv
can make two byte sequences equal.

Everything that asks the proof moves together: unchanged_limits in
`diff --json` and verifier.json, the text and PR comment, the unchanged
limits a partial comparison may carry (#808) and the shared
plugin-reference limits check leaves out (#714). check's boundary result
still cannot name a limit, so it now refuses these comparisons with
unchanged_limits_not_representable, as it does for a limit at its own path;
its rows, decision and violations do not move. No schema, member, reason
code or check id is added. The metadata value is not coerced.

tests/test_linked_unchanged_limits.py holds every layout (directory link,
per-skill link, file link, a chain of file links) to the direct result on
diff, verify, the PR comment and check; keeps the negative controls refused
(skill added or edited behind the link, link retargeted to an identical
copy, link text rewritten to land on the same file, link replaced by a
directory or the reverse, a later hop retargeted, working-tree-only
retargets and edits); and holds the proof to exactly the links the reader
reads through, across dangling, looping, absolute, escaping, over-long,
intermediate and nested links. The #812 coverage case that pinned the
refusal is retired, and the static-only allowlist follows its two call
sites down the file.

* Address review cycle 1 on linked unchanged limits (#822)

The STABILITY migration note said that where the other routes now compare
past a limit reached through an in-tree link, `check` refuses with
`unchanged_limits_not_representable` and its rows stay empty. That holds for
an instruction file's limit, which `check` cannot leave out, but not for a
plugin-reference limit: `_without_shared_plugin_reference_limits` asks the
same unchanged proof, so a `parse_failed` `plugin.json` that is a file link
to an unchanged target is now left out exactly as one at its own path has
been since #714. `check` then compares and publishes the removed `deny` row
in its boundary result and its control envelope's `capability_rows`, where
it refused with `base_inventory_incomplete` / `head_inventory_incomplete`
and no row, and `diff`, `verify` and the PR comment, which withheld
`plugins/demo` as `partial` (#808), are `comparable` with the limit in
`unchanged_limits`. `check`'s decision, violations and control state, and
`verify`'s control state and next action, do not move.

The note now splits `check` by whether it may leave the limit out, adds the
`partial` to `comparable` move, and says "every row's value" where it said
"every row". The CHANGELOG entry mirrors it, the distribution-surfaces row
and `docs/host-boundary-support.md` no longer say any change behind the link
refuses the comparison (an edited plugin manifest keeps it `partial`), and
`tests/test_linked_unchanged_limits.py` pins the plugin manifest at its own
path and behind a file link, on every route, with the edited-target control.

The `blob_path_unchanged` and `_reader_path` docstrings now state the proof
as necessary, not sufficient: it follows the links on the way to the path,
not every condition the reader puts on reading a whole linked directory, and
a side whose reader does not read the path carries no limit there. A test
holds a head that adds a link inside the linked directory to a refusal on
every route. The two pinned call-site lines in `cli/verify/git.py` move with
the docstrings.

* Address review cycle 4 on agent launches in CI (#823)

Four review cycles each found a new shell form (quoted words, $(...),
here-docs, comments, reserved words) that the run: reader mis-read. The
cycle-4 finding was the fourth: a `#` comment was split as commands, so an
apostrophe in a comment hid a launch and a command line in a comment
invented one. Parsing arbitrary shell cannot converge, so the reader now
claims only forms it reads exactly and names every other one as a limit.

- A run: step is an agent launch only when it is one line of plain words
  (letters, digits and `_ . / : = , % + -`, separated by spaces or tabs),
  run by bash, sh or no declared shell:, whose program's file name is
  `claude` with -p/--print or `codex` followed by exec. Every POSIX shell
  runs such text as exactly those words; a cross-check of 2251 generated
  texts in sh, bash, dash and ksh agrees on every one.
- Any other run: that mentions claude or codex as a word of its own is an
  `unread_agent_runs[]` entry (job, step, agent) on the workflow grant: a
  non-blocking coverage issue in audit --host that publishes none of its
  text, is never compared, so it gives no row, and never says whether the
  step starts an agent. It takes no gain from another launch; only when an
  unread step goes and a read launch is added in the same job is a rule
  that launch meets named and not claimed, since it may be that step
  rewritten.
- claude_args and codex-args are read only as a plain list of words (the
  same characters and parentheses, across blanks and newlines, with no
  --settings or --mcp-config flag). Any other value, a ${{ }} expression
  included, is `unread_arguments`: published only as a digest, so an edit
  is a changed row, and read for no rule; a rule a launch gains where that
  input was unread before is named and not claimed.
- A codex --config override publishes its key; its value is <redacted>
  under env, headers or a secret-named key, as written for sandbox_mode,
  default_permissions, approval_policy and model, and a digest otherwise.
  The word after a secret-named word such as --token is <redacted> and the
  value is then published redacted, a named limit.
- The shell tokenizer, the command-substitution and here-doc scanners, the
  reserved-word reader and the shell-quote and string-argv emulations are
  removed, with the expression-prefix reading of argument inputs. JSON is
  read only in the settings and mcp_config inputs.

Host-grants 0.7 is unreleased, so its schema changes in place: the launch
`unresolved_reason` is only inputs_not_a_mapping, a setting may be
unread_arguments, and the workflow grant adds unread_agent_runs. The
support page, STABILITY, CHANGELOG, the current contract page and
llms-full.txt describe the tightened reader.

* Address review cycle 5 on agent launches in CI (#823)

C5-F1: the move rule took a launch that only stopped being read to have
left its job. With job a running `claude -p --dangerously-skip-permissions
Review` in the base, quoting its prompt (an unread step) or running it
through `npx` while job b added the same plain launch gave a `changed`
row saying the launch "moved between jobs (a → b)" and "already met that
rule in the job it left", with no widening signal. Job a still runs it.

- `same_launch` now refuses a move while the losing job may still run
  the launch in a form this reader does not read: an unread step of that
  agent stands at a step label the lost launch held, the job has more
  unread steps of that agent than before (the launch may have moved to
  another index), or one of its read launches of that agent holds an
  expression or an unread argument input in an input the rule is read
  from. The review's M5 and M6 are now `widened` and name "a step no
  longer declares an agent launch this audit reads (a/steps[0])"; a real
  move, a rename, a swap, and a move beside an unread step the job
  already had stay moves.

C5-F2: docs/integrations.md said an unread `run:` is a non-widening row
that diff and the PR comment show. It gives no row; only an unread
argument input is a row. The sentence is split accordingly.

Nonblocking items fixed:
- `_job_secrets` walks each container once, without recursion. A job
  `env` holding itself through a YAML alias raised RecursionError, and a
  ten-way fan-out eight levels deep did not finish; both now read in
  well under a second.
- A `settings` or `mcp_config` value that is neither a JSON object nor a
  plain file path (path characters, and a `${{ }}` only as a plain
  context reference) publishes only a `<withheld:…>` digest. A comment
  line before JSON published the `env` value that JSON held.
- The last of a repeated `--permission-mode` counts, in `claude_args` and
  in a plain `run:`, as parse-sdk-options and the CLI keep it.
- The docs/distribution-surfaces.md capability_diff row no longer says an
  unread argument input or an unresolved launch is named only by the
  inventory and audit --host: a row reporting its launch names it too.

The support page, STABILITY (Direction and What is withheld) and the
CHANGELOG entry say the same.

Tests: tests/test_workflow_agent_launches.py grows to 290 cases. The five
new move-rule widening cases, the alias test, the two json-or-path cases
and the four --permission-mode cases fail on the previous head; the
moved-beside-an-unread-step case and an end-to-end diff/audit test of M5
and M6 are added beside them.

* Address review cycle 6 on agent launches in CI (#823)

C6-F1: the cycle 5 move guard only refused a move when an unread step of
that agent stood at a step label the lost launch held, or the job had
more unread steps than before. Merging job a's `npm i -g
@anthropic-ai/claude-code` step into `claude -p
--dangerously-skip-permissions Review` while job b added that plain
step (V1), or removing an unread `echo "claude"` step while the launch
became a quoted, unread step at another index (V2), kept the unread
count and missed every held label, so the row said the bypass "moved
between jobs (a/steps[1] -> b/steps[0])" and gave no widening signal.
Job a may still run the launch. An unread step carries no text that
tells which launch it is, so `may_still_meet` now refuses the move
whenever the losing job has any unread step of that agent, wherever it
stands and whether or not it was there before (option (a) of the
review). V1 and V2 are now `widened`, with the "may still start an
agent" caveat, and name "a step no longer declares an agent launch this
audit reads". The cycle 5 guard that kept a move beside an unread step
the job already had (`claude mcp add x`) now widens, in the safe
direction; it moves to the widening cases, and a move beside a step that
names no agent stays a move.

C6-F2: `_codex_override` caught only `TOMLDecodeError` and
`RecursionError`, and `tomllib` raises a plain `ValueError` for an
integer past Python's 4300-digit limit, so `run: codex exec -c
sandbox_mode=<5000 digits> Review` crashed `diff`, `audit --host` and
`check` (exit 1) and `verify` (exit 4, internal_error). It now catches
`ValueError`, which `TOMLDecodeError` subclasses; such a value is read
as text and selects no sandbox, so the row is `changed`. `-c
default_permissions=<4400 digits>` and `codex e
--config=sandbox_mode=<4301 digits>` are covered too.

Wording: the support page, STABILITY (Direction and What is not read)
and the CHANGELOG say a launch has not left a job that keeps any named
unread step of that agent or a launch of it with an unread input the
rule is read from. The support page and STABILITY also name the one
case where a launch that stops being read is worded as moved rather
than as no longer declaring a launch: another job adds the same launch
while the job it left keeps none of those, as when the launch became a
script in the same change.

Nonblocking items fixed:
- docs/integrations.md: "One that gains a rule" read as the unread
  `run:` of the sentence before it; it now says "An agent launch that
  gains a rule".
- The HostWorkflowAgentSettingV7 docstring, published as the 0.7
  schema's description, states that a `settings` or `mcp_config` value
  that neither starts like a JSON object nor is a plain file path
  publishes only a `<withheld:...>` digest. The 0.7 inventory and
  baseline schema files are regenerated.

* Address review cycle 7 on agent launches in CI (#823)

C7-F1: `may_still_meet` looked for unread steps only under the losing
job's name, which holds nothing once that job is renamed or removed, and
`job_left` accepted any job that no longer exists. So renaming job a to
a2 while quoting its `claude -p --dangerously-skip-permissions Review`
(R6), or running it throu…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant