fix: compare application wiring past unrelated links and submodules; read google.adk.Agent - #877
Conversation
`diff --application` at the root scope took the unscoped archive route, which refuses every symlink and gitlink. On 2026-09-25, 15 of the first 27 runs against open third-party SDK/ADK PRs exited 2 on a path the reader never opens: `CLAUDE.md -> AGENTS.md`, a linked skill directory, a `VERSION` link leaving the tree, a vendored submodule. - Materialize every scope, the root included, through the scoped verified materializer, which recreates links rather than refusing them and packs the tree instead of the history. - Never read a Python input through a link: a target this scope already reads is compared at its own path, and any other in-scope target is a gap over the link's path. Reading the alias made one agent two ambiguous ones and hid the real file's change. - Record gitlinks behind an opt-in `archive_tree(record_gitlinks=True)`, materialized as the empty directory an unpopulated checkout leaves. An unchanged gitlink commit is named in limits; any other is a coverage gap over its path. Other archive callers still refuse. - Stop turning the host-configuration census's link count into application coverage gaps. - Read `google.adk.Agent`, the package-root re-export, as an ADK agent constructor (`from google.adk import Agent`). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pengfei-threemoonslab
left a comment
There was a problem hiding this comment.
Reviewed all nine changed files at 9f325c3b23afd1d6c5536c39c159377cde327ab2, including the shared materializer's other callers, discovery filtering, per-binding uncertainty propagation, and the ADK constructor change. I found two reproducible correctness issues in coverage handling; I recommend fixing both before merge. Details and reproduction results are attached inline.
Validation against an isolated archive of this exact commit:
- 758 tests passed:
test_application_diff{,_review,_reach}.py,test_scoped_base_tree.py,test_capability_diff_partial_clone.py,test_adapter_static_only.py,test_google_adk.py,test_verification_git_snapshot.py,test_host_file_links.py, andtest_host_link_read_through.py. - Ruff passed for all changed Python files;
git diff --checkpassed. - Additional small committed-repository reproductions confirmed both findings. The dangling-link case was also run against base
a430e81a, where it correctly returnspartialandnot_established.
The existing regression tests pass, but neither case below is covered. The full repository test suite was not run locally. This is a code review, not a release-verifier or merge-authorization result.
| # Python input this scope already reads, it is compared at its own path, so | ||
| # the link adds nothing; reading it again made one agent two ambiguous ones | ||
| # and hid that file's changes. Any other target is named as a gap below. | ||
| linked = {p.relative_to(root).as_posix(): p for p in python_files if p.is_symlink()} |
There was a problem hiding this comment.
[P1] Preserve gaps for dangling Python symlinks
linked is built from _candidate_files(root), but its filesystem walker only returns paths for which path.is_file() is true. A dangling in-scope link such as app/agent.py -> missing.py is therefore absent before this code runs and never reaches the new Linked Python input gap. Removing host_discovery_incomplete_paths also removes the previous fallback. Reproduced with a base containing app/agent.py that binds lookup, then replacing that file with the link above: diff --application --scope app --base <base> --head <head> --json now returns comparison_status: "compared", change: "removed", uncertainty: {}, and no head coverage gaps. Base a430e81a returns partial / not_established for the same case. This turns an unread application input into conclusive negative evidence, contrary to the documented gap behavior for other in-scope link targets. Census Python symlink entries without following/filtering their targets, and retain a source-attributed gap for unresolved targets. Please add a regression covering this file-to-dangling-link transition.
There was a problem hiding this comment.
Confirmed and fixed in 3a0c00c. On 9f325c3b, your repro gives compared with a definite removed row. Main a430e81a gives partial / not_established (Discovery could not read agent.py.).
The same regression covered one more case: a source directory replaced by a dangling or absolute link (app/tools -> /opt/tools). Main caught it through the same host-census gap that this PR removed.
Fix. observe() now censuses every link under the scope itself, with os.walk and without following links, filtered by discovery's _skip_part. It no longer depends on discovery's inventory, which drops links that don't resolve.
- A
*.pylink whose target isn't a Python input the scope already reads is always a gap over its own path. That covers dangling targets, targets leaving the scope or tree, and non-Python targets. An alias of an input the scope reads is still compared at the target's own path. - A link to a directory outside the scope that holds Python is a gap.
- A link that resolves to nothing in the tree is a gap where the other side reads source at or beneath its path. A later step,
_reconcile_unresolved_links, checks this. An unchangedagent/VERSION -> ../../VERSIONbeside the application still changes nothing, and a test pins that.
Regressions in tests/test_application_diff_reach.py:
test_python_file_replaced_by_a_dangling_link_is_not_a_removal, in both directions. It assertsnot_establishedwithuncertainty == {"head": ["Linked Python input: agent.py"]}.test_source_directory_replaced_by_an_unresolved_link_is_not_a_removal: absolute, leaving-tree and dangling targets, in both directions.test_directory_link_leaving_the_scope_with_python_is_a_gap.test_python_link_to_anything_else_is_a_gap_over_its_path: I had pinned the leaving-scope case as no gap, which was wrong; it is now a gap.
Run against an archive of 9f325c3b with the new file copied in, 13 of these cases fail. All 31 in the file pass at 3a0c00c7.
| if scope == ".": | ||
| result.submodules[path] = commit | ||
| elif PurePosixPath(path).is_relative_to(scope): | ||
| result.submodules[PurePosixPath(path).relative_to(scope).as_posix()] = commit |
There was a problem hiding this comment.
[P2] Make a submodule at the selected scope cover all relative bindings
When the gitlink path equals scope, relative_to(scope).as_posix() is ".". _reconcile_submodules() then emits a gap with source: ".", but Observations.absence_gaps() only matches exact paths or path + "/" prefixes; a relative binding such as agent.py matches neither "." nor "./". Reproduced with a base whose app is a gitlink and a head that vendors ordinary app/agent.py source binding lookup, using --scope app: the overall result is partial, yet the row says change: "added" with uncertainty: {} despite the unread base submodule. The reverse transition similarly risks a definite removal. Represent this as a scope-wide gap (source=None) or explicitly treat "." as covering every relative path, and test both submodule-to-directory and directory-to-submodule transitions at the selected scope.
There was a problem hiding this comment.
Confirmed and fixed in 3a0c00c. On 9f325c3b, a base whose app is a gitlink compared with --scope app against a head holding real app/agent.py gave partial overall, but the row was added with uncertainty: {}. Main refuses the case outright (exit 2).
Fix. In _reconcile_submodules, a gitlink at the selected scope itself (relative path ".") now gives a scope-wide gap with source=None, so absence_gaps() covers every relative binding path. The message reads Submodule content is not read (commit …): the selected scope rather than : ..
Regression. test_submodule_at_the_selected_scope_covers_the_whole_scope covers both the submodule→directory and directory→submodule transitions. Each asserts not_established, the uncertainty on the side holding the gitlink, and a gap whose source is None. Both fail on 9f325c3b and pass now.
…ovals PR #877 review found two cases where unread input became conclusive negative evidence: - Discovery drops a path that does not resolve, so replacing `agent.py` with a dangling link read as a definite removal, `compared`, with no gap. The dropped host-census gap had been the only fallback, and it also covered a source directory replaced by an absolute or dangling link. Census every link under the scope directly, without following it: gap a `*.py` link that does not alias an input the scope already reads, a directory link holding Python outside the scope, and a link resolving to nothing in the tree where the other side reads source at or beneath it. An unchanged `agent/VERSION -> ../../VERSION` still changes nothing. - A gitlink at the selected scope itself gave a gap with source ".", which covers no relative binding path, so `agent.py` read as a definite addition or removal. Make that gap scope-wide. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-a855cb # Conflicts: # CHANGELOG.md # docs/distribution-surfaces.md
_materialize_isolated_tree spawned one `git cat-file blob <oid>` per file in both its entries and links loops. Once `diff --application --scope .` materializes the whole tree (#877), that dominated the run: one side of TencentCloud/CubeSandbox took ~180 s to archive (the #686 cost class). Blobs are now read by _isolated_blobs: one `cat-file --batch-check` types and sizes every object (a missing or non-blob object refuses with a ConfigError before any content is read), then `cat-file --batch` runs of at most 64 MiB each return the content, with strict framing checks. Only full object IDs reach the batch. The link-text reads in _scope_through_boundary_links use the same reader. Every blob is still hashed against its tree entry's oid before it is written; path handling, containment and escape checks, blobs-before-links order, placeholders and the final digest comparison are unchanged. The reads go through _run_git_dir (now accepting input=) and the existing _run_process boundary, so no subprocess call site is added; only the two line pins in test_adapter_static_only.py moved. diff --application --scope . with #877: CubeSandbox 412 s -> 46 s, dlt 203 s -> 45 s, identical rows.
… in the agent's function (#865) (#903) * Read a Google ADK tool a local factory builds, and a tools list built in the agent's function (#865) visulate/visulate-for-oracle#526 bound `save_memory_tool = create_save_memory_tool()` in its root agent's builder, where the factory returns `FunctionTool` around a nested function; the reader stopped at the local assignment. A factory call bound once and unconditionally (in the agent's function or at a module's top level), or inline in `tools=[...]`, is now followed as syntax to the factory's one unconditional `return` of `FunctionTool(inner)`, a plain function, a name bound once to one of those, or another factory's call, up to four deep; the tool is the wrapped function, with the factory recorded in `import_path`. A factory that returns from several places or under a condition, recurses, is decorated or a generator; a wrapped function that is decorated, a parameter, or changed or handed on (Visulate's delegate sets `__name__`); a tool changed where it is bound; and a factory held outside the read scope stay named with why (`factory_return`, or the resolver's reason). A third-party factory keeps the answer it had. On visulate#526 at `ai-agent`, the two memory tools become not_established candidate additions beside the nine named delegates. MuhammadVT/smart-assignment#46 built `tools = [...]` in the agent's function with a conditional `tools.append(...)`: read as a dynamic tools expression, it lost the unconditional tools. Such a list with append/extend/insert/`+=` is read member by member; an addition under a condition or in a loop is named on the agent and never read as bound; any other use keeps it dynamic. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Resolve repository-local imported tools for Google ADK and OpenAI Agents SDK (#864) (#879) * feat: resolve repository-local imported tools for ADK and SDK readers (#864) A tool an agent binds from another module was unresolved at the module boundary, so a PR adding one produced no row. A shared resolver (inputs/python_imports.py) follows a tools-list reference through static imports, package re-exports, module-qualified access and plain aliases to one function definition, reading only regular .py files inside the read directory through the bounded input reader, never importing or running them. Every stop is a named reason. - Google ADK: imported names, module.function, alias = function, and FunctionTool/LongRunningFunctionTool wrappers (inline, assigned, or built in the imported module). One tool per definition; same-named functions in different modules stay distinct via binding locators. - OpenAI Agents SDK: names and module.function reaching the SDK's @function_tool, with guard evidence for imported definitions. - diff --application: rows carry import_path evidence (modules, lines, digests) outside compared meaning; unresolved references are scoped to their agent with the reason. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): import resolution never produces a false or complete answer (review round 1) Adversarial review of the shared resolver found answers that read as complete while wrong: - One SDK agent binding two same-named definitions from different modules kept the last and reported a false CHANGED row. Both readers now bind neither, whatever the list order, and name both definitions. - A name the function building the agent binds itself (a local import, a parameter) was resolved through the module's binding. It is now a named stop (`local_binding`); a nested function that is the only definition of its name is that definition. - `import a.b` then `a.b.f` read `a/__init__`'s own `b`. It now reads the submodule, and several `import a.x` statements are not a rebinding. - An ADK wrapper warning shared by two agents scoped its gap to one of them. Every unresolved-reference record now gets its own gap. - `scan` counted one definition twice when an import reached a module another configured source also reads, making `{tool: ...}` selectors ambiguous. The catalog keeps one observation, and the binding graph reaches it through the exact definition locator the reader resolved when the edge's own source has none. - The module-binding walk climbed a parent chain per node; it is now linear (a 744 KB nested module: 32 s -> 3.7 s). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): read a reference where it is used (review round 2) - A builder's own `from support import lookup` is followed like a module-level import (`ImportResolver.resolve_local_import`) instead of stopping, for bare names and ADK `FunctionTool(func=...)` arguments alike. - `nonlocal` follows the outer function's binding; a name a scope binds more than once is a named stop. - `tools.lookup = ...` / `setattr(tools, "lookup", ...)` in the module that binds `tools.lookup` is a named stop. - A factory's own `toolset = McpToolset(...)` / `tool = FunctionTool(...)` is read through the existing toolset and wrapper paths again, so the MCP endpoint and toolset checks return. - A module-level ADK agent binds the module-level `def`, not a nested one the flat function map happened to keep last (pre-existing on main). - An SDK list variable bound twice anywhere in the file, or changed in place, is dynamic. - The `scan` dedupe runs once where sources are loaded, so guard association and inventory completion see the same tools; a source an inventory completes keeps its observation; `./` spellings match. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): read a list and a flat name where they are bound (review round 3) - SDK tool lists are read through the scope that binds the name at the construction. They are read only when that scope binds the name once, to a literal list, and no change site anywhere in the file (a method call, a subscript store, `global` / `nonlocal`) changes that binding. Each change site is indexed once, so the reader is linear. A module-level list's names are resolved where the list is written, not in the building function's scope (R3-1, R3-4). - A flat function, wrapper or toolset map answers a module-level reference only when its entry is the module's single top-level binding. Otherwise the resolver answers (a module import wins over a nested `def`). If the resolver cannot establish the binding, the same-named definition is named, with a tool-scoped issue, and its row is never established (R3-5). - A wrapper's `func` is read where the wrapper is written. - An enclosing package `__init__.py` that reassigns the definition is a named stop (R3-3). `_dotted` is iterative, so a chain thousands of attributes deep no longer crashes (R3-2). - The scan dedupe drops the removed copy's guard evidence too (R3-7). - The CHANGELOG states the inventory-completion ambiguity (R3-6). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): a guess is scoped to its tool, and a list's second handle counts (review round 4) - A guessed ADK binding no longer marks the agent's list incomplete. - That had dropped every per-tool `scan` finding of the agent for one name (R4-2). `scan` is back to the round-3 behavior: a shadowed, medium-confidence definition. The PR #400 tests pass unmodified again. - The comparison reads `tool_issues` as a tool-scoped gap and an agent-scoped gap. The row is `not_established` on whichever side it is present, added and removed included (R4-1). - `x = FunctionTool(func=x)` right after `def x` wraps that `def` and is not a guess. That was noise on byte-identical files. - The SDK list reader counts `alias = TOOLS` and `TOOLS` passed to any call but a read-only builtin or logging method as a change (R4-3). - An attribute of the resolved name reassigned in an enclosing package's `__init__.py`, or in a module it imports relatively, is a named stop. This covers an alias spelled from a package above (R4-4). - Same-file symbol lookup is a dict, not a scan of every tool per reference. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): scope a guess to its candidates; read list arguments and patches precisely (review round 5) - A guessed ADK binding's reason covers the tool it bound. It also covers every tool the module's bindings of that name could give the agent: a `def`, an import, `x = other`, or `x = FunctionTool(func=f)`. It covers the whole agent (`ANY_TOOL`) only when one of those cannot be named. - A guess that binds a name the agent already lists still leaves a gap (R5-1). - A `try: import … except ImportError: def …` fallback no longer hides the agent's other changes (R5-2). - A literal list passed to a call counts as changed unless the call is: - a read-only builtin, `pprint`, or a logging method; - an SDK `Agent` or a copy (`clone`, `replace`) reading its own `tools=`; - a function the module can resolve that leaves that parameter alone. Only names bound to a literal list are followed into a callee, so no other module is read for an unrelated call (R5-3). - The package-patch check counts only attributes rooted at an import. It parses the modules an `__init__` imports on a separate bounded budget, so a package that re-exports many modules no longer exhausts the resolution budget (R5-4, R5-5). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): only reads keep a list literal; follow a guess's candidates (review round 6) - A literal tool list stays readable only while every use of it is a read. Allowed reads: iteration, indexing, comparison, a truth test, formatting, a read-only builtin or logging method, an agent's or copy's own `tools=`, or a function whose every use of that parameter is such a read. `+=`, a return, a tuple, `*args`, `**kwargs` or storing it in another container makes it dynamic (R6-1). - A guess's candidate tool names are followed to their definitions: an import through the resolver, or by its imported name when the module is outside the scope; `x = other` and `FunctionTool(func=f)` through the name they spell. A candidate that cannot be followed, or a wildcard import, scopes the reason to the whole agent (R6-2, R6-3). - A module that a package `__init__` imports relatively and that cannot be read (a link, a missing file) is a named stop, cached with the package (R6-4). A module read for patches is the same object a later resolution uses (R6-5). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): spreads and tests are reads; a package's routine imports never stop (review round 7) - A literal list is still read when it is spread (`[*COMMON, x]`, `print(*TOOLS)`), tested (`TOOLS or …` in a condition), or read through a dict method (`.get`, `.keys`, `.values`, `.items`). `x or y` whose value is the list is judged by its own use. A `globals()` or `vars()` call in the module makes its lists dynamic (R7-3, R7-1 b4). - The package-patch scan treats an import above the scope as the read's boundary, as every import does. It skips an optional module imported under `except ImportError`. A link, or a missing module that is not optional, still stops (R7-2). - Docs: drop a duplicated sentence. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): code that runs first is checked, and unread code keeps a binding named (review round 8) - Every module on a resolution chain (the agent's own file included), the `__init__.py` of every package enclosing one, and every in-scope module those import are checked for a reassignment of an attribute named like a step of the chain. `import patches` in the agent's file, or in a module the chain re-exports through, is now a named stop (R8-2). The defining module handing its function on (`registry.lookup = lookup`) does not count, and every location is kept, so one never hides another. - A relative import in any of those modules that climbs above the scope runs code that is not read. It no longer disappears: the tool is named and its row is `not_established` with that import in the reason, in both readers; for `scan` the ADK module stays at medium (R8-1, which round 7 had made silent). An import under `if TYPE_CHECKING:` never runs and is skipped. - `sys.modules`, however spelled, and importing a module by `__name__` reach its lists like `globals()` does: an SDK list there is dynamic (R8-3). - A redirecting package `__getattr__` stays a documented residual (R8-4): a gap for every hook on a chain would make #864's own attest rows `not_established`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): every spelling of unread code keeps the caveat; a lazy package hook is proven (review round 9) - A caveat survives a `FunctionTool(func=f)` wrapper the imported module builds (R9-1), and an absolute import spelled through a directory above the scope (`from svc.patches import ...` with scope `svc/app`) is a caveat like the relative one (R9-2). - The defining module's hand-on exemption no longer covers a patch through its own import (`import tools as _me; _me.lookup = ...`, R9-3), and only `typing`'s `TYPE_CHECKING` skips a block (R9-4). - Ordinary code no longer stops a resolution (R9-5): every location an ambiguous import could mean is read, a generated `*_pb2` module is the boundary and any other missing relative module a caveat, and the patch scan reads up to 1024 modules. - A package `__getattr__` is established only when every return it can reach for the name gives that submodule (`import_module(f".{name}", __name__)`, `from . import name`); attest's two lazy loaders stay established, a redirecting hook is named (R8-4). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): the repository says what is its own code; harden the lazy-hook idiom (review round 10) - Whether an absolute import no file in the scope provides is the application's own code is read from the repository: the compared commit's tree for `diff --application` (set per side), the checkout for `scan`. A module or regular package at the root or under `src/` is; a directory without `__init__.py` only when it holds the named submodule. `from common.patches import ...` with scope `svc/app` is a caveat (R10-1), and SDK apps under `agents/` import the SDK again (R10-3: the round-9 ancestor-name rule had caveated every tool there). - The scope spelled from the repository root (`svc.app.tools` with scope `svc/app`) is read inside the scope, so it resolves and a patch module it names is checked instead of guessed. - The lazy-hook idiom requires an undecorated hook that never rebinds its parameter, `importlib` / `import_module` bound only by importing them, and a returned local bound exactly once, counting `for`, `with`, walrus and `except` targets (R10-2). - A generated `_version` module is the boundary like `*_pb2` (R10-4); the patch-scan budget message names what it scans. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): every directory above the scope is an import root; unread links and subverted hooks are named (review round 11) - An absolute import is looked up at the repository root, under `src/`, and under every directory between the root and the scope, so `backend/app` importing `common` from `backend/common` is a caveat (R11-1). - A root entry that is a symbolic link or a submodule, spelled by the import, is unread code: a caveat (R11-2). `scan` outside a checkout reads the three directories above the scope instead of nothing (R11-3). - A store into `sys.modules` in code that runs first is a named stop, and a package hook is not trusted when the package rebinds `__name__`, patches `importlib`, or stores into `sys.modules` or `globals()` other than the idiom's own cache (R11-4). - The scope spelled from an import root is read inside the scope only through a regular package (R11-5). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): a package is not an import root; a sys.modules store is read by its key (review round 12) - A directory between the repository root and the scope that holds an `__init__.py` is imported through its parent, never from the path: SDK apps under a regular `app/agents/` package import the SDK again, and `app/types.py` no longer shadows the standard library (R12-1, a round-11 regression). - A store into `sys.modules` (`[...] =`, `setdefault`, `__setitem__`, `update`) is a named stop when its key names a module on the chain or is built on `__name__`, a caveat when it is computed (a plugin loader's `spec.name`), and nothing when it names another module (R12-2, R12-3). - `globals().update(...)` / `setdefault` in a package disqualifies its hook (R12-3). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(#864): packages above the scope are read; sys.modules and globals() by allow-list (review round 13) - Every directory between the repository root and the scope is an import root again, package or not: a service run from `backend/` with a stray `backend/__init__.py` names `common.patches` (R13-1, a round-12 regression). A standard-library name, and the scope's own package on the way to it unless it holds the name imported, are not the repository's code, so SDK apps under `app/agents/` still import the SDK (R13-5). - The `__init__.py` of every package above the scope, and each module it imports, are read through the repository layout (the commit's tree, or the checkout) for the same reassignments (R13-2). - `sys.modules` and `globals()` are read by allow-list. A store whose key names a module on the chain, a package above one, or the framework's own modules is a named stop; `__name__` plus a literal is that module's own name; any other use is a caveat; a module rebinding its own name through them is a reassignment; a change to `__path__` is a caveat (R13-3, R13-4, R13-6). A package that rebinds `__getattr__` or `__path__` does not have a trusted hook. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read what the packages above the scope import, and every spelling of a module's own object (#879 review, round 14) - Above the scope, follow each import the way the in-scope reader does: every package on the way to the module and each submodule named (`from .hooks import patches`, `import svc.lib.util`). A reassignment there counts only when rooted at the scope or at something that cannot be found; guarded and generated imports are exempt. - A standard-library name is exempt only at a root that is a regular package; interpreter-preloaded modules are always exempt. - Allow-list reads: comparisons, spreads, iteration, pkgutil.iter_modules(__path__), namespace keywords (get_type_hints(globalns=globals())), and patch.dict/setitem by key. - The module's own object through an alias, sys.modules.get, import_module(__name__), __dict__/vars() stores, a computed setattr, or handed to a function is a reassignment or caveat, never silent. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * A local named vars or globals is a variable, not the builtin namespace (#879 review, round 14 corpus) siada-cli's DebugUtils.dump binds `vars = stack[-2][-3]` and iterates it; the bare-name rule read it as the builtin handed on and named a real definition change not_established. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Only a two-part patch on another module is exempt above the scope; read the scope's modules an ancestor imports and the module object anywhere (#879 review, round 15) - R15-1: above the scope, a patch is exempt only when it sets one attribute of another module file outside the scope that its package binds nothing else under. A longer path, a name imported from a module, a package attribute shadowing the submodule, or a module alias (`_t = tools`) counts, in both readers. - R15-2: a scope module that an ancestor __init__.py imports is read like the chain's own. - R15-3: a module file wins over a same-named directory without __init__.py. - R15-4: a link's blob text is never read as source. - R15-5: the module object used anywhere but an attribute, a plain alias, a comparison or a reader is a caveat. So are __dict__/vars().update, getattr(m, "__dict__") stores, f_globals, builtins.globals under another name, a reader of the module's own under a reader's name, `globals = globals`, and __path__ changes above the scope (pkgutil.extend_path aside). get_type_hints(fn, globals()) is a read. - R15-6: above-scope files are read in batched `cat-file --batch` per directory (1100 modules: 76 s -> 3 s). The read bound is named once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Read a frame's namespace by allow-list; batch a directory only while it is small (#879 review, round 16) - R16-1: `sys._getframe(1).f_globals.get("__name__")` in a logging helper is a read. A frame's namespace is read by the same allow-list as globals(); any other use is a caveat. - R16-2: above-scope reads batch a directory's modules in one `cat-file --batch` only while the directory holds at most 16 MB of Python. A directory of generated or vendored modules is read file by file, as asked (peak RSS 305 MB -> 87 MB on the 60 MB case). - Docs: `sys.path` / `sys.meta_path` changes that make a same-named module elsewhere the one imported are not followed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Bind the located definition, prove readers by their binding, check the lazy hook's import (#879 PR review) - The resolved locator decides between same-named definitions even when the edge's own source holds exactly one. After source deduplication, a `tools.lookup` imported as `shared_lookup` no longer binds the local `lookup`. - A read-only reader is proven by its binding: - a bare builtin only when nothing binds it and no star import may; - a standard-library reader only when imported from that module (`json.dumps`, `from pprint import pprint`); - a logging method only on `logging` or a `getLogger()` logger. A `print` imported from the application, or an `.info()` on its own object, makes the list dynamic. - A lazy `__getattr__` is the submodule idiom only when its alias imports the submodule asked for: `from . import alternate as memory` is not. perf: read materialized Git tree blobs through batched cat-file (#878) _materialize_isolated_tree spawned one `git cat-file blob <oid>` per file in both its entries and links loops. Once `diff --application --scope .` materializes the whole tree (#877), that dominated the run: one side of TencentCloud/CubeSandbox took ~180 s to archive (the #686 cost class). Blobs are now read by _isolated_blobs: one `cat-file --batch-check` types and sizes every object (a missing or non-blob object refuses with a ConfigError before any content is read), then `cat-file --batch` runs of at most 64 MiB each return the content, with strict framing checks. Only full object IDs reach the batch. The link-text reads in _scope_through_boundary_links use the same reader. Every blob is still hashed against its tree entry's oid before it is written; path handling, containment and escape checks, blobs-before-links order, placeholders and the final digest comparison are unchanged. The reads go through _run_git_dir (now accepting input=) and the existing _run_process boundary, so no subprocess call site is added; only the two line pins in test_adapter_static_only.py moved. diff --application --scope . with #877: CubeSandbox 412 s -> 46 s, dlt 203 s -> 45 s, identical rows. fix: compare application wiring past unrelated links and submodules; read google.adk.Agent (#877) * fix: compare application wiring past unrelated links and submodules `diff --application` at the root scope took the unscoped archive route, which refuses every symlink and gitlink. On 2026-09-25, 15 of the first 27 runs against open third-party SDK/ADK PRs exited 2 on a path the reader never opens: `CLAUDE.md -> AGENTS.md`, a linked skill directory, a `VERSION` link leaving the tree, a vendored submodule. - Materialize every scope, the root included, through the scoped verified materializer, which recreates links rather than refusing them and packs the tree instead of the history. - Never read a Python input through a link: a target this scope already reads is compared at its own path, and any other in-scope target is a gap over the link's path. Reading the alias made one agent two ambiguous ones and hid the real file's change. - Record gitlinks behind an opt-in `archive_tree(record_gitlinks=True)`, materialized as the empty directory an unpopulated checkout leaves. An unchanged gitlink commit is named in limits; any other is a coverage gap over its path. Other archive callers still refuse. - Stop turning the host-configuration census's link count into application coverage gaps. - Read `google.adk.Agent`, the package-root re-export, as an ADK agent constructor (`from google.adk import Agent`). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix: keep unread links and scope-level submodules from reading as removals PR #877 review found two cases where unread input became conclusive negative evidence: - Discovery drops a path that does not resolve, so replacing `agent.py` with a dangling link read as a definite removal, `compared`, with no gap. The dropped host-census gap had been the only fallback, and it also covered a source directory replaced by an absolute or dangling link. Census every link under the scope directly, without following it: gap a `*.py` link that does not alias an input the scope already reads, a directory link holding Python outside the scope, and a link resolving to nothing in the tree where the other side reads source at or beneath it. An unchanged `agent/VERSION -> ../../VERSION` still changes nothing. - A gitlink at the selected scope itself gave a gap with source ".", which covers no relative binding path, so `agent.py` read as a definite addition or removal. Make that gap scope-wide. fix(diff --application): key SDK identity on imports; unobserved agents are not removals (#873) * fix(diff --application): key SDK identity on imports; unobserved agents are not removals On speechmatics/speechmatics-academy#142, a LiveKit voice agent moved its tools into an Agent subclass that passes them through super().__init__. `diff --application` printed `compared` with four false REMOVED rows. Two defects: - Framework identity. Discovery counted any bare `@function_tool` as the OpenAI Agents SDK, and the SDK reader recognized `function_tool` and `Agent` by spelling alone, so `livekit.agents` symbols were read as the SDK's. Both now key on import provenance: the absolute `agents` / `openai_agents` package, or an unimported name (unless a foreign wildcard could supply it). Relative imports are no framework's signal. - Unobserved is not removed. An agent observed on one side whose file on the other side still assigns its name (or passes it as `name=`), via a construction no reader supports (subclass, factory, Agent[Ctx], clone), now records a scoped coverage gap for that agent: `partial`, rows `not_established`. Handoff-only references do not count as an observed construction. A genuinely deleted agent stays an established removal. Regression tests fail on the unfixed tree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * fix(diff --application): resolve SDK names per scope; imports keep an agent named Addresses the two P2 findings on #873. - Import provenance was flattened across lexical scopes: one file-wide import map, accepted if any origin was the SDK, so a LiveKit `Builder(...)` became an SDK agent when a sibling function imported the SDK under the same alias. `_SdkNames` now resolves each spelling in the scope that uses it (nearest binding scope, class bodies skipped for nested code, global/nonlocal obeyed; decorators and defaults in the enclosing scope). An import there must be the SDK's; a parameter or local assignment is not. `function_tool` is decided per decorator node, not as a file-wide set of spellings. The walk is iterative. - The missing-agent check ignored import bindings, so `from agent_factory import agent` (or `exported as agent`) still gave a definite removal. Import aliases now count as the file still naming the agent. Regression tests for both fail on the previous head. feat: compare application agent wiring without prior setup (#871) * feat: compare application agent wiring without prior setup * fix: address application comparison review findings * test: normalize colored CLI errors in review regressions fix(ci): prevent checkout module shadowing in non-GitHub recipes (#870) * fix(ci): prevent checkout module shadowing in installation recipes * test(ci): cover quoted and qualified Python installation commands Read agent launches in CI workflows (#823) (#850) * Read agent launches in CI workflows (#823) The workflow grant read triggers, token permissions, reusable-workflow secrets (#693) and step `uses:` references (#771), and nothing about how a coding agent is launched inside a job. Changing a claude-code-action's `claude_args` from `--allowedTools "Read"` to `--permission-mode bypassPermissions --allowedTools "Bash(*)"`, adding a `claude -p --permission-mode acceptEdits` run step, or checking out the pull request head in a `pull_request_target` job each gave "No static host-grant changes detected", and a move to `issue_comment` with `pull-requests: write` gave a row with no agent context. The workflow grant now lists, as text that is never executed, fetched or evaluated: - `agent_launches[]`: a step whose `uses:` is anthropics/claude-code-action, anthropics/claude-code-base-action or openai/codex-action (any ref, any case) with the documented permission inputs it sets, or a `run:` that is one literal simple command starting with `claude -p/--print` or `codex exec`, with its documented permission flags under their primary spelling. The prompt, `--model` and undocumented flags are not compared. `job_secrets` names the secrets the job references, as context only. - `checkout_refs[]`: each actions/checkout step's `with.ref`, null for the default. Each job's multiset of launches and refs is compared, never the step label, so a rename or reorder is quiet; a difference is one `changed` row on the existing workflow row naming `job/step` and both values. Direction is claimed only by documented rules a job's launches gain, read from literal values: bypassed permission checks (either spelling), bypassed approvals and sandbox, a danger-full-access sandbox, `safety-strategy: unsafe`, or a `*` user gate. Those raise `workflow_agent_widened_<added|changed>` and make the row widened; every other edit, including `--allowedTools "Bash(*)"` (#824's to rate), is `changed`. A workflow row whose workflow runs an agent ends its `why` with the untrusted-input trigger, write scopes, secrets and pull request checkout beside each agent step; it is a note and moves no direction. A compound `run:`, an expansion or an expression is `unresolved`, publishes none of its text and records a non-blocking coverage issue naming `job/step`, as an unread secret value does (#693): coverage stays complete, adding one is a row that claims no effect, and an edit inside one is not reported. Scripts, composite (#701) and unknown actions, and agents reached through npx/timeout/sudo are listed as unread surfaces. Every value, ref and secret name goes through the #802 label redaction; a rewritten one is null with `redacted` and is neither published nor compared. Contract 40 and host-grants 0.6 shipped in 1.1.0, so host-grants inventory, baseline and drift move to 0.7 (the 0.6 schema files are untouched) and the runtime contract to 41. A 0.4-0.6 baseline holding a workflow grant is incomparable (`baseline_workflow_agent_launches_unavailable`); one without a workflow stays comparable. Verifier 0.20 and capability diff 0.3 do not move. No check id is added or removed and `check` decides as before. Docs: the support page (tables, rules, note, limits, a new unread-surfaces bullet), a STABILITY "Migration Note: Unreleased", CHANGELOG `## Unreleased` above 1.1.0, the agent contract page, the distribution-surfaces `capability_diff` row and its parity comment, the Stop hook note, version tables and pins, and a rebuilt llms-full.txt. The pilot ledger's source-tree column was re-measured on the Route H fixture: identical cells to the 1.1.0 engine apart from contract 41 and inventory 0.7, with byte-identical diff rows. Tests: new tests/test_workflow_agent_launches.py covers the four reproduction cases, what is and is not read, every unsupported shape, direction rules and non-rules, the acceptance's negative controls, redaction with a CLI canary sweep across every published output, the 0.6 baseline migration, schema validation, and the same row on diff, verify, the PR comment, check, the control envelope and the Stop hook. Existing tests move to 0.7/41. The host-config and cold-start replays reproduce their committed outcomes and run-of-record scores, and the sample goldens are unchanged. Closes #823 * Read no widening rule from an agent input that holds an expression (#823) The support page and STABILITY say a documented rule is read from literal values only, so a value holding `${{ }}` never meets one. The `*` user gate did not apply that: `allowed_non_write_users: "${{ vars.USERS }}, *"` raised workflow_agent_widened_changed. Every rule now skips a value holding an expression, as the claude_args and codex-args rules already did, and a test pins the gate case. * Address review cycle 1 on agent launches in CI (#823) F1. `claude_args` and `codex-args` were split with the `run:` shell tokenizer, which gave up on a newline or an unquoted `(`, so a widening written the way the actions document it gave `changed`: `claude_args: |` on several lines, `--allowedTools Bash(git:*) --dangerously-skip-permissions`, a `# comment` line, or a multi-line `codex-args`. Each input is now split the way its action splits it. The Claude actions' parse-sdk-options.ts drops full `#` lines, makes `()|&;<>` literal and reads the rest with shell-quote (newlines are whitespace, `$NAME` is empty, an unquoted `#` ends the input, a `--` word is always a flag); codex-action reads a JSON array of strings or string-argv. The `run:` tokenizer is kept for `run:` steps only. The published `claude_args` is the text the action parses, so a full-line comment is neither published nor compared. F2. Any URL path made a whole setting `redacted`, so a `--dangerously-skip-permissions` beside `https://example.com/style-guide` gave no row, and every `plugin_marketplaces` value was never compared, under a limit that wrongly called it credential-shaped. The documented rules are now decided from the declared text when the workflow is read, before anything is withheld, and published on the launch as `widening_rules` (rule and setting), which the comparator keys on, so redaction never hides a rule. A URL publishes its scheme and host with `<redacted-path>`, as an MCP server URL does (#723), and the rest of the value is published and compared. Other credential-shaped text in a setting or checkout ref (a token shape, an assignment, a bearer or header value, URL userinfo) is published redacted and makes the workflow a blocking limit through `_uncompared_workflow_text`, as a redacted step reference does (#767). F3. JSON in `settings`, `mcp_config`, `--settings`, `--mcp-config` or any argument word was published verbatim, env values and apiKeyHelper included. A JSON object now publishes what `.claude/settings.json` and `.mcp.json` publish: key names, with `env`/`headers` values, apiKeyHelper and secret-named values `<redacted>`, as canonical JSON. A codex `--config` override under env, headers or a secret-named key publishes `<redacted>`. Text that starts like JSON and does not parse is withheld (`unparsed_json`, a non-blocking limit). The agent-launch canary sweep now carries JSON-shaped canaries and their digests across diff, audit, check, verify, the PR comment and verifier.json, and a second sweep covers the refused credential-shaped case. Nonblocking: a rule gained where the job launched that agent before only in an unread form is named and not claimed (the `unknown_before` rule); a quoted word starting with `#` no longer makes a `run:` compound; `anthropics/claude-code-action/base-action` is read as the base action; the STABILITY note says a prompt after a variadic flag is compared; and `__all__ =[` is spaced. Host-grants 0.7 is unreleased, so `widening_rules`, `unparsed_json` and the base-action agent value extend it in place; the 0.7 schema files are regenerated. The support page, STABILITY migration note, CHANGELOG Unreleased entry, contract summary, integrations Stop-hook sentence and the capability_diff distribution-surface row say the same. * Address review cycle 1 on agent launches in CI (#823) The second review of #850 at 25c13ce8 found two P1 and four P2 defects. The branch is rebased onto origin/main daa4ad5f. Attached JSON values are withheld (F1). A --settings={...} or --mcp-config={...} word inside claude_args, and codex's attached -c<override> / -c=<override>, published their env, header and apiKeyHelper values verbatim, because only a word starting with "{" was withheld. _withheld_words now splits a --name=value word and withholds the value, and reads codex's attached -c through _withheld_config, as clap reads it. The canary sweep carries both spellings. Redacted prose no longer refuses the comparison (F2). The #802 label redaction rewrites ordinary prose ("never print bearer tokens", "Authorization: headers"), and a redacted agent setting made the workflow a blocking limit. That hid every row beside it and made every check incomparable while the workflow existed. The rules are already read from the declared text, so a redacted setting is now compared by its published text and its widening_rules. uncompared_agent_launch_texts names it as a non-blocking limit, and the row cell shows the redacted text. A redacted checkout ref still refuses, as a redacted step reference does (#767), because it names the code a job runs. A renamed job's launch moves its rules (F3). _agent_rule_gains keyed a rule on its job, so renaming a job that launches a bypassing agent was a widening. agent_rule_gains now pairs a rule one job gains with the same rule another job lost, when the launch that met it left that job: the job no longer launches that agent, or the same launch (agent_launch_key less the job) now runs in the gaining job. The why names the move. A second job gaining a rule, or a different launch gaining one while the first job still launches that agent, still widens. An expression no longer turns off every rule, and the row says what it leaves unread (F4). Rules are read from literal text a ${{ }} expression cannot reach: - the words of claude_args or codex-args before the first expression, less the word it touches and any quoted run still open at it; - the elements of a JSON-array codex-args before the one holding it; - the gate entries that hold none. A setting holding an expression is published with holds_expression (the unreleased 0.7 schema extends in place), and a row that changes it says the text the expression reaches is not read. A gain where the job's launch held an expression before, in the input the rule is read from, is named and not claimed, as unknown_before is. That also fixes a false widening at 25c13ce8: replacing --model ${{ vars.CLAUDE_MODEL }} with --model opus beside --dangerously-skip-permissions was reported as gaining the bypass. Two non-blocking fixes: - _published_value reads each expression as one word, so an expression in a URL's userinfo is withheld with it rather than garbling the URL and publishing its path. - A codex --config value that starts like a table and does not parse, as string-argv leaves a quoted one, is withheld as unparsed_json. Rebase (F5). CHANGELOG keeps #853's #778 line beside #823's under github_action row. llms.txt and ai-search-summary state the source tree as contract 41, unreleased, ahead of the published v1.1.0 (contract 40), as test_public_surface_contract requires while the two differ. The pilot ledger's source-tree column was re-taken on the rebased tree, through ./shipgate beside the engine of e3c6cb0c, the commit v1.1.0 was cut from: - the only differences are contract 40 -> 41 and inventory schema 0.6 -> 0.7; - diff rows are byte-identical; - check JSON differs only in the launcher path its next action names. The support page, the STABILITY preamble and migration note, the CHANGELOG entry, agent-contract-current, the Stop hook sentence in integrations, the capability_diff row in distribution-surfaces with its parity comment, the 0.7 schema files and llms-full.txt are updated to match. * Address review cycle 2 on agent launches in CI (#823) The third review of #850 at c7f67550 found one P1 and one P2 defect and four P3 notes. The branch is on origin/main 44b9e05d; no rebase was needed. A JSON setting publishes its shape, not its free text (C2-F1). _withheld_json published the whole _redact_secret_values tree. That tree is the host readers' digest input, not what they publish: it keeps every string outside env, headers and secret-named keys. So an mcp-remote --header "Authorization: Bearer ..." argument in --mcp-config, and a hook's curl command in settings, reached diff text, diff --json, audit --host --json, the PR comment and verifier.json. _json_shape now keeps key names, numbers, booleans and null, redacts what the host readers redact, and replaces each other string with <withheld:...>, a 12-hex digest of redacted_config_sha256 for that string, so an edit to it is still a changed row. It keeps only the strings a host reader publishes: - a permissions.allow/ask/deny rule and a documented Claude Code setting's value (defaultMode, the switches, enabledMcpjsonServers entries); - an MCP server's command name and its URL's scheme and host, followed by the digest when they drop a command's arguments or a URL's query. A codex --config table or array is read under its key path, so mcp_servers.gh={command="gh", ...} keeps its command name. The canary sweep adds the mcp-remote header and the hook command in every spelling (action input, claude_args, CLI flag, codex -c table). Re-running the reviewer's two repositories through ./shipgate gives 0 canaries in every output. The documented bypasses written through listed inputs widen (C2-F2). - Claude Code settings written as JSON meet bypass_permissions when their defaultMode is bypassPermissions, read by claude_setting_values as the settings reader reads .claude/settings.json. This covers the action's settings input, a --settings value in claude_args, and the CLI's --settings flag. A path is not read, and a settings value holding an expression meets none. - openai/codex-action's permission-profile: :danger-full-access meets danger_full_access. ":danger-full-access" is Codex's reserved name for its built-in full-access profile (BUILT_IN_PERMISSION_PROFILE_DANGER_FULL_ACCESS). Mode inputs are now (input, value, rule) triples. settings and permission-profile join _RULE_SETTINGS, so a gain after an expression in either is named and not claimed, as for claude_args. P3 notes: - allowed_bots: "*" now reads "accepts runs triggered by any bot" through agent_rule_text. - Which of two gaining jobs a moved rule goes to no longer depends on declaration order: the same launch arriving is matched before a job that merely stopped launching the agent. The support page, the STABILITY migration note, the contract summary, the schema docstrings and the CHANGELOG entry now say what a structured value publishes instead of claiming it publishes what the host readers would, and list the two rules. host-grants 0.7 is extended in place; it is unreleased. * Address review cycle 3 on agent launches in CI (#823) Rebased onto origin/main 01777037 (#852, #821), which had already moved the unreleased runtime contract 40 -> 41, verifier 0.20 -> 0.21 and capability diff 0.3 -> 0.4. This change now extends contract 41 in place instead of minting it: one comment in schemas/contract.py names #821 and #823, and the texts that said verifier 0.20 and capability diff 0.3 were unchanged now say capability_diff registry row keeps #852's two added roots and both paragraphs; STABILITY keeps the #821, #823 and #827 notes, with #821's "host-grants stays 0.6" and #827's version sentence corrected for a tree that also carries host-grants 0.7; CHANGELOG keeps both Unreleased entries. llms-full.txt is rebuilt. The pilot ledger's source-tree column was re-measured on the rebased tree beside e3c6cb0c: identical cells except contract 40/41 and inventory schema 0.6/0.7, identical diff text, rows, check JSON and inventory apart from its schema version, and diff --json / verifier.json differing only in #821's schema versions and coverage members. A codex exec step that selects the full-access sandbox through --config was a changed row while --sandbox danger-full-access widened, and -sdanger-full-access was not read at all. The flag reader now reads a short flag's attached value as clap does (-s<mode>, -s=<mode>, -c<override>, -c=<override>), and a --config override that sets sandbox_mode to danger-full-access, or default_permissions to :danger-full-access (what codex-action's permission-profile input passes the CLI), meets danger_full_access, its value read as the CLI's parse_overrides reads it. As the CLI resolves them, the last override of a key counts, default_permissions outranks sandbox_mode, and a --sandbox flag outranks both, so an override beside --sandbox meets none. In codex-args such an override meets none, because the action appends its own --sandbox or default_permissions selection after codex-args. The support page and the STABILITY Direction bullet list the spellings. Withholding a quoted URL inside a JSON-array codex-args element no longer drops the element's escaped closing quote, so the published value stays the JSON array the support page describes. * Address review cycle 2 on agent launches in CI (#823) An agent CLI inside a double-quoted $(...) or a backtick substitution, or after a shell reserved word, got no launch, no row and no coverage issue: gh pr comment --body "$(claude -p --dangerously-skip-permissions ...)" read as "No static host-grant changes detected", while the support page said such a run: is listed as unresolved once for each agent CLI it starts at the head of a command. The word splitter reads a double-quoted substitution as one word and a backtick as no operator, and the head reader took then, do, { and ! for the command. _command_substitutions now lists the text of each $(...) and backtick substitution outside single quotes, double-quoted ones included, and an agent CLI heading a command inside one is an unresolved launch (shell_expansion for a single top-level command, compound_command otherwise), publishing none of its text. An escaped \$, a single-quoted '$(...)', a comment and $((...)) arithmetic are not substitutions, so echo "claude -p ..." stays unlisted. The scanner keeps an explicit stack and reads each nested substitution as `_` in the one around it, so no nesting depth recurses or re-splits text. When the run's quoting does not balance, each line's substitutions are read as well. A command's head is now read after shell reserved words (!, time [-p], {, if, then, elif, else, while, until, do, function NAME), and a command after one is not a simple command, so its launch is compound_command and never read. The compound_command and shell_expansion limit phrases, the launch schema description (host-grants 0.7 schemas regenerated), the support page and the STABILITY bullet say so. A read launch that became a form this audit does not recognise (npx, a path, codex options before exec) was worded "a step no longer launches an agent", though the step still starts one. The removed case now reads "a step no longer declares an agent launch this audit reads", and a row whose workflow still exists adds that the step may still start an agent in a way this audit does not read, naming those forms. codex with an option before exec is listed under Known unread surfaces and in STABILITY's "What is not read". Tests cover each shape the review named, the negative controls, the diff and audit --host route of the reproduction, the reworded removal for npx, a path and codex root options, and 20000 nested substitutions. * Address review cycle 3 on agent launches in CI (#823) A step that drafts a PR comment in a quoted here-doc, such as cat > comment.md <<'EOF' with "Reproduce locally with `claude -p ...`" in its body, was listed as an unresolved agent launch: a "high changed" row saying a step now launches an agent, and a GitHub coverage issue. The shell passes a here-doc's body to its command as input and, with a quoted delimiter, expands nothing in it, so the step starts no agent. The phantom also hid a real gain: a --dangerously-skip-permissions step added beside it was "not counted as a widening" because the job "launched the agent in a form this audit does not read" before. The substitution scanner read backticks and $( across the whole run:, and the word splitter read a body line starting with claude -p as a command head. _here_documents now removes each here-doc's body and closing line before the run: is split or scanned. <<, <<- and a quoted, escaped or partly quoted delimiter are read outside quotes, comments and $((...)) arithmetic, inside a $(...) or backtick substitution too; <<< is a here-string; bodies start after the opening line and follow in order when one line opens several. A body is never a command. Only an unquoted <<EOF body's substitutions, which the shell runs, are read, with quotes and # in it literal, so '$(claude -p ...)' there is listed and was missed before. A here-doc with no closing line is left in the text and read as ordinary lines, so a misread << hides no later command. Closing lines are looked up by content, so many here-docs cost one pass. Every shape was checked against bash with stub claude and codex functions. A launch that only became a form this audit does not read no longer counts as having left its job. The moved-between-jobs rule now takes a rule as moved only when the job that met it no longer exists, or when the same launch now runs in the gaining job and that job's own launches of the agent still run there or in the job it left (a moved step, a swap). Job a going from claude -p --dangerously-skip-permissions to npx @anthropic-ai/claude-code -p while job b's launch is edited into a bypassing one is now widened, where it was "moved between jobs". A step moved and edited while its job remains is claimed too. A checkout step on one side only, such as a default checkout in an added job, is worded "a step now declares a checkout" or "a step no longer declares a checkout" rather than "a checkout's declared ref changed". The support page, STABILITY's migration note and the CHANGELOG entry say which here-doc text is read, name the two shapes still read as lines, and state the narrowed move rule. * Name an unchanged instruction limit reached through an in-tree link instead of refusing (#822) An instruction file whose limit this entry cannot resolve, such as a SKILL.md whose metadata holds a non-string value, is named as an unchanged limit when a change leaves it alone (#721). Reached through an in-tree link the reader reads through (#700) - `.claude/skills -> ../.agents/skills`, a per-skill link or a file link - the same untouched file refused the whole comparison: diff, verify and the manifest-free PR comment printed `base_inventory_incomplete; head_inventory_incomplete` and no row, hiding a removed deny rule beside it. `unchanged_limits` asked blob_path_unchanged whether the source, the path the link is read under, was one regular file in Git, and a path through a link never is. blob_path_unchanged now resolves the path on each side the way the reader reaches it: from Git tree entries for the base and a commit head, and for a working-tree head without following any link (each component's own entry, each link's own text, the file's unfiltered hash). It holds only when both resolutions are equal - every link at the same path with the same text, every other component a directory - and the file they land on is the same regular-file blob at the same in-tree path. A link is followed only under the rules the reader and the base archive already use (#700, #711): a relative text that lands inside the tree after normalization, directories above where it lands, at most eight links, and links at one component of the path. A path no link reaches still takes the one `ls-tree` it took before, and blob object IDs are still compared, so no filter or textconv can make two byte sequences equal. Everything that asks the proof moves together: unchanged_limits in `diff --json` and verifier.json, the text and PR comment, the unchanged limits a partial comparison may carry (#808) and the shared plugin-reference limits check leaves out (#714). check's boundary result still cannot name a limit, so it now refuses these comparisons with unchanged_limits_not_representable, as it does for a limit at its own path; its rows, decision and violations do not move. No schema, member, reason code or check id is added. The metadata value is not coerced. tests/test_linked_unchanged_limits.py holds every layout (directory link, per-skill link, file link, a chain of file links) to the direct result on diff, verify, the PR comment and check; keeps the negative controls refused (skill added or edited behind the link, link retargeted to an identical copy, link text rewritten to land on the same file, link replaced by a directory or the reverse, a later hop retargeted, working-tree-only retargets and edits); and holds the proof to exactly the links the reader reads through, across dangling, looping, absolute, escaping, over-long, intermediate and nested links. The #812 coverage case that pinned the refusal is retired, and the static-only allowlist follows its two call sites down the file. * Address review cycle 1 on linked unchanged limits (#822) The STABILITY migration note said that where the other routes now compare past a limit reached through an in-tree link, `check` refuses with `unchanged_limits_not_representable` and its rows stay empty. That holds for an instruction file's limit, which `check` cannot leave out, but not for a plugin-reference limit: `_without_shared_plugin_reference_limits` asks the same unchanged proof, so a `parse_failed` `plugin.json` that is a file link to an unchanged target is now left out exactly as one at its own path has been since #714. `check` then compares and publishes the removed `deny` row in its boundary result and its control envelope's `capability_rows`, where it refused with `base_inventory_incomplete` / `head_inventory_incomplete` and no row, and `diff`, `verify` and the PR comment, which withheld `plugins/demo` as `partial` (#808), are `comparable` with the limit in `unchanged_limits`. `check`'s decision, violations and control state, and `verify`'s control state and next action, do not move. The note now splits `check` by whether it may leave the limit out, adds the `partial` to `comparable` move, and says "every row's value" where it said "every row". The CHANGELOG entry mirrors it, the distribution-surfaces row and `docs/host-boundary-support.md` no longer say any change behind the link refuses the comparison (an edited plugin manifest keeps it `partial`), and `tests/test_linked_unchanged_limits.py` pins the plugin manifest at its own path and behind a file link, on every route, with the edited-target control. The `blob_path_unchanged` and `_reader_path` docstrings now state the proof as necessary, not sufficient: it follows the links on the way to the path, not every condition the reader puts on reading a whole linked directory, and a side whose reader does not read the path carries no limit there. A test holds a head that adds a link inside the linked directory to a refusal on every route. The two pinned call-site lines in `cli/verify/git.py` move with the docstrings. * Address review cycle 4 on agent launches in CI (#823) Four review cycles each found a new shell form (quoted words, $(...), here-docs, comments, reserved words) that the run: reader mis-read. The cycle-4 finding was the fourth: a `#` comment was split as commands, so an apostrophe in a comment hid a launch and a command line in a comment invented one. Parsing arbitrary shell cannot converge, so the reader now claims only forms it reads exactly and names every other one as a limit. - A run: step is an agent launch only when it is one line of plain words (letters, digits and `_ . / : = , % + -`, separated by spaces or tabs), run by bash, sh or no declared shell:, whose program's file name is `claude` with -p/--print or `codex` followed by exec. Every POSIX shell runs such text as exactly those words; a cross-check of 2251 generated texts in sh, bash, dash and ksh agrees on every one. - Any other run: that mentions claude or codex as a word of its own is an `unread_agent_runs[]` entry (job, step, agent) on the workflow grant: a non-blocking coverage issue in audit --host that publishes none of its text, is never compared, so it gives no row, and never says whether the step starts an agent. It takes no gain from another launch; only when an unread step goes and a read launch is added in the same job is a rule that launch meets named and not claimed, since it may be that step rewritten. - claude_args and codex-args are read only as a plain list of words (the same characters and parentheses, across blanks and newlines, with no --settings or --mcp-config flag). Any other value, a ${{ }} expression included, is `unread_arguments`: published only as a digest, so an edit is a changed row, and read for no rule; a rule a launch gains where that input was unread before is named and not claimed. - A codex --config override publishes its key; its value is <redacted> under env, headers or a secret-named key, as written for sandbox_mode, default_permissions, approval_policy and model, and a digest otherwise. The word after a secret-named word such as --token is <redacted> and the value is then published redacted, a named limit. - The shell tokenizer, the command-substitution and here-doc scanners, the reserved-word reader and the shell-quote and string-argv emulations are removed, with the expression-prefix reading of argument inputs. JSON is read only in the settings and mcp_config inputs. Host-grants 0.7 is unreleased, so its schema changes in place: the launch `unresolved_reason` is only inputs_not_a_mapping, a setting may be unread_arguments, and the workflow grant adds unread_agent_runs. The support page, STABILITY, CHANGELOG, the current contract page and llms-full.txt describe the tightened reader. * Address review cycle 5 on agent launches in CI (#823) C5-F1: the move rule took a launch that only stopped being read to have left its job. With job a running `claude -p --dangerously-skip-permissions Review` in the base, quoting its prompt (an unread step) or running it through `npx` while job b added the same plain launch gave a `changed` row saying the launch "moved between jobs (a → b)" and "already met that rule in the job it left", with no widening signal. Job a still runs it. - `same_launch` now refuses a move while the losing job may still run the launch in a form this reader does not read: an unread step of that agent stands at a step label the lost launch held, the job has more unread steps of that agent than before (the launch may have moved to another index), or one of its read launches of that agent holds an expression or an unread argument input in an input the rule is read from. The review's M5 and M6 are now `widened` and name "a step no longer declares an agent launch this audit reads (a/steps[0])"; a real move, a rename, a swap, and a move beside an unread step the job already had stay moves. C5-F2: docs/integrations.md said an unread `run:` is a non-widening row that diff and the PR comment show. It gives no row; only an unread argument input is a row. The sentence is split accordingly. Nonblocking items fixed: - `_job_secrets` walks each container once, without recursion. A job `env` holding itself through a YAML alias raised RecursionError, and a ten-way fan-out eight levels deep did not finish; both now read in well under a second. - A `settings` or `mcp_config` value that is neither a JSON object nor a plain file path (path characters, and a `${{ }}` only as a plain context reference) publishes only a `<withheld:…>` digest. A comment line before JSON published the `env` value that JSON held. - The last of a repeated `--permission-mode` counts, in `claude_args` and in a plain `run:`, as parse-sdk-options and the CLI keep it. - The docs/distribution-surfaces.md capability_diff row no longer says an unread argument input or an unresolved launch is named only by the inventory and audit --host: a row reporting its launch names it too. The support page, STABILITY (Direction and What is withheld) and the CHANGELOG entry say the same. Tests: tests/test_workflow_agent_launches.py grows to 290 cases. The five new move-rule widening cases, the alias test, the two json-or-path cases and the four --permission-mode cases fail on the previous head; the moved-beside-an-unread-step case and an end-to-end diff/audit test of M5 and M6 are added beside them. * Address review cycle 6 on agent launches in CI (#823) C6-F1: the cycle 5 move guard only refused a move when an unread step of that agent stood at a step label the lost launch held, or the job had more unread steps than before. Merging job a's `npm i -g @anthropic-ai/claude-code` step into `claude -p --dangerously-skip-permissions Review` while job b added that plain step (V1), or removing an unread `echo "claude"` step while the launch became a quoted, unread step at another index (V2), kept the unread count and missed every held label, so the row said the bypass "moved between jobs (a/steps[1] -> b/steps[0])" and gave no widening signal. Job a may still run the launch. An unread step carries no text that tells which launch it is, so `may_still_meet` now refuses the move whenever the losing job has any unread step of that agent, wherever it stands and whether or not it was there before (option (a) of the review). V1 and V2 are now `widened`, with the "may still start an agent" caveat, and name "a step no longer declares an agent launch this audit reads". The cycle 5 guard that kept a move beside an unread step the job already had (`claude mcp add x`) now widens, in the safe direction; it moves to the widening cases, and a move beside a step that names no agent stays a move. C6-F2: `_codex_override` caught only `TOMLDecodeError` and `RecursionError`, and `tomllib` raises a plain `ValueError` for an integer past Python's 4300-digit limit, so `run: codex exec -c sandbox_mode=<5000 digits> Review` crashed `diff`, `audit --host` and `check` (exit 1) and `verify` (exit 4, internal_error). It now catches `ValueError`, which `TOMLDecodeError` subclasses; such a value is read as text and selects no sandbox, so the row is `changed`. `-c default_permissions=<4400 digits>` and `codex e --config=sandbox_mode=<4301 digits>` are covered too. Wording: the support page, STABILITY (Direction and What is not read) and the CHANGELOG say a launch has not left a job that keeps any named unread step of that agent or a launch of it with an unread input the rule is read from. The support page and STABILITY also name the one case where a launch that stops being read is worded as moved rather than as no longer declaring a launch: another job adds the same launch while the job it left keeps none of those, as when the launch became a script in the same change. Nonblocking items fixed: - docs/integrations.md: "One that gains a rule" read as the unread `run:` of the sentence before it; it now says "An agent launch that gains a rule". - The HostWorkflowAgentSettingV7 docstring, published as the 0.7 schema's description, states that a `settings` or `mcp_config` value that neither starts like a JSON object nor is a plain file path publishes only a `<withheld:...>` digest. The 0.7 inventory and baseline schema files are regenerated. * Address review cycle 7 on agent launches in CI (#823) C7-F1: `may_still_meet` looked for unread steps only under the losing job's name, which holds nothing once that job is renamed or removed, and `job_left` accepted any job that no longer exists. So renaming job a to a2 while quoting its `claude -p --dangerously-skip-permissions Review` (R6), or running it throu…
Summary
Fixes two reach bugs in
diff --application(#871). They were found on 2026-09-25 against open third-party PRs that edit OpenAI Agents SDK / Google ADK tool wiring (the 134-PR corpus for #868).1. The root scope refused any repository that has a symlink or submodule.
At
--scope ., 15 of the first 27 runs exited 2 withGit tree contains unsupported external binding at <path>. None of the paths was application source:CLAUDE.md -> AGENTS.md: dlt#4417, jaeger#9636, omnigent#6611, weave#7948.claude/skills/…or.agents/skills/…directory: asterinas#3834, vllm#57322, kagent#2788agent/VERSION -> ../../VERSION: CubeSandbox#1508.pylintrc: tensorflow#128063The root scope passed
scope=None, which sends it down the unscoped archive route. That route refuses every link and gitlink, and it also packs the full commit history.2.
from google.adk import Agentwas not read as an agent.The ADK reader's constructor table lacked
google.adk.Agent, which is the package root's re-export ofgoogle.adk.agents.llm_agent.Agent. As a result, CubeSandbox#1508'sroot_agentgavenot_established("No supported application agents were established").What changed
cli/application_diff.py). Every scope goes through the scoped verified materializer, which already recreates links instead of refusing them (A symlink or submodule anywhere in the base tree makesshipgate diffrefuse the whole repository — 4 of 12 public repositories, includingCLAUDE.md -> AGENTS.md#688/Read through an in-tree symlink at a boundary path:CLAUDE.md -> AGENTS.mdand.claude/skills -> ../.agents/skillsstill leave the inventory incomplete #700). It also packs only the tree, not the history.observe()builds its own list of every link under the scope, without following links and including dangling ones. It no longer relies on discovery's inventory, which drops a path that doesn't resolve.CLAUDE.md -> AGENTS.md, or a linked skill directory the scope already reads at its own path.*.pylink whose target is a Python input the scope already reads is compared at the target's path. Reading the alias as well turned one agent into two ambiguous ones and hid the real file's change; main does this at a narrow scope. Any other*.pylink is a gap over its own path: dangling, leaving the scope or the repository, or landing on something that is not a Python input._reconcile_unresolved_links). Replacingagent.py, or a source directory, with a dangling or absolute link is thereforenot_established, never a removal. An unchangedagent/VERSION -> ../../VERSIONchanges nothing.archive_treegains an opt-inrecord_gitlinks(cli/verify/git.py). The default still refuses, so the host route andverifyare unchanged; a test pins this.{path: commit}.limitsand does not make the comparisonpartial. This mirrors the unchanged-limit reasoning of An unchanged partial or experimental surface refuses every diff in the repository #721/An unchanged instruction limit reached through an in-tree link refuses the whole comparison instead of being named unchanged #822.source=None). The result ispartial, nevercompared.host_discovery_incomplete_pathscounts every link that could conceal a host path, so each recreated link would otherwise readDiscovery could not read CLAUDE.mdand turn the comparisonpartial. No test pinned it. The link census above replaces the one thing it did catch for applications, a dangling link where source used to be.google.adk.Agentis added toAGENT_CLASS_NAMES(inputs/google_adk.py). It covers bothfrom google.adk import Agentandimport google.adk as adk; adk.Agent(...).application_comparison_schema_versionstays"0.1". No new JSON field was added: submodules surface through the existinglimitsandcoverage_gaps.Reviewer notes
Review round 1 (
3a0c00c7). Both findings were real regressions from the first commit, and both are fixed:*.pylink read as a definite removal. So did a source directory replaced by an unresolved link, a case the review didn't list; main caught both through the host-census gap.".", which covered no binding path.13 new regression cases fail on
9f325c3band pass now. Details are in the review threads.Merged
main(761d5c0f), bringing in fix(diff --application): key SDK identity on imports; unobserved agents are not removals #873. OnlyCHANGELOG.mdanddocs/distribution-surfaces.mdconflicted, and both sides' text was kept. fix(diff --application): key SDK identity on imports; unobserved agents are not removals #873's_unobserved_agent_gapsruns after_align_exact_moves; this PR's submodule and unresolved-link steps run before it.Pinned decision changed.
test_gitlink_refusal_is_not_a_successful_comparisonexpected exit 2. It is nowtest_gitlink_is_not_a_successful_comparison, and it keeps the intent: a gitlink never yieldscompared. It now asserts apartialresult with a named gap, and text that says "not a no-change result".Static-only pins moved. Two
git.pyallowlist line numbers intests/test_adapter_static_only.pyshifted with the new lines (2872→2897, 3276→3304). The call sites themselves are unchanged.Root-scope cost. With no refusal, the root scope now materializes every blob, at one
git cat-fileper blob. That took 180 s for CubeSandbox's 3,485 files on a loaded machine. Main already paid that cost for any repository without links, plus a history walk. Batching the blob reads is a separate follow-up to the shared materializer.Docs.
docs/application-comparison.mdgains a Links and submodules section.comparednow notes that an unchanged submodule may appear inlimits.application_diffrow indocs/distribution-surfaces.mdnames the new test file.Type
Verification
CI is authoritative for
python -m ruff check .,python -m compileall -q src tests, andpython -m pytest.Additional local checks run:
tests/test_application_diff_reach.py, plus the rewritten gitlink test: red on main.git archive a430e81awith these tests copied in: 17 failed.unsupported external binding at CLAUDE.md / .claude/skills/review / agent/VERSION / .pylintrc / docs/latest / helper.py / vendor/core (160000)),not_establishedforgoogle.adk.Agent, andTypeErrorfor the new keyword.pythonpath = ["src"]overridesPYTHONPATH, so an archive of main is the only honest way to get this red run.test_application_diff{,_review,_reach,_identity}.py,test_scoped_base_tree.py,test_capability_diff_partial_clone.py,test_google_adk.py,test_host_file_links.py,test_host_link_read_through.py,test_adapter_static_only.py,test_distribution_surface_parity.py,test_public_surface_contract.py, plus Ruff. The full-suite run was at9f325c3b: Full suite (pytest -n auto -m 'not perf') had 3 failures, none from this change:test_check_unmodelled_host_config_keys[local_settings_enabled_plugins-{worktree,git_range}]fails the same way on agit archiveof main (the fixture's owngit commit -qm baseexits 1 in this environment), andtest_workflow_label_redaction::…linear_time[scheme-chars]missed its 2.0 s bound by 8 ms under-n autoload and passes serially.a430e81avs this branch, root scope:agent/VERSION, 120000)partial:cube_code_agentestablished; gap names the sibling-module toolrun_python_in_cube(#864)CLAUDE.md, 120000)partial: Python census past--max-python-files1000; two malformedmcp.jsoncandidates.claude/skills/aster-code-review, 120000)partial: new SDK agent in.agents/skills/aster-code-review/acr/…/backend.pywith dynamictools=toolstemporalio/bridge/sdk-core, 160000)partial:Submodule content is not read (unchanged commit 85b71d7ecd4f): temporalio/bridge/sdk-core; partial from unrelated test-file agents with dynamic tools and unsupported langchainCLAUDE.md, 120000; also two gitlinks)partialfrom one unsupportedmcp_server_sourcecandidate; gitlinksidlandjaeger-uinamed as unchangedThe five root-scope rows were measured at
9f325c3b, before the review fix, and were not re-run.761d5c0f. Status, rows, agent counts and limits were compared against the stored maina430e81aresults. Each difference was re-run on main6dbb445a(with fix(diff --application): key SDK identity on imports; unobserved agents are not removals #873) to attribute it.a430e81a. This includes tensorflow#128063, the one PR that was correct and complete.6dbb445a. These are fix(diff --application): key SDK identity on imports; unobserved agents are not removals #873's effects: speechmatics-academy#142, dspy#10478/#10401, dapr-agents#809, launchdarkly python-ai-sdk#102, Upsonic#1, phantom-transition#3.not_established→partial. The ADK agent is now established, and a gap names the Google ADK: resolve repository-local imported functions and module-qualified tool bindings #864 sibling tool.not_established→partial. It usesfrom google.adk import Agent. Three agents are now established, with a correctexpense_agent → register_expense: addedrow; the built-inrequest_inputis named as unresolved..cursor/rules/*.mdclink on both mains; nowpartial, from unsupported crewai/langchain files and a movedvendor/stackone-ai-nodesubmodule gap.Release-readiness notes
docs/checks.md(none)STABILITY.md(no schema change; application comparison stays0.1)Related: #864, #865, #867, #868, #871.
🤖 Generated with Claude Code