From 9b0dc02020198f7bfa2f646aa192bbf137203959 Mon Sep 17 00:00:00 2001 From: Pengfei Hu Date: Mon, 28 Sep 2026 13:52:56 -0700 Subject: [PATCH 1/2] Read a Google ADK tool a local factory builds, and a tools list built in the agent's function (#865) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit visulate/visulate-for-oracle#526 bound `save_memory_tool = create_save_memory_tool()` in its root agent's builder, where the factory returns `FunctionTool` around a nested function; the reader stopped at the local assignment. A factory call bound once and unconditionally (in the agent's function or at a module's top level), or inline in `tools=[...]`, is now followed as syntax to the factory's one unconditional `return` of `FunctionTool(inner)`, a plain function, a name bound once to one of those, or another factory's call, up to four deep; the tool is the wrapped function, with the factory recorded in `import_path`. A factory that returns from several places or under a condition, recurses, is decorated or a generator; a wrapped function that is decorated, a parameter, or changed or handed on (Visulate's delegate sets `__name__`); a tool changed where it is bound; and a factory held outside the read scope stay named with why (`factory_return`, or the resolver's reason). A third-party factory keeps the answer it had. On visulate#526 at `ai-agent`, the two memory tools become not_established candidate additions beside the nine named delegates. MuhammadVT/smart-assignment#46 built `tools = [...]` in the agent's function with a conditional `tools.append(...)`: read as a dynamic tools expression, it lost the unconditional tools. Such a list with append/extend/insert/`+=` is read member by member; an addition under a condition or in a loop is named on the agent and never read as bound; any other use keeps it dynamic. Co-Authored-By: Claude Opus 5.5 Resolve repository-local imported tools for Google ADK and OpenAI Agents SDK (#864) (#879) * feat: resolve repository-local imported tools for ADK and SDK readers (#864) A tool an agent binds from another module was unresolved at the module boundary, so a PR adding one produced no row. A shared resolver (inputs/python_imports.py) follows a tools-list reference through static imports, package re-exports, module-qualified access and plain aliases to one function definition, reading only regular .py files inside the read directory through the bounded input reader, never importing or running them. Every stop is a named reason. - Google ADK: imported names, module.function, alias = function, and FunctionTool/LongRunningFunctionTool wrappers (inline, assigned, or built in the imported module). One tool per definition; same-named functions in different modules stay distinct via binding locators. - OpenAI Agents SDK: names and module.function reaching the SDK's @function_tool, with guard evidence for imported definitions. - diff --application: rows carry import_path evidence (modules, lines, digests) outside compared meaning; unresolved references are scoped to their agent with the reason. Co-Authored-By: Claude Opus 5.5 * fix(#864): import resolution never produces a false or complete answer (review round 1) Adversarial review of the shared resolver found answers that read as complete while wrong: - One SDK agent binding two same-named definitions from different modules kept the last and reported a false CHANGED row. Both readers now bind neither, whatever the list order, and name both definitions. - A name the function building the agent binds itself (a local import, a parameter) was resolved through the module's binding. It is now a named stop (`local_binding`); a nested function that is the only definition of its name is that definition. - `import a.b` then `a.b.f` read `a/__init__`'s own `b`. It now reads the submodule, and several `import a.x` statements are not a rebinding. - An ADK wrapper warning shared by two agents scoped its gap to one of them. Every unresolved-reference record now gets its own gap. - `scan` counted one definition twice when an import reached a module another configured source also reads, making `{tool: ...}` selectors ambiguous. The catalog keeps one observation, and the binding graph reaches it through the exact definition locator the reader resolved when the edge's own source has none. - The module-binding walk climbed a parent chain per node; it is now linear (a 744 KB nested module: 32 s -> 3.7 s). Co-Authored-By: Claude Opus 5.5 * fix(#864): read a reference where it is used (review round 2) - A builder's own `from support import lookup` is followed like a module-level import (`ImportResolver.resolve_local_import`) instead of stopping, for bare names and ADK `FunctionTool(func=...)` arguments alike. - `nonlocal` follows the outer function's binding; a name a scope binds more than once is a named stop. - `tools.lookup = ...` / `setattr(tools, "lookup", ...)` in the module that binds `tools.lookup` is a named stop. - A factory's own `toolset = McpToolset(...)` / `tool = FunctionTool(...)` is read through the existing toolset and wrapper paths again, so the MCP endpoint and toolset checks return. - A module-level ADK agent binds the module-level `def`, not a nested one the flat function map happened to keep last (pre-existing on main). - An SDK list variable bound twice anywhere in the file, or changed in place, is dynamic. - The `scan` dedupe runs once where sources are loaded, so guard association and inventory completion see the same tools; a source an inventory completes keeps its observation; `./` spellings match. Co-Authored-By: Claude Opus 5.5 * fix(#864): read a list and a flat name where they are bound (review round 3) - SDK tool lists are read through the scope that binds the name at the construction. They are read only when that scope binds the name once, to a literal list, and no change site anywhere in the file (a method call, a subscript store, `global` / `nonlocal`) changes that binding. Each change site is indexed once, so the reader is linear. A module-level list's names are resolved where the list is written, not in the building function's scope (R3-1, R3-4). - A flat function, wrapper or toolset map answers a module-level reference only when its entry is the module's single top-level binding. Otherwise the resolver answers (a module import wins over a nested `def`). If the resolver cannot establish the binding, the same-named definition is named, with a tool-scoped issue, and its row is never established (R3-5). - A wrapper's `func` is read where the wrapper is written. - An enclosing package `__init__.py` that reassigns the definition is a named stop (R3-3). `_dotted` is iterative, so a chain thousands of attributes deep no longer crashes (R3-2). - The scan dedupe drops the removed copy's guard evidence too (R3-7). - The CHANGELOG states the inventory-completion ambiguity (R3-6). Co-Authored-By: Claude Opus 5.5 * fix(#864): a guess is scoped to its tool, and a list's second handle counts (review round 4) - A guessed ADK binding no longer marks the agent's list incomplete. - That had dropped every per-tool `scan` finding of the agent for one name (R4-2). `scan` is back to the round-3 behavior: a shadowed, medium-confidence definition. The PR #400 tests pass unmodified again. - The comparison reads `tool_issues` as a tool-scoped gap and an agent-scoped gap. The row is `not_established` on whichever side it is present, added and removed included (R4-1). - `x = FunctionTool(func=x)` right after `def x` wraps that `def` and is not a guess. That was noise on byte-identical files. - The SDK list reader counts `alias = TOOLS` and `TOOLS` passed to any call but a read-only builtin or logging method as a change (R4-3). - An attribute of the resolved name reassigned in an enclosing package's `__init__.py`, or in a module it imports relatively, is a named stop. This covers an alias spelled from a package above (R4-4). - Same-file symbol lookup is a dict, not a scan of every tool per reference. Co-Authored-By: Claude Opus 5.5 * fix(#864): scope a guess to its candidates; read list arguments and patches precisely (review round 5) - A guessed ADK binding's reason covers the tool it bound. It also covers every tool the module's bindings of that name could give the agent: a `def`, an import, `x = other`, or `x = FunctionTool(func=f)`. It covers the whole agent (`ANY_TOOL`) only when one of those cannot be named. - A guess that binds a name the agent already lists still leaves a gap (R5-1). - A `try: import … except ImportError: def …` fallback no longer hides the agent's other changes (R5-2). - A literal list passed to a call counts as changed unless the call is: - a read-only builtin, `pprint`, or a logging method; - an SDK `Agent` or a copy (`clone`, `replace`) reading its own `tools=`; - a function the module can resolve that leaves that parameter alone. Only names bound to a literal list are followed into a callee, so no other module is read for an unrelated call (R5-3). - The package-patch check counts only attributes rooted at an import. It parses the modules an `__init__` imports on a separate bounded budget, so a package that re-exports many modules no longer exhausts the resolution budget (R5-4, R5-5). Co-Authored-By: Claude Opus 5.5 * fix(#864): only reads keep a list literal; follow a guess's candidates (review round 6) - A literal tool list stays readable only while every use of it is a read. Allowed reads: iteration, indexing, comparison, a truth test, formatting, a read-only builtin or logging method, an agent's or copy's own `tools=`, or a function whose every use of that parameter is such a read. `+=`, a return, a tuple, `*args`, `**kwargs` or storing it in another container makes it dynamic (R6-1). - A guess's candidate tool names are followed to their definitions: an import through the resolver, or by its imported name when the module is outside the scope; `x = other` and `FunctionTool(func=f)` through the name they spell. A candidate that cannot be followed, or a wildcard import, scopes the reason to the whole agent (R6-2, R6-3). - A module that a package `__init__` imports relatively and that cannot be read (a link, a missing file) is a named stop, cached with the package (R6-4). A module read for patches is the same object a later resolution uses (R6-5). Co-Authored-By: Claude Opus 5.5 * fix(#864): spreads and tests are reads; a package's routine imports never stop (review round 7) - A literal list is still read when it is spread (`[*COMMON, x]`, `print(*TOOLS)`), tested (`TOOLS or …` in a condition), or read through a dict method (`.get`, `.keys`, `.values`, `.items`). `x or y` whose value is the list is judged by its own use. A `globals()` or `vars()` call in the module makes its lists dynamic (R7-3, R7-1 b4). - The package-patch scan treats an import above the scope as the read's boundary, as every import does. It skips an optional module imported under `except ImportError`. A link, or a missing module that is not optional, still stops (R7-2). - Docs: drop a duplicated sentence. Co-Authored-By: Claude Opus 5.5 * fix(#864): code that runs first is checked, and unread code keeps a binding named (review round 8) - Every module on a resolution chain (the agent's own file included), the `__init__.py` of every package enclosing one, and every in-scope module those import are checked for a reassignment of an attribute named like a step of the chain. `import patches` in the agent's file, or in a module the chain re-exports through, is now a named stop (R8-2). The defining module handing its function on (`registry.lookup = lookup`) does not count, and every location is kept, so one never hides another. - A relative import in any of those modules that climbs above the scope runs code that is not read. It no longer disappears: the tool is named and its row is `not_established` with that import in the reason, in both readers; for `scan` the ADK module stays at medium (R8-1, which round 7 had made silent). An import under `if TYPE_CHECKING:` never runs and is skipped. - `sys.modules`, however spelled, and importing a module by `__name__` reach its lists like `globals()` does: an SDK list there is dynamic (R8-3). - A redirecting package `__getattr__` stays a documented residual (R8-4): a gap for every hook on a chain would make #864's own attest rows `not_established`. Co-Authored-By: Claude Opus 5.5 * fix(#864): every spelling of unread code keeps the caveat; a lazy package hook is proven (review round 9) - A caveat survives a `FunctionTool(func=f)` wrapper the imported module builds (R9-1), and an absolute import spelled through a directory above the scope (`from svc.patches import ...` with scope `svc/app`) is a caveat like the relative one (R9-2). - The defining module's hand-on exemption no longer covers a patch through its own import (`import tools as _me; _me.lookup = ...`, R9-3), and only `typing`'s `TYPE_CHECKING` skips a block (R9-4). - Ordinary code no longer stops a resolution (R9-5): every location an ambiguous import could mean is read, a generated `*_pb2` module is the boundary and any other missing relative module a caveat, and the patch scan reads up to 1024 modules. - A package `__getattr__` is established only when every return it can reach for the name gives that submodule (`import_module(f".{name}", __name__)`, `from . import name`); attest's two lazy loaders stay established, a redirecting hook is named (R8-4). Co-Authored-By: Claude Opus 5.5 * fix(#864): the repository says what is its own code; harden the lazy-hook idiom (review round 10) - Whether an absolute import no file in the scope provides is the application's own code is read from the repository: the compared commit's tree for `diff --application` (set per side), the checkout for `scan`. A module or regular package at the root or under `src/` is; a directory without `__init__.py` only when it holds the named submodule. `from common.patches import ...` with scope `svc/app` is a caveat (R10-1), and SDK apps under `agents/` import the SDK again (R10-3: the round-9 ancestor-name rule had caveated every tool there). - The scope spelled from the repository root (`svc.app.tools` with scope `svc/app`) is read inside the scope, so it resolves and a patch module it names is checked instead of guessed. - The lazy-hook idiom requires an undecorated hook that never rebinds its parameter, `importlib` / `import_module` bound only by importing them, and a returned local bound exactly once, counting `for`, `with`, walrus and `except` targets (R10-2). - A generated `_version` module is the boundary like `*_pb2` (R10-4); the patch-scan budget message names what it scans. Co-Authored-By: Claude Opus 5.5 * fix(#864): every directory above the scope is an import root; unread links and subverted hooks are named (review round 11) - An absolute import is looked up at the repository root, under `src/`, and under every directory between the root and the scope, so `backend/app` importing `common` from `backend/common` is a caveat (R11-1). - A root entry that is a symbolic link or a submodule, spelled by the import, is unread code: a caveat (R11-2). `scan` outside a checkout reads the three directories above the scope instead of nothing (R11-3). - A store into `sys.modules` in code that runs first is a named stop, and a package hook is not trusted when the package rebinds `__name__`, patches `importlib`, or stores into `sys.modules` or `globals()` other than the idiom's own cache (R11-4). - The scope spelled from an import root is read inside the scope only through a regular package (R11-5). Co-Authored-By: Claude Opus 5.5 * fix(#864): a package is not an import root; a sys.modules store is read by its key (review round 12) - A directory between the repository root and the scope that holds an `__init__.py` is imported through its parent, never from the path: SDK apps under a regular `app/agents/` package import the SDK again, and `app/types.py` no longer shadows the standard library (R12-1, a round-11 regression). - A store into `sys.modules` (`[...] =`, `setdefault`, `__setitem__`, `update`) is a named stop when its key names a module on the chain or is built on `__name__`, a caveat when it is computed (a plugin loader's `spec.name`), and nothing when it names another module (R12-2, R12-3). - `globals().update(...)` / `setdefault` in a package disqualifies its hook (R12-3). Co-Authored-By: Claude Opus 5.5 * fix(#864): packages above the scope are read; sys.modules and globals() by allow-list (review round 13) - Every directory between the repository root and the scope is an import root again, package or not: a service run from `backend/` with a stray `backend/__init__.py` names `common.patches` (R13-1, a round-12 regression). A standard-library name, and the scope's own package on the way to it unless it holds the name imported, are not the repository's code, so SDK apps under `app/agents/` still import the SDK (R13-5). - The `__init__.py` of every package above the scope, and each module it imports, are read through the repository layout (the commit's tree, or the checkout) for the same reassignments (R13-2). - `sys.modules` and `globals()` are read by allow-list. A store whose key names a module on the chain, a package above one, or the framework's own modules is a named stop; `__name__` plus a literal is that module's own name; any other use is a caveat; a module rebinding its own name through them is a reassignment; a change to `__path__` is a caveat (R13-3, R13-4, R13-6). A package that rebinds `__getattr__` or `__path__` does not have a trusted hook. Co-Authored-By: Claude Opus 5.5 * Read what the packages above the scope import, and every spelling of a module's own object (#879 review, round 14) - Above the scope, follow each import the way the in-scope reader does: every package on the way to the module and each submodule named (`from .hooks import patches`, `import svc.lib.util`). A reassignment there counts only when rooted at the scope or at something that cannot be found; guarded and generated imports are exempt. - A standard-library name is exempt only at a root that is a regular package; interpreter-preloaded modules are always exempt. - Allow-list reads: comparisons, spreads, iteration, pkgutil.iter_modules(__path__), namespace keywords (get_type_hints(globalns=globals())), and patch.dict/setitem by key. - The module's own object through an alias, sys.modules.get, import_module(__name__), __dict__/vars() stores, a computed setattr, or handed to a function is a reassignment or caveat, never silent. Co-Authored-By: Claude Opus 5.5 * A local named vars or globals is a variable, not the builtin namespace (#879 review, round 14 corpus) siada-cli's DebugUtils.dump binds `vars = stack[-2][-3]` and iterates it; the bare-name rule read it as the builtin handed on and named a real definition change not_established. Co-Authored-By: Claude Opus 5.5 * Only a two-part patch on another module is exempt above the scope; read the scope's modules an ancestor imports and the module object anywhere (#879 review, round 15) - R15-1: above the scope, a patch is exempt only when it sets one attribute of another module file outside the scope that its package binds nothing else under. A longer path, a name imported from a module, a package attribute shadowing the submodule, or a module alias (`_t = tools`) counts, in both readers. - R15-2: a scope module that an ancestor __init__.py imports is read like the chain's own. - R15-3: a module file wins over a same-named directory without __init__.py. - R15-4: a link's blob text is never read as source. - R15-5: the module object used anywhere but an attribute, a plain alias, a comparison or a reader is a caveat. So are __dict__/vars().update, getattr(m, "__dict__") stores, f_globals, builtins.globals under another name, a reader of the module's own under a reader's name, `globals = globals`, and __path__ changes above the scope (pkgutil.extend_path aside). get_type_hints(fn, globals()) is a read. - R15-6: above-scope files are read in batched `cat-file --batch` per directory (1100 modules: 76 s -> 3 s). The read bound is named once. Co-Authored-By: Claude Opus 5.5 * Read a frame's namespace by allow-list; batch a directory only while it is small (#879 review, round 16) - R16-1: `sys._getframe(1).f_globals.get("__name__")` in a logging helper is a read. A frame's namespace is read by the same allow-list as globals(); any other use is a caveat. - R16-2: above-scope reads batch a directory's modules in one `cat-file --batch` only while the directory holds at most 16 MB of Python. A directory of generated or vendored modules is read file by file, as asked (peak RSS 305 MB -> 87 MB on the 60 MB case). - Docs: `sys.path` / `sys.meta_path` changes that make a same-named module elsewhere the one imported are not followed. Co-Authored-By: Claude Opus 5.5 * Bind the located definition, prove readers by their binding, check the lazy hook's import (#879 PR review) - The resolved locator decides between same-named definitions even when the edge's own source holds exactly one. After source deduplication, a `tools.lookup` imported as `shared_lookup` no longer binds the local `lookup`. - A read-only reader is proven by its binding: - a bare builtin only when nothing binds it and no star import may; - a standard-library reader only when imported from that module (`json.dumps`, `from pprint import pprint`); - a logging method only on `logging` or a `getLogger()` logger. A `print` imported from the application, or an `.info()` on its own object, makes the list dynamic. - A lazy `__getattr__` is the submodule idiom only when its alias imports the submodule asked for: `from . import alternate as memory` is not. perf: read materialized Git tree blobs through batched cat-file (#878) _materialize_isolated_tree spawned one `git cat-file blob ` per file in both its entries and links loops. Once `diff --application --scope .` materializes the whole tree (#877), that dominated the run: one side of TencentCloud/CubeSandbox took ~180 s to archive (the #686 cost class). Blobs are now read by _isolated_blobs: one `cat-file --batch-check` types and sizes every object (a missing or non-blob object refuses with a ConfigError before any content is read), then `cat-file --batch` runs of at most 64 MiB each return the content, with strict framing checks. Only full object IDs reach the batch. The link-text reads in _scope_through_boundary_links use the same reader. Every blob is still hashed against its tree entry's oid before it is written; path handling, containment and escape checks, blobs-before-links order, placeholders and the final digest comparison are unchanged. The reads go through _run_git_dir (now accepting input=) and the existing _run_process boundary, so no subprocess call site is added; only the two line pins in test_adapter_static_only.py moved. diff --application --scope . with #877: CubeSandbox 412 s -> 46 s, dlt 203 s -> 45 s, identical rows. fix: compare application wiring past unrelated links and submodules; read google.adk.Agent (#877) * fix: compare application wiring past unrelated links and submodules `diff --application` at the root scope took the unscoped archive route, which refuses every symlink and gitlink. On 2026-09-25, 15 of the first 27 runs against open third-party SDK/ADK PRs exited 2 on a path the reader never opens: `CLAUDE.md -> AGENTS.md`, a linked skill directory, a `VERSION` link leaving the tree, a vendored submodule. - Materialize every scope, the root included, through the scoped verified materializer, which recreates links rather than refusing them and packs the tree instead of the history. - Never read a Python input through a link: a target this scope already reads is compared at its own path, and any other in-scope target is a gap over the link's path. Reading the alias made one agent two ambiguous ones and hid the real file's change. - Record gitlinks behind an opt-in `archive_tree(record_gitlinks=True)`, materialized as the empty directory an unpopulated checkout leaves. An unchanged gitlink commit is named in limits; any other is a coverage gap over its path. Other archive callers still refuse. - Stop turning the host-configuration census's link count into application coverage gaps. - Read `google.adk.Agent`, the package-root re-export, as an ADK agent constructor (`from google.adk import Agent`). Co-Authored-By: Claude Opus 5.5 * fix: keep unread links and scope-level submodules from reading as removals PR #877 review found two cases where unread input became conclusive negative evidence: - Discovery drops a path that does not resolve, so replacing `agent.py` with a dangling link read as a definite removal, `compared`, with no gap. The dropped host-census gap had been the only fallback, and it also covered a source directory replaced by an absolute or dangling link. Census every link under the scope directly, without following it: gap a `*.py` link that does not alias an input the scope already reads, a directory link holding Python outside the scope, and a link resolving to nothing in the tree where the other side reads source at or beneath it. An unchanged `agent/VERSION -> ../../VERSION` still changes nothing. - A gitlink at the selected scope itself gave a gap with source ".", which covers no relative binding path, so `agent.py` read as a definite addition or removal. Make that gap scope-wide. fix(diff --application): key SDK identity on imports; unobserved agents are not removals (#873) * fix(diff --application): key SDK identity on imports; unobserved agents are not removals On speechmatics/speechmatics-academy#142, a LiveKit voice agent moved its tools into an Agent subclass that passes them through super().__init__. `diff --application` printed `compared` with four false REMOVED rows. Two defects: - Framework identity. Discovery counted any bare `@function_tool` as the OpenAI Agents SDK, and the SDK reader recognized `function_tool` and `Agent` by spelling alone, so `livekit.agents` symbols were read as the SDK's. Both now key on import provenance: the absolute `agents` / `openai_agents` package, or an unimported name (unless a foreign wildcard could supply it). Relative imports are no framework's signal. - Unobserved is not removed. An agent observed on one side whose file on the other side still assigns its name (or passes it as `name=`), via a construction no reader supports (subclass, factory, Agent[Ctx], clone), now records a scoped coverage gap for that agent: `partial`, rows `not_established`. Handoff-only references do not count as an observed construction. A genuinely deleted agent stays an established removal. Regression tests fail on the unfixed tree. Co-Authored-By: Claude Opus 5.5 * fix(diff --application): resolve SDK names per scope; imports keep an agent named Addresses the two P2 findings on #873. - Import provenance was flattened across lexical scopes: one file-wide import map, accepted if any origin was the SDK, so a LiveKit `Builder(...)` became an SDK agent when a sibling function imported the SDK under the same alias. `_SdkNames` now resolves each spelling in the scope that uses it (nearest binding scope, class bodies skipped for nested code, global/nonlocal obeyed; decorators and defaults in the enclosing scope). An import there must be the SDK's; a parameter or local assignment is not. `function_tool` is decided per decorator node, not as a file-wide set of spellings. The walk is iterative. - The missing-agent check ignored import bindings, so `from agent_factory import agent` (or `exported as agent`) still gave a definite removal. Import aliases now count as the file still naming the agent. Regression tests for both fail on the previous head. feat: compare application agent wiring without prior setup (#871) * feat: compare application agent wiring without prior setup * fix: address application comparison review findings * test: normalize colored CLI errors in review regressions fix(ci): prevent checkout module shadowing in non-GitHub recipes (#870) * fix(ci): prevent checkout module shadowing in installation recipes * test(ci): cover quoted and qualified Python installation commands Read agent launches in CI workflows (#823) (#850) * Read agent launches in CI workflows (#823) The workflow grant read triggers, token permissions, reusable-workflow secrets (#693) and step `uses:` references (#771), and nothing about how a coding agent is launched inside a job. Changing a claude-code-action's `claude_args` from `--allowedTools "Read"` to `--permission-mode bypassPermissions --allowedTools "Bash(*)"`, adding a `claude -p --permission-mode acceptEdits` run step, or checking out the pull request head in a `pull_request_target` job each gave "No static host-grant changes detected", and a move to `issue_comment` with `pull-requests: write` gave a row with no agent context. The workflow grant now lists, as text that is never executed, fetched or evaluated: - `agent_launches[]`: a step whose `uses:` is anthropics/claude-code-action, anthropics/claude-code-base-action or openai/codex-action (any ref, any case) with the documented permission inputs it sets, or a `run:` that is one literal simple command starting with `claude -p/--print` or `codex exec`, with its documented permission flags under their primary spelling. The prompt, `--model` and undocumented flags are not compared. `job_secrets` names the secrets the job references, as context only. - `checkout_refs[]`: each actions/checkout step's `with.ref`, null for the default. Each job's multiset of launches and refs is compared, never the step label, so a rename or reorder is quiet; a difference is one `changed` row on the existing workflow row naming `job/step` and both values. Direction is claimed only by documented rules a job's launches gain, read from literal values: bypassed permission checks (either spelling), bypassed approvals and sandbox, a danger-full-access sandbox, `safety-strategy: unsafe`, or a `*` user gate. Those raise `workflow_agent_widened_` and make the row widened; every other edit, including `--allowedTools "Bash(*)"` (#824's to rate), is `changed`. A workflow row whose workflow runs an agent ends its `why` with the untrusted-input trigger, write scopes, secrets and pull request checkout beside each agent step; it is a note and moves no direction. A compound `run:`, an expansion or an expression is `unresolved`, publishes none of its text and records a non-blocking coverage issue naming `job/step`, as an unread secret value does (#693): coverage stays complete, adding one is a row that claims no effect, and an edit inside one is not reported. Scripts, composite (#701) and unknown actions, and agents reached through npx/timeout/sudo are listed as unread surfaces. Every value, ref and secret name goes through the #802 label redaction; a rewritten one is null with `redacted` and is neither published nor compared. Contract 40 and host-grants 0.6 shipped in 1.1.0, so host-grants inventory, baseline and drift move to 0.7 (the 0.6 schema files are untouched) and the runtime contract to 41. A 0.4-0.6 baseline holding a workflow grant is incomparable (`baseline_workflow_agent_launches_unavailable`); one without a workflow stays comparable. Verifier 0.20 and capability diff 0.3 do not move. No check id is added or removed and `check` decides as before. Docs: the support page (tables, rules, note, limits, a new unread-surfaces bullet), a STABILITY "Migration Note: Unreleased", CHANGELOG `## Unreleased` above 1.1.0, the agent contract page, the distribution-surfaces `capability_diff` row and its parity comment, the Stop hook note, version tables and pins, and a rebuilt llms-full.txt. The pilot ledger's source-tree column was re-measured on the Route H fixture: identical cells to the 1.1.0 engine apart from contract 41 and inventory 0.7, with byte-identical diff rows. Tests: new tests/test_workflow_agent_launches.py covers the four reproduction cases, what is and is not read, every unsupported shape, direction rules and non-rules, the acceptance's negative controls, redaction with a CLI canary sweep across every published output, the 0.6 baseline migration, schema validation, and the same row on diff, verify, the PR comment, check, the control envelope and the Stop hook. Existing tests move to 0.7/41. The host-config and cold-start replays reproduce their committed outcomes and run-of-record scores, and the sample goldens are unchanged. Closes #823 * Read no widening rule from an agent input that holds an expression (#823) The support page and STABILITY say a documented rule is read from literal values only, so a value holding `${{ }}` never meets one. The `*` user gate did not apply that: `allowed_non_write_users: "${{ vars.USERS }}, *"` raised workflow_agent_widened_changed. Every rule now skips a value holding an expression, as the claude_args and codex-args rules already did, and a test pins the gate case. * Address review cycle 1 on agent launches in CI (#823) F1. `claude_args` and `codex-args` were split with the `run:` shell tokenizer, which gave up on a newline or an unquoted `(`, so a widening written the way the actions document it gave `changed`: `claude_args: |` on several lines, `--allowedTools Bash(git:*) --dangerously-skip-permissions`, a `# comment` line, or a multi-line `codex-args`. Each input is now split the way its action splits it. The Claude actions' parse-sdk-options.ts drops full `#` lines, makes `()|&;<>` literal and reads the rest with shell-quote (newlines are whitespace, `$NAME` is empty, an unquoted `#` ends the input, a `--` word is always a flag); codex-action reads a JSON array of strings or string-argv. The `run:` tokenizer is kept for `run:` steps only. The published `claude_args` is the text the action parses, so a full-line comment is neither published nor compared. F2. Any URL path made a whole setting `redacted`, so a `--dangerously-skip-permissions` beside `https://example.com/style-guide` gave no row, and every `plugin_marketplaces` value was never compared, under a limit that wrongly called it credential-shaped. The documented rules are now decided from the declared text when the workflow is read, before anything is withheld, and published on the launch as `widening_rules` (rule and setting), which the comparator keys on, so redaction never hides a rule. A URL publishes its scheme and host with ``, as an MCP server URL does (#723), and the rest of the value is published and compared. Other credential-shaped text in a setting or checkout ref (a token shape, an assignment, a bearer or header value, URL userinfo) is published redacted and makes the workflow a blocking limit through `_uncompared_workflow_text`, as a redacted step reference does (#767). F3. JSON in `settings`, `mcp_config`, `--settings`, `--mcp-config` or any argument word was published verbatim, env values and apiKeyHelper included. A JSON object now publishes what `.claude/settings.json` and `.mcp.json` publish: key names, with `env`/`headers` values, apiKeyHelper and secret-named values ``, as canonical JSON. A codex `--config` override under env, headers or a secret-named key publishes ``. Text that starts like JSON and does not parse is withheld (`unparsed_json`, a non-blocking limit). The agent-launch canary sweep now carries JSON-shaped canaries and their digests across diff, audit, check, verify, the PR comment and verifier.json, and a second sweep covers the refused credential-shaped case. Nonblocking: a rule gained where the job launched that agent before only in an unread form is named and not claimed (the `unknown_before` rule); a quoted word starting with `#` no longer makes a `run:` compound; `anthropics/claude-code-action/base-action` is read as the base action; the STABILITY note says a prompt after a variadic flag is compared; and `__all__ =[` is spaced. Host-grants 0.7 is unreleased, so `widening_rules`, `unparsed_json` and the base-action agent value extend it in place; the 0.7 schema files are regenerated. The support page, STABILITY migration note, CHANGELOG Unreleased entry, contract summary, integrations Stop-hook sentence and the capability_diff distribution-surface row say the same. * Address review cycle 1 on agent launches in CI (#823) The second review of #850 at 25c13ce8 found two P1 and four P2 defects. The branch is rebased onto origin/main daa4ad5f. Attached JSON values are withheld (F1). A --settings={...} or --mcp-config={...} word inside claude_args, and codex's attached -c / -c=, published their env, header and apiKeyHelper values verbatim, because only a word starting with "{" was withheld. _withheld_words now splits a --name=value word and withholds the value, and reads codex's attached -c through _withheld_config, as clap reads it. The canary sweep carries both spellings. Redacted prose no longer refuses the comparison (F2). The #802 label redaction rewrites ordinary prose ("never print bearer tokens", "Authorization: headers"), and a redacted agent setting made the workflow a blocking limit. That hid every row beside it and made every check incomparable while the workflow existed. The rules are already read from the declared text, so a redacted setting is now compared by its published text and its widening_rules. uncompared_agent_launch_texts names it as a non-blocking limit, and the row cell shows the redacted text. A redacted checkout ref still refuses, as a redacted step reference does (#767), because it names the code a job runs. A renamed job's launch moves its rules (F3). _agent_rule_gains keyed a rule on its job, so renaming a job that launches a bypassing agent was a widening. agent_rule_gains now pairs a rule one job gains with the same rule another job lost, when the launch that met it left that job: the job no longer launches that agent, or the same launch (agent_launch_key less the job) now runs in the gaining job. The why names the move. A second job gaining a rule, or a different launch gaining one while the first job still launches that agent, still widens. An expression no longer turns off every rule, and the row says what it leaves unread (F4). Rules are read from literal text a ${{ }} expression cannot reach: - the words of claude_args or codex-args before the first expression, less the word it touches and any quoted run still open at it; - the elements of a JSON-array codex-args before the one holding it; - the gate entries that hold none. A setting holding an expression is published with holds_expression (the unreleased 0.7 schema extends in place), and a row that changes it says the text the expression reaches is not read. A gain where the job's launch held an expression before, in the input the rule is read from, is named and not claimed, as unknown_before is. That also fixes a false widening at 25c13ce8: replacing --model ${{ vars.CLAUDE_MODEL }} with --model opus beside --dangerously-skip-permissions was reported as gaining the bypass. Two non-blocking fixes: - _published_value reads each expression as one word, so an expression in a URL's userinfo is withheld with it rather than garbling the URL and publishing its path. - A codex --config value that starts like a table and does not parse, as string-argv leaves a quoted one, is withheld as unparsed_json. Rebase (F5). CHANGELOG keeps #853's #778 line beside #823's under github_action row. llms.txt and ai-search-summary state the source tree as contract 41, unreleased, ahead of the published v1.1.0 (contract 40), as test_public_surface_contract requires while the two differ. The pilot ledger's source-tree column was re-taken on the rebased tree, through ./shipgate beside the engine of e3c6cb0c, the commit v1.1.0 was cut from: - the only differences are contract 40 -> 41 and inventory schema 0.6 -> 0.7; - diff rows are byte-identical; - check JSON differs only in the launcher path its next action names. The support page, the STABILITY preamble and migration note, the CHANGELOG entry, agent-contract-current, the Stop hook sentence in integrations, the capability_diff row in distribution-surfaces with its parity comment, the 0.7 schema files and llms-full.txt are updated to match. * Address review cycle 2 on agent launches in CI (#823) The third review of #850 at c7f67550 found one P1 and one P2 defect and four P3 notes. The branch is on origin/main 44b9e05d; no rebase was needed. A JSON setting publishes its shape, not its free text (C2-F1). _withheld_json published the whole _redact_secret_values tree. That tree is the host readers' digest input, not what they publish: it keeps every string outside env, headers and secret-named keys. So an mcp-remote --header "Authorization: Bearer ..." argument in --mcp-config, and a hook's curl command in settings, reached diff text, diff --json, audit --host --json, the PR comment and verifier.json. _json_shape now keeps key names, numbers, booleans and null, redacts what the host readers redact, and replaces each other string with , a 12-hex digest of redacted_config_sha256 for that string, so an edit to it is still a changed row. It keeps only the strings a host reader publishes: - a permissions.allow/ask/deny rule and a documented Claude Code setting's value (defaultMode, the switches, enabledMcpjsonServers entries); - an MCP server's command name and its URL's scheme and host, followed by the digest when they drop a command's arguments or a URL's query. A codex --config table or array is read under its key path, so mcp_servers.gh={command="gh", ...} keeps its command name. The canary sweep adds the mcp-remote header and the hook command in every spelling (action input, claude_args, CLI flag, codex -c table). Re-running the reviewer's two repositories through ./shipgate gives 0 canaries in every output. The documented bypasses written through listed inputs widen (C2-F2). - Claude Code settings written as JSON meet bypass_permissions when their defaultMode is bypassPermissions, read by claude_setting_values as the settings reader reads .claude/settings.json. This covers the action's settings input, a --settings value in claude_args, and the CLI's --settings flag. A path is not read, and a settings value holding an expression meets none. - openai/codex-action's permission-profile: :danger-full-access meets danger_full_access. ":danger-full-access" is Codex's reserved name for its built-in full-access profile (BUILT_IN_PERMISSION_PROFILE_DANGER_FULL_ACCESS). Mode inputs are now (input, value, rule) triples. settings and permission-profile join _RULE_SETTINGS, so a gain after an expression in either is named and not claimed, as for claude_args. P3 notes: - allowed_bots: "*" now reads "accepts runs triggered by any bot" through agent_rule_text. - Which of two gaining jobs a moved rule goes to no longer depends on declaration order: the same launch arriving is matched before a job that merely stopped launching the agent. The support page, the STABILITY migration note, the contract summary, the schema docstrings and the CHANGELOG entry now say what a structured value publishes instead of claiming it publishes what the host readers would, and list the two rules. host-grants 0.7 is extended in place; it is unreleased. * Address review cycle 3 on agent launches in CI (#823) Rebased onto origin/main 01777037 (#852, #821), which had already moved the unreleased runtime contract 40 -> 41, verifier 0.20 -> 0.21 and capability diff 0.3 -> 0.4. This change now extends contract 41 in place instead of minting it: one comment in schemas/contract.py names #821 and #823, and the texts that said verifier 0.20 and capability diff 0.3 were unchanged now say capability_diff registry row keeps #852's two added roots and both paragraphs; STABILITY keeps the #821, #823 and #827 notes, with #821's "host-grants stays 0.6" and #827's version sentence corrected for a tree that also carries host-grants 0.7; CHANGELOG keeps both Unreleased entries. llms-full.txt is rebuilt. The pilot ledger's source-tree column was re-measured on the rebased tree beside e3c6cb0c: identical cells except contract 40/41 and inventory schema 0.6/0.7, identical diff text, rows, check JSON and inventory apart from its schema version, and diff --json / verifier.json differing only in #821's schema versions and coverage members. A codex exec step that selects the full-access sandbox through --config was a changed row while --sandbox danger-full-access widened, and -sdanger-full-access was not read at all. The flag reader now reads a short flag's attached value as clap does (-s, -s=, -c, -c=), and a --config override that sets sandbox_mode to danger-full-access, or default_permissions to :danger-full-access (what codex-action's permission-profile input passes the CLI), meets danger_full_access, its value read as the CLI's parse_overrides reads it. As the CLI resolves them, the last override of a key counts, default_permissions outranks sandbox_mode, and a --sandbox flag outranks both, so an override beside --sandbox meets none. In codex-args such an override meets none, because the action appends its own --sandbox or default_permissions selection after codex-args. The support page and the STABILITY Direction bullet list the spellings. Withholding a quoted URL inside a JSON-array codex-args element no longer drops the element's escaped closing quote, so the published value stays the JSON array the support page describes. * Address review cycle 2 on agent launches in CI (#823) An agent CLI inside a double-quoted $(...) or a backtick substitution, or after a shell reserved word, got no launch, no row and no coverage issue: gh pr comment --body "$(claude -p --dangerously-skip-permissions ...)" read as "No static host-grant changes detected", while the support page said such a run: is listed as unresolved once for each agent CLI it starts at the head of a command. The word splitter reads a double-quoted substitution as one word and a backtick as no operator, and the head reader took then, do, { and ! for the command. _command_substitutions now lists the text of each $(...) and backtick substitution outside single quotes, double-quoted ones included, and an agent CLI heading a command inside one is an unresolved launch (shell_expansion for a single top-level command, compound_command otherwise), publishing none of its text. An escaped \$, a single-quoted '$(...)', a comment and $((...)) arithmetic are not substitutions, so echo "claude -p ..." stays unlisted. The scanner keeps an explicit stack and reads each nested substitution as `_` in the one around it, so no nesting depth recurses or re-splits text. When the run's quoting does not balance, each line's substitutions are read as well. A command's head is now read after shell reserved words (!, time [-p], {, if, then, elif, else, while, until, do, function NAME), and a command after one is not a simple command, so its launch is compound_command and never read. The compound_command and shell_expansion limit phrases, the launch schema description (host-grants 0.7 schemas regenerated), the support page and the STABILITY bullet say so. A read launch that became a form this audit does not recognise (npx, a path, codex options before exec) was worded "a step no longer launches an agent", though the step still starts one. The removed case now reads "a step no longer declares an agent launch this audit reads", and a row whose workflow still exists adds that the step may still start an agent in a way this audit does not read, naming those forms. codex with an option before exec is listed under Known unread surfaces and in STABILITY's "What is not read". Tests cover each shape the review named, the negative controls, the diff and audit --host route of the reproduction, the reworded removal for npx, a path and codex root options, and 20000 nested substitutions. * Address review cycle 3 on agent launches in CI (#823) A step that drafts a PR comment in a quoted here-doc, such as cat > comment.md <<'EOF' with "Reproduce locally with `claude -p ...`" in its body, was listed as an unresolved agent launch: a "high changed" row saying a step now launches an agent, and a GitHub coverage issue. The shell passes a here-doc's body to its command as input and, with a quoted delimiter, expands nothing in it, so the step starts no agent. The phantom also hid a real gain: a --dangerously-skip-permissions step added beside it was "not counted as a widening" because the job "launched the agent in a form this audit does not read" before. The substitution scanner read backticks and $( across the whole run:, and the word splitter read a body line starting with claude -p as a command head. _here_documents now removes each here-doc's body and closing line before the run: is split or scanned. <<, <<- and a quoted, escaped or partly quoted delimiter are read outside quotes, comments and $((...)) arithmetic, inside a $(...) or backtick substitution too; <<< is a here-string; bodies start after the opening line and follow in order when one line opens several. A body is never a command. Only an unquoted < ../.agents/skills`, a per-skill link or a file link - the same untouched file refused the whole comparison: diff, verify and the manifest-free PR comment printed `base_inventory_incomplete; head_inventory_incomplete` and no row, hiding a removed deny rule beside it. `unchanged_limits` asked blob_path_unchanged whether the source, the path the link is read under, was one regular file in Git, and a path through a link never is. blob_path_unchanged now resolves the path on each side the way the reader reaches it: from Git tree entries for the base and a commit head, and for a working-tree head without following any link (each component's own entry, each link's own text, the file's unfiltered hash). It holds only when both resolutions are equal - every link at the same path with the same text, every other component a directory - and the file they land on is the same regular-file blob at the same in-tree path. A link is followed only under the rules the reader and the base archive already use (#700, #711): a relative text that lands inside the tree after normalization, directories above where it lands, at most eight links, and links at one component of the path. A path no link reaches still takes the one `ls-tree` it took before, and blob object IDs are still compared, so no filter or textconv can make two byte sequences equal. Everything that asks the proof moves together: unchanged_limits in `diff --json` and verifier.json, the text and PR comment, the unchanged limits a partial comparison may carry (#808) and the shared plugin-reference limits check leaves out (#714). check's boundary result still cannot name a limit, so it now refuses these comparisons with unchanged_limits_not_representable, as it does for a limit at its own path; its rows, decision and violations do not move. No schema, member, reason code or check id is added. The metadata value is not coerced. tests/test_linked_unchanged_limits.py holds every layout (directory link, per-skill link, file link, a chain of file links) to the direct result on diff, verify, the PR comment and check; keeps the negative controls refused (skill added or edited behind the link, link retargeted to an identical copy, link text rewritten to land on the same file, link replaced by a directory or the reverse, a later hop retargeted, working-tree-only retargets and edits); and holds the proof to exactly the links the reader reads through, across dangling, looping, absolute, escaping, over-long, intermediate and nested links. The #812 coverage case that pinned the refusal is retired, and the static-only allowlist follows its two call sites down the file. * Address review cycle 1 on linked unchanged limits (#822) The STABILITY migration note said that where the other routes now compare past a limit reached through an in-tree link, `check` refuses with `unchanged_limits_not_representable` and its rows stay empty. That holds for an instruction file's limit, which `check` cannot leave out, but not for a plugin-reference limit: `_without_shared_plugin_reference_limits` asks the same unchanged proof, so a `parse_failed` `plugin.json` that is a file link to an unchanged target is now left out exactly as one at its own path has been since #714. `check` then compares and publishes the removed `deny` row in its boundary result and its control envelope's `capability_rows`, where it refused with `base_inventory_incomplete` / `head_inventory_incomplete` and no row, and `diff`, `verify` and the PR comment, which withheld `plugins/demo` as `partial` (#808), are `comparable` with the limit in `unchanged_limits`. `check`'s decision, violations and control state, and `verify`'s control state and next action, do not move. The note now splits `check` by whether it may leave the limit out, adds the `partial` to `comparable` move, and says "every row's value" where it said "every row". The CHANGELOG entry mirrors it, the distribution-surfaces row and `docs/host-boundary-support.md` no longer say any change behind the link refuses the comparison (an edited plugin manifest keeps it `partial`), and `tests/test_linked_unchanged_limits.py` pins the plugin manifest at its own path and behind a file link, on every route, with the edited-target control. The `blob_path_unchanged` and `_reader_path` docstrings now state the proof as necessary, not sufficient: it follows the links on the way to the path, not every condition the reader puts on reading a whole linked directory, and a side whose reader does not read the path carries no limit there. A test holds a head that adds a link inside the linked directory to a refusal on every route. The two pinned call-site lines in `cli/verify/git.py` move with the docstrings. * Address review cycle 4 on agent launches in CI (#823) Four review cycles each found a new shell form (quoted words, $(...), here-docs, comments, reserved words) that the run: reader mis-read. The cycle-4 finding was the fourth: a `#` comment was split as commands, so an apostrophe in a comment hid a launch and a command line in a comment invented one. Parsing arbitrary shell cannot converge, so the reader now claims only forms it reads exactly and names every other one as a limit. - A run: step is an agent launch only when it is one line of plain words (letters, digits and `_ . / : = , % + -`, separated by spaces or tabs), run by bash, sh or no declared shell:, whose program's file name is `claude` with -p/--print or `codex` followed by exec. Every POSIX shell runs such text as exactly those words; a cross-check of 2251 generated texts in sh, bash, dash and ksh agrees on every one. - Any other run: that mentions claude or codex as a word of its own is an `unread_agent_runs[]` entry (job, step, agent) on the workflow grant: a non-blocking coverage issue in audit --host that publishes none of its text, is never compared, so it gives no row, and never says whether the step starts an agent. It takes no gain from another launch; only when an unread step goes and a read launch is added in the same job is a rule that launch meets named and not claimed, since it may be that step rewritten. - claude_args and codex-args are read only as a plain list of words (the same characters and parentheses, across blanks and newlines, with no --settings or --mcp-config flag). Any other value, a ${{ }} expression included, is `unread_arguments`: published only as a digest, so an edit is a changed row, and read for no rule; a rule a launch gains where that input was unread before is named and not claimed. - A codex --config override publishes its key; its value is under env, headers or a secret-named key, as written for sandbox_mode, default_permissions, approval_policy and model, and a digest otherwise. The word after a secret-named word such as --token is and the value is then published redacted, a named limit. - The shell tokenizer, the command-substitution and here-doc scanners, the reserved-word reader and the shell-quote and string-argv emulations are removed, with the expression-prefix reading of argument inputs. JSON is read only in the settings and mcp_config inputs. Host-grants 0.7 is unreleased, so its schema changes in place: the launch `unresolved_reason` is only inputs_not_a_mapping, a setting may be unread_arguments, and the workflow grant adds unread_agent_runs. The support page, STABILITY, CHANGELOG, the current contract page and llms-full.txt describe the tightened reader. * Address review cycle 5 on agent launches in CI (#823) C5-F1: the move rule took a launch that only stopped being read to have left its job. With job a running `claude -p --dangerously-skip-permissions Review` in the base, quoting its prompt (an unread step) or running it through `npx` while job b added the same plain launch gave a `changed` row saying the launch "moved between jobs (a → b)" and "already met that rule in the job it left", with no widening signal. Job a still runs it. - `same_launch` now refuses a move while the losing job may still run the launch in a form this reader does not read: an unread step of that agent stands at a step label the lost launch held, the job has more unread steps of that agent than before (the launch may have moved to another index), or one of its read launches of that agent holds an expression or an unread argument input in an input the rule is read from. The review's M5 and M6 are now `widened` and name "a step no longer declares an agent launch this audit reads (a/steps[0])"; a real move, a rename, a swap, and a move beside an unread step the job already had stay moves. C5-F2: docs/integrations.md said an unread `run:` is a non-widening row that diff and the PR comment show. It gives no row; only an unread argument input is a row. The sentence is split accordingly. Nonblocking items fixed: - `_job_secrets` walks each container once, without recursion. A job `env` holding itself through a YAML alias raised RecursionError, and a ten-way fan-out eight levels deep did not finish; both now read in well under a second. - A `settings` or `mcp_config` value that is neither a JSON object nor a plain file path (path characters, and a `${{ }}` only as a plain context reference) publishes only a `` digest. A comment line before JSON published the `env` value that JSON held. - The last of a repeated `--permission-mode` counts, in `claude_args` and in a plain `run:`, as parse-sdk-options and the CLI keep it. - The docs/distribution-surfaces.md capability_diff row no longer says an unread argument input or an unresolved launch is named only by the inventory and audit --host: a row reporting its launch names it too. The support page, STABILITY (Direction and What is withheld) and the CHANGELOG entry say the same. Tests: tests/test_workflow_agent_launches.py grows to 290 cases. The five new move-rule widening cases, the alias test, the two json-or-path cases and the four --permission-mode cases fail on the previous head; the moved-beside-an-unread-step case and an end-to-end diff/audit test of M5 and M6 are added beside them. * Address review cycle 6 on agent launches in CI (#823) C6-F1: the cycle 5 move guard only refused a move when an unread step of that agent stood at a step label the lost launch held, or the job had more unread steps than before. Merging job a's `npm i -g @anthropic-ai/claude-code` step into `claude -p --dangerously-skip-permissions Review` while job b added that plain step (V1), or removing an unread `echo "claude"` step while the launch became a quoted, unread step at another index (V2), kept the unread count and missed every held label, so the row said the bypass "moved between jobs (a/steps[1] -> b/steps[0])" and gave no widening signal. Job a may still run the launch. An unread step carries no text that tells which launch it is, so `may_still_meet` now refuses the move whenever the losing job has any unread step of that agent, wherever it stands and whether or not it was there before (option (a) of the review). V1 and V2 are now `widened`, with the "may still start an agent" caveat, and name "a step no longer declares an agent launch this audit reads". The cycle 5 guard that kept a move beside an unread step the job already had (`claude mcp add x`) now widens, in the safe direction; it moves to the widening cases, and a move beside a step that names no agent stays a move. C6-F2: `_codex_override` caught only `TOMLDecodeError` and `RecursionError`, and `tomllib` raises a plain `ValueError` for an integer past Python's 4300-digit limit, so `run: codex exec -c sandbox_mode=<5000 digits> Review` crashed `diff`, `audit --host` and `check` (exit 1) and `verify` (exit 4, internal_error). It now catches `ValueError`, which `TOMLDecodeError` subclasses; such a value is read as text and selects no sandbox, so the row is `changed`. `-c default_permissions=<4400 digits>` and `codex e --config=sandbox_mode=<4301 digits>` are covered too. Wording: the support page, STABILITY (Direction and What is not read) and the CHANGELOG say a launch has not left a job that keeps any named unread step of that agent or a launch of it with an unread input the rule is read from. The support page and STABILITY also name the one case where a launch that stops being read is worded as moved rather than as no longer declaring a launch: another job adds the same launch while the job it left keeps none of those, as when the launch became a script in the same change. Nonblocking items fixed: - docs/integrations.md: "One that gains a rule" read as the unread `run:` of the sentence before it; it now says "An agent launch that gains a rule". - The HostWorkflowAgentSettingV7 docstring, published as the 0.7 schema's description, states that a `settings` or `mcp_config` value that neither starts like a JSON object nor is a plain file path publishes only a `` digest. The 0.7 inventory and baseline schema files are regenerated. * Address review cycle 7 on agent launches in CI (#823) C7-F1: `may_still_meet` looked for unread steps only under the losing job's name, which holds nothing once that job is renamed or removed, and `job_left` accepted any job that no longer exists. So renaming job a to a2 while quoting its `claude -p --dangerously-skip-permissions Review` (R6), or running it through `npx` (R7), or removing a while an existing job c gained the quoted launch (R3), as job b added the plain launch, was one `low changed` row saying the bypass "moved between jobs (a/steps[0] -> b/steps[0])", with no widening signal. The launch may still run, unread, under a2 or in c. `may_still_meet` now also refuses the move while any job other than the losing and the receiving one has more unread steps of that agent, or read launches of it holding unread text in an input the rule is read from, than it had at the base; a job new at the head counts from zero. `job_left` applies the same check, so both move paths refuse alike. R6, R7, R3, the same while the losing job remains, and a rename into an unread argument input are widenings; a plain rename, a rename with the install step the job keeps beside the launch, and a rename beside another job that keeps its unread step stay moves. Two jobs renamed at once, one of them holding an unread step, cannot be told apart from R6 and now claim the gain, in the safe direction; the support page, STABILITY, the CHANGELOG and the `AgentRuleGains` docstring say so. C7-F2: `_EXPRESSION_RE` (`\$\{\{(.*?)\}\}`) scanned to the end of the text from every unterminated `${{`, so a job `env:` or an agent action's `settings` input holding a few hundred KB of them stalled `diff`, `audit --host` and `check` (92 s at 300 KB, 1070 s at 1 MB). The pattern now ends at `}}` or the end of the text, as `_EXPRESSION_SPAN_RE` does, and only a closed expression names a secret, so what `job_secrets` publishes is unchanged. 300 KB and 1 MB now take 0.8 s and 1.0 s end to end; a new case bounds about 360 KB of them, and the settings input, to 5 s. Nonblocking, fixed: - A declared `shell:` was read whenever its first word was `bash` or `sh`, so `bash -c 'claude -p --dangerously-skip-permissions Review' {0}` beside `run: claude -p Review` read the step and missed the template. A template is now read only as `bash`/`sh` running the script alone: `set` flags, `-l`, `-i`, `-r`, `-o`/`-O` with a name other than `noexec`, and `--noprofile`, `--norc`, `--posix`, `--login`, `--restricted`, `--noediting`, `--verbose`, then `{0}` last. Anything else (`-c`, `-s`, `-n`, `--rcfile`, words after `{0}`) leaves the step an unread, named limit. The schema description of `unread_agent_runs` says so, and the 0.7 inventory and baseline schema files are regenerated. - The coverage line for a workflow that changed with no row said "redacted values such as env values and apiKeyHelper are not compared", which named nothing a workflow holds. A workflow's line now says "text this entry does not read, such as a step's env or an unread agent step, is not compared; audit --host names each unread agent step". Other files keep their note. The capability_diff row in docs/distribution-surfaces.md and its parity comment are updated. Name an unchanged instruction limit reached through an in-tree link instead of refusing (#822) (#863) * Name an unchanged instruction limit reached through an in-tree link instead of refusing (#822) An instruction file whose limit this entry cannot resolve, such as a SKILL.md whose metadata holds a non-string value, is named as an unchanged limit when a change leaves it alone (#721). Reached through an in-tree link the reader reads through (#700) - `.claude/skills -> ../.agents/skills`, a per-skill link or a file link - the same untouched file refused the whole comparison: diff, verify and the manifest-free PR comment printed `base_inventory_incomplete; head_inventory_incomplete` and no row, hiding a removed deny rule beside it. `unchanged_limits` asked blob_path_unchanged whether the source, the path the link is read under, was one regular file in Git, and a path through a link never is. blob_path_unchanged now resolves the path on each side the way the reader reaches it: from Git tree entries for the base and a commit head, and for a working-tree head without following any link (each component's own entry, each link's own text, the file's unfiltered hash). It holds only when both resolutions are equal - every link at the same path with the same text, every other component a directory - and the file they land on is the same regular-file blob at the same in-tree path. A link is followed only under the rules the reader and the base archive already use (#700, #711): a relative text that lands inside the tree after normalization, directories above where it lands, at most eight links, and links at one component of the path. A path no link reaches still takes the one `ls-tree` it took before, and blob object IDs are still compared, so no filter or textconv can make two byte sequences equal. Everything that asks the proof moves together: unchanged_limits in `diff --json` and verifier.json, the text and PR comment, the unchanged limits a partial comparison may carry (#808) and the shared plugin-reference limits check leaves out (#714). check's boundary result still cannot name a limit, so it now refuses these comparisons with unchanged_limits_not_representable, as it does for a limit at its own path; its rows, decision and violations do not move. No schema, member, reason code or check id is added. The metadata value is not coerced. tests/test_linked_unchanged_limits.py holds every layout (directory link, per-skill link, file link, a chain of file links) to the direct result on diff, verify, the PR comment and check; keeps the negative controls refused (skill added or edited behind the link, link retargeted to an identical copy, link text rewritten to land on the same file, link replaced by a directory or the reverse, a later hop retargeted, working-tree-only retargets and edits); and holds the proof to exactly the links the reader reads through, across dangling, looping, absolute, escaping, over-long, intermediate and nested links. The #812 coverage case that pinned the refusal is retired, and the static-only allowlist follows its two call sites down the file. * Address review cycle 1 on linked unchanged limits (#822) The STABILITY migration note said that where the other routes now compare past a limit reached through an in-tree link, `check` refuses with `unchanged_limits_not_representable` and its rows stay empty. That holds for an instruction file's limit, which `check` cannot leave out, but not for a plugin-reference limit: `_without_shared_plugin_reference_limits` asks the same unchanged proof, so a `parse_failed` `plugin.json` that is a file link to an unchanged target is now left out exactly as one at its own path has been since #714. `check` then compares and publishes the removed `deny` row in its boundary result and its control envelope's `capability_rows`, where it refused with `base_inventory_incomplete` / `head_inventory_incomplete` and no row, and `diff`, `verify` and the PR comment, which withheld `plugins/demo` as `partial` (#808), are `comparable` with the limit in `unchanged_limits`. `check`'s decision, violations and control state, and `verify`'s control state and next action, do not move. The note now splits `check` by whether it may leave the limit out, adds the `partial` to `comparable` move, and says "every row's value" where it said "every row". The CHANGELOG entry mirrors it, the distribution-surfaces row and `docs/host-boundary-support.md` no longer say any change behind the link refuses the comparison (an edited plugin manifest keeps it `partial`), and `tests/test_linked_unchanged_limits.py` pins the plugin manifest at its own path and behind a file link, on every route, with the edited-target control. The `blob_path_unchanged` and `_reader_path` docstrings now state the proof as necessary, not sufficient: it follows the links on the way to the path, not every condition the reader puts on reading a whole linked directory, and a side whose reader does not read the path carries no limit there. A test holds a head that adds a link inside the linked directory to a refusal on every route. The two pinned call-site lines in `cli/verify/git.py` move with the docstrings. Build(deps-dev): Bump coverage from 7.16.0 to 7.16.1 (#843) Bumps [coverage](https://github.com/coveragepy/coveragepy) from 7.16.0 to 7.16.1. - [Release notes](https://github.com/coveragepy/coveragepy/releases) - [Changelog](https://github.com/coveragepy/coveragepy/blob/main/CHANGES.rst) - [Commits](https://github.com/coveragepy/coveragepy/compare/7.16.0...7.16.1) --- updated-dependencies: - dependency-name: coverage dependency-version: 7.16.1 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Pengfei Hu Build(deps-dev): Bump platformdirs from 4.11.7 to 4.11.11 (#842) Bumps [platformdirs](https://github.com/tox-dev/platformdirs) from 4.11.7 to 4.11.11. - [Release notes](https://github.com/tox-dev/platformdirs/releases) - [Changelog](https://github.com/tox-dev/platformdirs/blob/main/docs/changelog.rst) - [Commits](https://github.com/tox-dev/platformdirs/compare/4.11.7...4.11.11) --- updated-dependencies: - dependency-name: platformdirs dependency-version: 4.11.8 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Pengfei Hu Build(deps-dev): Bump filelock from 3.32.5 to 4.0.1 (#841) Bumps [filelock](https://github.com/tox-dev/py-filelock) from 3.32.5 to 4.0.1. - [Release notes](https://github.com/tox-dev/py-filelock/releases) - [Changelog](https://github.com/tox-dev/filelock/blob/main/docs/changelog.rst) - [Commits](https://github.com/tox-dev/py-filelock/compare/3.32.5...4.0.1) --- updated-dependencies: - dependency-name: filelock dependency-version: 3.32.7 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Pengfei Hu Build(deps-dev): Update weasyprint requirement (#707) Updates the requirements on [weasyprint](https://github.com/Kozea/WeasyPrint) to permit the latest version. - [Release notes](https://github.com/Kozea/WeasyPrint/releases) - [Changelog](https://github.com/Kozea/WeasyPrint/blob/main/docs/changelog.rst) - [Commits](https://github.com/Kozea/WeasyPrint/compare/v69.0...v70.0) --- updated-dependencies: - dependency-name: weasyprint dependency-version: '70.0' dependency-type: direct:development ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Pengfei Hu Publish hook matcher, command summary and timeout, and MCP launch arguments (#819) (#851) * Publish hook matcher, command summary and timeout, and MCP launch arguments (#819) A hook row read `PostToolUse → PostToolUse` whether the edit was to the hook's matcher, its command or its timeout, and an MCP server whose version pin moved to `@latest` read as a change "in a detail this output does not show": the grants carried none of it, and only `config_sha256` saw the edit. Host-grants inventory, baseline and drift move to 0.7 and the runtime contract to 41. A hook grant adds `handlers[]` (each handler's group `matcher`, its `type`, a `command` summary `{env_keys, argv0, args, omitted_args}` and its `timeout`) and `omitted_handlers`; an `mcp_server` grant adds `args` and `omitted_args`. A declaration outside the documented hooks shape publishes `handlers: null` and its row says the detail is not shown. Plugin-selected and Codex hooks keep their loading basis. Every published word passes through the #802 label redaction, with `Bearer` values replaced first; the value after a credential-named flag, of an `env`-style `NAME=value` word and a generated-looking word are ``; leading shell assignments keep only their names; a home path is written from `~`. Words are cut at 80 characters, eight words follow `argv0`, twelve MCP arguments and sixteen handlers are listed, and the text counts the rest. The members display what `config_sha256` already binds, so grant equality and every inventory digest leave them out: a change is a row exactly when it was one before, a 0.6 baseline compares with no new row or reason and its digest still verifies, and `audit --host --save-baseline` may now replace a 0.6 baseline (older ones are still refused). The shared capability rows render the difference for hooks as they do for MCP servers: `PostToolUse: matcher Edit → Edit|Write|Bash`, `PostToolUse: timeout 10 → 600`, `PreToolUse: handler 2 timeout 5 → 50`, `docs: args -y example-mcp-server@1.2.3 → -y example-mcp-server@latest`, in `diff`, `verify` text, the PR comment and `check` text, and as `review.changes[].change` in `diff --json` and `verifier.json`. Row values, the row count, `check`'s boundary result and the envelope's `capability_rows` are unchanged; verifier 0.20 and capability diff 0.3 do not move. Measured against the prepared 1.1.0 commit e3c6cb0c on the 80 vendored benchmark cases: rows byte-identical on all 80, 42 entries on 35 cases gain detail, and all 7 changed hook/MCP entries that were content-free now name their field. The benchmark replays reproduce their run-of-record scores, and the route readiness dry run was rerun for the contract bump (identical cells apart from the two version numbers). Tests: tests/test_hook_mcp_detail_fields.py covers the issue's four fixtures on every route, the published v0.7 schemas, redaction of a token in a command, a secret positional argument, env-style assignments and an over-length command, display-only equality and digests, 0.6 baseline compatibility and re-save, loading basis, and the out-of-shape limit. * Address review cycle 1 on hook and MCP detail fields (#819) Refs #819 A header credential after any scheme other than Bearer was published. The label rule replaces only the first word after `Authorization:`, and the detail pre-pass special-cased only `Bearer`, so `Authorization: Basic ` printed as `Authorization: ` in hook commands and MCP arguments on every route, and `Authorization: Bot `, `X-Auth-Token: ` and `api-key: ` printed their token. `_detail_label` now runs the label rule and then replaces the whole value of a credential written `Name: value`, the scheme included, up to the closing quote or the end of the text: `Authorization`, `Proxy-Authorization`, `Cookie`, `Set-Cookie` and any header or key whose name is, or ends in, a credential word (`X-Auth-Token`, `api-key`, `X-API-Key`, a JSON `"token":`). All three schemes now publish `Authorization: `. The name only starts where a run of name characters starts, so the scan stays linear. The Bearer pre-pass is no longer needed and is gone. `audit --host --save-baseline` wrote hook commands and MCP arguments into the committed baseline, including ones read from `~/.claude/settings.json`, `~/.cursor/mcp.json` or managed settings under `--scope local-static`, or from a git-ignored `.claude/settings.local.json` in either scope. Those values were never in the repository, and a short positional password such as `-p hunter2` matches no word rule. A saved baseline now holds each grant as comparisons read it (`compared_grant`), with no `handlers`, `omitted_handlers`, `args` or `omitted_args`. `HostGrantsBaselineV7` keeps the 0.6 snapshot, which forbids those members, so a saved 0.7 baseline's grants are exactly a 0.6 baseline's. Nothing read them before. `inventory_sha256` is unchanged: the reviewer's fake-HOME reproduction gives the same `5e406f56cd73…` digest before and after, with 4 leaked values before and 0 after. A comparison between two commits still renders both sides' detail through `host_comparison_baseline`, which applies the saved baseline's checks to the full normalized inventory and is never saved. MCP-argument redaction is now at least what the digest's input redacts. The word after any item the digest's list rule treats as a credential marker is ``, with or without dashes (`token X`, `password X`); that rule is factored into `_is_list_secret_marker`, which the digest and the display share. `--auth X` is redacted, as the hook string rule already did. A flag name is read with every character but letters and digits removed, so `--brave_api_key` matches the `apikey` ending. From the non-blocking notes: `-u`/`--user`/`-U`/`--proxy-user` values keep the user and lose the password after `:`. A timeout of `5` becoming `5.0` now reads `timeout 5 → 5.0`, because handler fields are compared as their JSON publishes them. Tests: the credential canary sweep adds a Basic header, a custom `X-Auth-Token` header and `-u user:password` in a hook, and a Basic header, an `api-key:` header, a bare `token`, `--auth` and `--brave_api_key` in MCP arguments. It now also checks the saved baseline. The per-argument rule test takes lists, so it pins the next-word cases the reviewer named. A new test checks that every argument the digest's input redacts is published redacted. A local-static baseline saved with a fake HOME holds none of its hook or MCP values while the inventory still shows them, and still compares with no drift. A repository baseline holds nothing from `.claude/settings.local.json`. `5` to `5.0` names both values. The baseline 0.7 schema is regenerated, which differs from 0.6 only in its version. STABILITY's migration note gains a Saved baselines bullet, its redaction bullet is restated, and "a hand edit to a baseline's copy" is gone because there is no copy. The CHANGELOG, host-boundary support doc, agent contract and distribution-surfaces row now say the same, and the row's redaction claim is narrowed to "permission-rule argument redaction". llms-full.txt is rebuilt. `diff --json` over the 80 vendored benchmark cases gives rows byte-identical to the prepared 1.1.0 commit on all 80. Review entries are identical to the previous head on all 80, and 42 entries on 35 cases differ from 1.1.0, as the CHANGELOG states. Rebased onto main after #853 moved the published-release pins to v1.1.0. The conflict resolutions landed in the rebased first commit. The README and quickstart quotes now include the `billing` server's `args`, which the published 1.1.0 does not print, so `test_host_diff_entry_docs` requires both pages to use the not-yet-released label. Each page names the published `1.1.0` and says what it prints instead. llms.txt and the AI-search summary name v1.1.0 (contract 40) as published and this tree's contract 41 as unreleased. The pilot ledger keeps main's 2026-09-22 v1.1.0 measurement and the #819 source-tree rerun beside the 1.1.0 release commit, and its source-tree column reads contract 41 and inventory schema 0.7. * Address review cycle 1 on hook and MCP detail fields (#819) Refs #819 Generated credentials that hold a `.`, `:` or `;` were published whole in hook commands and MCP arguments. The generated-key test ran only on a word made entirely of the base64 alphabet, so one separator defeated it: a SendGrid `SG..` key, a Telegram `:` token, an Airtable `pat.` token, a Discord bot token, a Mapbox `sk..` token and an Azure `AccountName=…;AccountKey=` connection string reached `diff`, `verify`, the PR comment, `verifier.json` and `check`. `_without_generated_runs` now tests each run of the base64 alphabet inside a word, split at every other character and at an `=` that separates a name from its value, and replaces each run that reads as a key. Once one run of a word is a key, every other run in it with a key's shape (20 or more characters of two classes) is replaced too, because a token's other parts are no less random and are often too short for the entropy test alone (a Mapbox signature). A run followed by `=` is an assignment's name and is kept, so the Azure string publishes `AccountName=acct;AccountKey=`, and the hex of a `sha256:`, `sha384:` or `sha512:` digest is kept as the pin it is. The whole-word test still runs first, so no word it redacted before is published now. A credential flag right after another credential-named flag published its value while the digest's input redacted it. In `--no-password --token abc`, the boolean `--no-password` took `--token` as its value, the next-word rule stopped there, and `abc` was published; the digest's list rule reads `--token` as naming `abc` and redacts it, so rotating `abc` changed the published arguments with no row. Which word is replaced now depends only on the word before it (`_credential_values`), so both words are ``. The same case in a hook command was also published, because the digest's string rule had already written `--token` as `` before the words were split. A word that follows a credential name in the command as written is now `` wherever the redacted words still hold it. An edit past the 12-argument bound read "no difference in … arguments". When either side declares more arguments than the grant publishes, the MCP sentence now says `the first 12 arguments` and names `an argument past the first 12` among what it does not show. A hook with more than 16 handlers says `of the first 16 handlers` and names `a handler past the first 16`, and one whose command has more than eight arguments names `a command argument past the first 8`. From the non-blocking notes: the digest's own string rule now runs on the text as written before any other pattern, so a known token shape that runs into the flag after it (`sk-…--password X`) no longer hides the value that rule redacts. `--secret-key X`, `--aws-access-key X` and `--pass X` are redacted as credential flags. STABILITY and the CHANGELOG now name the shapes that still publish (`-p hunter2`, `-phunter2`, `--key hunter2`, an e-mail address) and the over-redaction of an image whose name ends in a credential word (`ghcr.io/org/auth:`). `-u` after a replaced word keeps its user-and-password rule, and a redaction marker after `-u` is no longer split at its `:`. Tests: the canary sweep adds all six joined token shapes, the chained flag, the access-key, secret-key and pass flags, and a glued `sk-…--password` value, across a new hook and a new MCP server, and checks every route and artifact as before. The word table pins each joined shape and the benign controls: a digest pin, an image tag, `@scope/pkg@1.0.14`, `pkg==1.10.1`, a dotted `$CLAUDE_PROJECT_DIR` path and a connection string without a key. The digest-superset invariant adds the reviewer's three chained cases and is checked over every list of up to four words from a vocabulary of flag shapes. A value-level test covers glued words in both an argument list and a hook command. Rotating the value after `--no-password --token` stays quiet and redacted. A 14-argument docker config whose tag is the last argument, a seventeenth handler and a command's tenth argument each name the bound. Against the previous head's engine, 26 of the new or changed test cases fail and the benign controls pass; all pass here. Re-running the 80 vendored benchmark cases through `diff` against the previous head gives identical rows, review entries, text and published hook and MCP detail (437 words) on all 80, so no real-world argument is redacted that was not before. The replay tests reproduce the run-of-record scores unchanged. * Address review cycle 2 on hook and MCP detail fields (#819) A hook timeout written as an integer too large for a float crashed every route that read the hook. `_hook_handlers` called `math.isfinite` on any int, which converts it to a float first, so a timeout of 1 followed by 400 zeros made `diff`, `diff --json`, `check` and `audit --host --json` exit 1 with OverflowError, and `verify` exit 4 with no PR comment and no verifier.json. `_hook_timeout` now publishes a finite float, or an integer of at most 80 digits, as the number it is, and any other value as its bounded text; an integer is never converted to a float, and its bit length bounds the digits before it is turned into text. The credential header rule ran on a hook's whole command, where an unquoted value runs to the end of the text, and `PWD` ends in `pwd`: `docker run --rm -v $PWD:/src --fix` published `-v $PWD:` and nothing after it, so the image moving to another registry with `--privileged` added read "no difference", and an added `echo auth: ok; curl ... | sh` hook printed only `echo auth: `. The digest's string rule and the label rule still run on the whole command; the header rule now runs on each word, as it already did for MCP arguments, and never takes a `$NAME` shell variable as a header name. A word that ends in a credential header name and its colon, as an unquoted `-H Authorization: Basic ` splits, takes the next word as its value, and the word after that when the next is an authentication scheme, so splitting at whitespace publishes no credential. docs/host-boundary-support.md said a change confined to the value after a credential-named flag is still a row. For a value the digest's own input redacts (after --token, --api-key or --password, a --password=... value, an X-Api-Key: header value) it is not compared and is no row, as on 1.1.0. That page, STABILITY and the CHANGELOG now say so, and keep "still a row" for positional tokens, generated keys, a header's words after its scheme, a flag the digest does not name, words past a bound and unpublished settings; the CHANGELOG says the display's redaction never hides a change. The review's nonblocking findings, all fixed: - `args` that is not a list on both sides no longer says arguments were compared: its entry reads "such as the command's path or arguments", as on main. - "handlers past the first N" names the bound of 16, not the other side's listed count. - Escaped JSON inside a double-quoted shell word (`{\"password\": \"...\"}`) is read by the header rule, which accepts a backslash before a quote, so its value is no longer published. - A POSIX shell's `-c` script (`bash -c "X=1; curl ... | sh"`) published `X=` and nothing after it. In a shell's script a leading assignment's value now ends where the shell ends it, at the first whitespace outside quotes and escapes, so it publishes `X= curl ... | sh`. Anywhere else a `NAME=value` word's value is still the rest of the word, since `docker run -e "FOO=a b"` sets FOO to `a b`, and a value holding a substitution, a parenthesis, a brace or an open quote is still replaced whole. - The digest's credential-assignment rule took time quadratic in a long run of name characters (40,000 characters of `password`: 1.6 seconds in the digest input, 6.6 seconds for one hook command's detail). A lookahead for the `=` and value every match needs removes it (now under 0.1 seconds); it matches exactly what it matched, with the same spans and groups over 200,000 random strings, so every config_sha256 is unchanged. A test pins both. The 80 vendored benchmark cases give the same rows and the same review entries at this commit as at the previous head, so the CHANGELOG's measurement stands, and the replay tests pass unchanged. * Address review cycle 2 on hook and MCP detail fields (#819) Three display rules took time quadratic in repository-controlled text, so one config file near the reader's 1 MiB bound stalled diff, check, verify and audit for minutes to over an hour: - the credential header rule's blanks around the colon backtracked into the value, whose first class also holds a blank ("token:" and 64,000 blanks took 20 seconds). They are possessive now; a value cannot end in a blank, so the rule matches exactly what it matched; - the shell-flag test, -[A-Za-z]*c[A-Za-z]*, backtracked over every c of a long option cluster. It now tests -[A-Za-z]+ and looks for the c apart; - the digest-pin test searched the whole word before every hex run for a sha256: prefix. It reads only the seven characters before the run. Timing each reader near the bound found three more of the same kind in the hook reader. shlex grows a word one character at a time, so a 1 MiB command took 16 seconds to split; _command_words is now a one-pass reader that returns exactly shlex's words. The leading-assignment loop copied the word list once per assignment, and the -c script's assignment loop copied the rest of the script once per assignment; both walk by index now. Near the bound every shape reads in under two seconds, and a test holds each under ten; a randomized test holds the rewritten rules to what they published before. In a shell's -c script an assignment's value now also ends at an unquoted ;, & or |, so `X=1;curl ... | sh` names curl instead of hiding it, and every word of the script that starts with an upper-case NAME=, or a quote and one, is read as an assignment rather than only the leading ones: `export DB_PASS=...` and `cd /x && DB_PASS=... ./run.sh` publish DB_PASS=, as the same words in a plain hook command already did. A < or > does not end a value, since a marker an earlier rule wrote there holds both. docs/quickstart.md said an edit confined to a credential-redacted argument says so. A value the digest's own input already redacts (after --token, --api-key or --password, or an X-Api-Key: header value) is not compared and prints no entry, as the other pages state; "says so" now covers only the command's path, any other redacted argument and what is past a bound. STABILITY and the CHANGELOG describe the script rule, say that a script after env or sudo is read as any other word, and name lower- and mixed-case assignments (db_pass=..., a connection string's Pwd=...) among what no rule recognises. Rebased onto #821 (#852), which also moved the runtime contract to 41. The resolution, in the commits that introduced each entry, extends contract v41 in place, and the CHANGELOG, STABILITY, docs/agent-contract-current.md and contract.py name each change's versions (verifier 0.21 and capability diff 0.4 are #821's, host-grants 0.7 is on the combined tree beside the 1.1.0 release commit: the same cells, the JSON differing only in schema versions, #821's coverage members, init's contract version and input id, and the added MCP grant's empty args. * Keep the 1 MiB reads within bound on a traced CI runner (#819) The new near-bound tests failed in CI's first suite shard: the shell-flag hook and the two many-assignment shapes took 10 to 13 seconds against a 10-second bound, where a laptop takes under two. The shard traces coverage on a shared runner, which multiplies every line of Python, and those shapes still read a 1 MiB word a character at a time in Python in several places. Those places now find the next character that matters with a pattern instead: the command splitter reads a piece at a time (a run of blanks, a quoted run, a lone quote, a run of anything else), the -c script walker and the value-end scan jump to the next quote, backslash, blank or separator, the generated-key test searches for each character class and counts letter and digit runs with patterns, and the generated-run scan reads only runs of twenty or more characters, the only ones it can replace. A 1 MiB word splits and scans about ten times faster. A second randomized test holds each scan to what its character-by-character form published. The many-assignment shapes spend their time in per-word Python that no pattern removes, so the bound is now 60 seconds: at this length the reviewed shapes took from over a minute (the hex runs) to over an hour (the header blanks). * Address review cycle 2 on hook and MCP detail fields (#819) Inside a shell's -c script, a credential word still hid the rest of the script. _published_word ran the whole label rule on the script, the credential header rule included, and a header's unquoted value runs to the end of the text it is found in. So `bash -c "echo token: ok; ./notify.sh"` and `bash -c "echo token: ok; curl ... | sh"` both published `echo token: `, and diff, verify, the PR comment and check called the change a detail they do not show. A docker `-v ~/.aws/credentials:...` mount hid the image and --privileged after it. The --flag=, -u and list rules read only the words outside the script, so `bash -c "curl -u admin:pw ..."`, `--api-key=X`, `--secret-key X` and `token X` published their values. A script is now read one shell word at a time. The string and label rules still run on the whole script first, then the assignment scan. After that, _script_words splits the script at whitespace, ;, &, |, parentheses and backticks outside quotes and escapes, and each word goes through the rules a hook command's word does: the header rule on that word alone, the --flag=, -u, env, generated-key and home-path rules, and the word after a credential name, an unquoted header name or -u (_credential_kinds, which _credential_values now reads too). The script is read both as written and after the string rule, as a hook command's words are, so `--no-password --token X` inside one publishes `--no-password `. A rewritten word loses its quotes unless it was one quoted word, and every other character is copied as written. Words are read only up to the 80-character bound, so a long script costs its first words. Three shapes the word rules published as written are now redacted: - a quoted credential assignment, which the digest's assignment rule does not read: pwsh's $env:API_KEY='x', node's process.env.TOKEN='x', an MCP --env=API_KEY='x' and one argument `export API_KEY='x'`; - curl's -u with its value glued on, -uuser:password; - the text a quote split from a URL. The string rule's URL ends at a quote, so `curl "https://x/a?token="abc` published https://x/abc once the word's quotes were removed. A URL is now read again on its word. The same rule covers a script URL that took `;X=` into its path and left X's quoted value glued to it; the randomized comparison below found that one. STABILITY, the CHANGELOG and docs/host-boundary-support.md describe the script rule and the three shapes. They also state the limit that remains: text after a blank inside a quoted URL is still published. They now say that a comparison against the working tree reads a git-ignored .claude/settings.local.json, so its commands can appear in local pr-comment.md and verifier.json, which the docs raised only for baselines. New tests pin the reviewer's shapes in a hook command and in MCP args, on every route, with canaries. A randomized test holds the script word scan to a character-by-character reading, and two more 1 MiB shapes hold the new scans linear. Comparing the previous head with this one on 73,000 generated scripts found no canary this head publishes that the previous one hid, and no command word left hidden. Both published all 204 commands and 82 argument lists in the vendored benchmark cases identically. * Address review cycle 3 on hook and MCP detail fields (#819) Inside a shell's -c script, a credential that only the as-written reading catches was published once earlier text in the script collapsed three or more words. _published_script read the words as written "two ahead" of the redacted words, on the premise that no rule adds or removes a word. The whole-script string rule does remove words: its URL takes an unquoted `?a&b&c&d;` and its credential assignment takes an unquoted `TOKEN=a|b|c|d`. The secrets set was not yet filled when the value was published, so `bash -c "curl https://x.invalid/?a&b&c&d; echo Authorization: Basic X"` published `Authorization: X`, and `--no-password --token X`, `--auth -u u:X` and `--auth token X` after either prefix published X, in hook commands and in MCP args alike, on every route. Every word of the script as written is now read before any word is published, so the two readings no longer need to keep step. The scan is linear. A near-1 MiB script of 500,000 words takes about a second. The word after a credential word was replaced across ;, | and &&, so it hid the next command's first word. `gh auth token | docker login ...` published `gh auth token | login ...` on both sides, and a docker -> podman edit read "no difference in the matcher, type, command summary or timeout". `echo token:; ./notify.sh` published `echo token:; `. _credential_kinds takes a starts_command test, and _script_command_words says of each script word whether a ;, &, |, newline, parenthesis or backtick comes before it. Both readings start afresh at each command. No digest redaction is lost: the string rule has already run on the whole script and still replaces what it reads across a separator (`--token |X` publishes `--token `). The same reset applies to a hook command's own words. A word that is only control operators (|, ||, &&, ;, &) starts a new command, so `gh auth token | docker login ghcr.io` keeps its pipe instead of reading `token docker`. An MCP server's args are not read by a shell, and the digest's list rule redacts whatever item follows `token`, `|` included, so those args are read as before. After sudo or env, or in a script run by a shell that is not POSIX such as pwsh -c, the script is still one word. An unquoted credential `Name:` word's value runs to the end of it: `sudo bash -c "echo token: ok; curl ... | sh"` publishes `echo token: `. STABILITY already stated the assignment consequence. It now states this header consequence too, and docs/host-boundary-support.md and the CHANGELOG limit "never hides the commands after it" to the script of a POSIX shell that is itself the command. They also say what the string rule still takes across a separator. Tests: each of `--no-password --token X`, `echo Authorization: Basic X`, `--auth -u u:X` and `--auth token X` behind no prefix, a `?a&b&c&d;` URL, a `?a&b&c;` URL and a `TOKEN=a|b|c|d` prefix, in a hook command and in MCP args. A route test finds none of the review's three canaries in diff text or JSON, verify text, pr-comment.md, verifier.json, check text or audit --host --json. A route test shows the docker -> podman edit named on every route. There are separator shapes (`echo token:; ./notify.sh`, `gh auth token; ./deploy.sh`, `&&`, `|`, parentheses, `Authorization: Basic; ./run.sh`) and shapes where the string rule takes a value across a separator (`|X`, a newline, `&X`). The top-level operator words are covered, and so is the MCP `token |` parity. Two near-1 MiB shapes guard the full pre-read in linear time. Against the previous engine, 20 of the new cases fail. * Address review cycle 4 on hook and MCP detail fields (#819) Stop publishing command and argument text. Four review cycles each found a credential inside free-form shell text that a redaction rule missed (a quoted word, a -c script, a separator, a here-document), and cycle 4 found another: a -u, --pass, --secret-key or token value inside a quoted word that is not the command's own script. Redacted shell text cannot be made safe by adding rules, so per the PM decision of 2026-09-23 none of it is published on any surface. - A hook handler's command is now {executable, sha256}: the last path segment of its first word, only when that is a plain token no redaction rule rewrites and not a shell reserved word (otherwise ), and the SHA-256 of the whole command as config_sha256's input holds it. The type, env_keys, argv0 and args members are gone. - An MCP server grant publishes package, at most one argument of a strict npm, PyPI or OCI shape that follows no flag but a package runner's own, and args_sha256, the digest of every argument with the package replaced by a marker, in place of args and omitted_args. - The matcher keeps the published-label redaction; a non-string matcher is . A timeout written as text is published only as a plain token. - Every free-text redaction rule this change had added is removed: the word, script, header, -u, flag-value and generated-key rules and the command splitter. The digest's assignment-rule lookahead stays. The members remain display only: each is a function of the configuration as config_sha256's input holds it, grant equality and the inventory digests leave them out, and a saved baseline holds none of them. A reorder of the published handlers now reads "the published handlers in a different order; a detail this output does not show may also differ, such as ...". Equal published handlers never establish equal handlers, since a setting such as async is not published. The PR comment no longer loses rows to one long entry. Its bound cut at the first line that did not fit, so a long hook entry hid every later row, the change count and the review question. An entry is now printed whole when the whole comment fits, and otherwise cut to the widest of 480, 240 and 120 characters at which it does, with a pointer to verifier.json. A comment written without a readiness report now points to verifier.json instead of a report.md that route does not write. Rows, row counts and digests are unchanged. diff --json on the 80 vendored benchmark cases gives byte-identical rows beside e3c6cb0c, and the seven content-free hook and MCP entries there now name their field. The README and quickstart answers are the published 1.1.0 ones again. STABILITY, the CHANGELOG entry, the host-boundary, contract, index, distribution-surface and pilot-ledger docs, the v0.7 schemas and llms-full.txt describe the smaller surface. * Address review cycle 5 on hook and MCP detail fields (#819) - A hook command whose first word is a URL no longer publishes the URL's host as its executable. The digest's input keeps a URL's host and drops its userinfo, query and path, so the last segment of `http://deploy:pw@build-cache.corp.internal?token=x` was the host. A first word holding `://`, as written or as that input holds it, now names no executable (``), as the STABILITY note already said. - When only one side's hook declaration is outside the documented shape, the row names that side and lists the other side's handlers, as an added hook's cell does: `PostToolUse: base matcher, command and timeout not shown (...); head (matcher Edit; command a.sh sha256:...)`. It used to say the declaration was malformed without naming a side, so a PR that repairs a hook block read as though the new block were the malformed one. Both sides outside the shape read as before. - A matcher longer than 1,024 characters is `` and never reaches the published-label redaction, whose jwt and database-URL patterns take quadratic time. Only the listed handlers are read for publishing, and a group's matcher once: the matcher was redacted once per handler, so a 100,000-character matcher over 2,000 handlers took 17 seconds to read and a 400,000-character one over 20,000 handlers did not finish in ten minutes. Both now read in about 0.1 seconds. - `args_sha256` digests the package's position beside the marked arguments, so a literal `` argument can no longer make two different argument lists digest alike. - `uvx --with` is no longer a package runner's flag: its value is an extra requirement, not the server, so the server's own pin is the published package. - A timeout published as text that reads as a finite number prints quoted, `timeout 5 → "5"`, instead of `timeout 5 → 5`. - docs/design-partner-pilot-results.md no longer describes a draft of this change that never shipped. The STABILITY note, CHANGELOG, host-boundary-support, the v0.7 inventory schema's descriptions and the tests follow. No command or argument text is published anywhere; every command digest is unchanged. * Address review cycle 6 on hook and MCP detail fields (#819) - The PR comment keeps every line 1.1.0 kept. With about 13 long hook entries the comment lost what 1.1.0's printed: each entry was cut to 120 characters and followed by its own 57-character verifier.json pointer, so a cut entry cost about 180 characters against 30 for `PreToolUse → PreToolUse`, the comment still did not fit, and its bound cut the coverage block, the review question, the reproduction, the advisory and the evidence line, and from 16 rows row headings too. The lines 1.1.0 printed now get their room first: the coverage block's budget and the agent instruction block are chosen with every entry in its shortest form, and the entries get only what is left. The first that fits is printed: every entry whole; every longer entry cut to the widest length of at least 60 characters at which the comment fits (bisection); entries in their shortest form, longest first; and that without its note. An entry's shortest form is a field-level difference cut after its name (`PreToolUse: …`, `docs: …`) or an added or removed grant's own row (`(absent) → PreToolUse`), printed only where shorter; none is longer than the entry 1.1.0 printed for the same row, and a permission rule's entry and a joined change are never shortened. One line after the rows, not one per entry, says entries were shortened and that verifier.json holds each whole. The omission line of a comment without a report is as long as 1.1.0's, so a comment 1.1.0 itself cut loses no line 1.1.0's cut kept. - A boolean timeout is published as the JSON boolean, a non-finite float as ``, and every timeout written as text prints quoted, so `true` → `"true"` and `Infinity` → `"inf"` read `timeout true → "true"` and `timeout → "inf"`, not "no difference". Both used to publish the word a string could spell. - The matcher's 1,024-character bound applies to the text as config_sha256's input holds it, so `Bash(TOKEN=<10 chars> x)` and the same rule with a 1,100-character value, which share a digest, publish the same matcher. The digest's string rule already runs over every matcher and is linear; the quadratic label redaction still never reads a long one. - A declaration whose command is not a string reads `the declaration is not a list of matcher groups whose hooks are objects and whose commands are strings`, as the STABILITY Shape bullet already stated. On the cycle-6 reproductions (two hook scripts moved under 14 and 13 events, and 16 async-only edits) every line of 1.1.0's comment, entries aside, is in this one. From 1 to 40 moved hooks, with and without a permission change, in both comment styles, every line 1.1.0 printed is kept, a coverage block that lists what 1.1.0's counted aside, and where 1.1.0's comment overflowed this one keeps at least as many lines. The v0.7 inventory schema, STABILITY, the CHANGELOG, host-boundary-support and the capability_diff registry row follow. Rows, row counts, digests and baselines are unchanged, and no command or argument text is published. Keep independently established host changes when one scope is incomparable (#808) (#861) A pull request that broke plugins/demo/.claude-plugin/plugin.json and also dropped a deny rule from .claude/settings.json printed "Cannot compare against main: head_inventory_incomplete" with no row on diff, verify and the manifest-free PR comment, because #714's unreadable-manifest limit refused the whole host comparison. Nothing was called safe, but the reviewer lost the one widening the change made outside the plugin. The plugin reader now records, as private snapshot facts, the plugin directory each plugin-reference issue is bounded by and the plugin directory of every inline hook selector. The shared comparator (diff, verify/PR, check) turns a refusal into comparison_status "partial" only when every refusing limit is a plugin-reference limit bounded by a directory below the repository root, every other blocking limit is an unchanged one proven as #721 proves it, the directory does not hold the project settings files, no outside marketplace declares inline hooks for a plugin inside it, and something outside it was read. Those directories are withheld on both sides (artifacts, grants, non-blocking issues, and a hook file another plugin also selects) and the rest goes through the same payload builder a comparable result uses. Any unproven condition keeps the refusal. Published in the unreleased verifier 0.21, capability diff 0.4 and contract v41: "partial" with the refusal's incomparable_reasons plus the rows, review and unchanged_limits established outside the directories, and coverage.items[].scope (reserved by #812) naming the withheld directory on each blocking_limit item. The models refuse a partial result with no reason, no coverage or an unscoped blocking limit, and a scope anywhere else; the legacy reader refuses a 0.20 artifact claiming either. Text output leads with the partial scope and reason before any row and never prints "No static host-grant changes detected" for a partial result. Authority does not move: partial is not comparable, so verify's control state, permissions, next action, merge verdict and exit code, the control envelope (projected as incomparable with no rows), check's decision and violations, the Stop hook, digests, baselines, drift, audit --host and host-grants 0.6 are unchanged. Only plugin-reference limits are scoped in this slice; unresolved skill structure and link limits stay refused and are pinned by tests. docs/distribution-surfaces.md's capability_diff row and tests/test_distribution_surface_parity.py are updated together, with STABILITY.md, CHANGELOG, agent contract, host-boundary, README and quickstart docs. Review: cycle 1 at 62b6ef2d had no P0-P2 findings; a follow-up pass at 877d1195 found one P2 (in a partial comparison a project settings change that moved a withheld plugin hook's loading basis was published as changed_without_grant_change), fixed at 2e4447ad: a partial comparison treats a withheld hook's loading basis as not shown to be unchanged, so such a settings file is changed_without_rows on every route. Cycle 2 at 2e4447ad had no P0-P2 findings. The branch was then brought up to date with main after #807 (#862) by a conflict-free merge that leaves the PR's own changes identical. Closes #808 Refs #714, #812, #821 Stop verify --preview in a configured repository publishing human_review_required for input capture (#807) (#862) In a repository with a committed shipgate.yaml, `verify --preview --json` answered agent_action_required with the documented verify command, but `--format control` answered human_review_required ("input directory capture is unavailable"), and every `agent control` refresh exited 4 with workspace_unverifiable. The preview's pointer bound the verification plan it records and took its workspace_identity from it, yet a preview runs no adapter, so that plan never carries an input-directory census and every reader refuses it. A preview pointer now binds the verifier route without the plan, and its identity is the working-tree identity a manifest-free preview already used (repository, HEAD, tree, snapshot_kind "worktree_overlay" and the overlay of paths differing from HEAD, output directory excluded under the shared #575/#804 containment rule). `--format control` and the refresh now agree with `--json` and are refused as workspace_changed after a tracked edit, new untracked file or removal, and current again once restored. Fail-closed is kept: a pointer that does bind a census-less plan (such as one a 1.1.0 configured preview left behind) is still workspace_unverifiable with a verify recovery, and `verification worker` still refuses to replay the preview's plan. Under Git configuration the worktree readers refuse (#813), the plan-less fallback identity used to declare no snapshot kind, so the refresh compared HEAD alone and read any later edit as current. Inside a repository that identity now always declares the worktree snapshot, and such a pointer is refused as workspace_unverifiable with the cause first and a review next action. This also covers a verify that stopped before building a plan (missing --config or --head). The plan-less workspace_changed refusal no longer claims paths "no longer have the content they had"; it names up to three paths that differ from HEAD, redacted with redact_text (#802 class). No new public surface: no command, schema version, report block, error kind, refusal code, exit code or minimum_control_contract_version moves. Docs updated in docs/agent-contract-current.md and docs/verification-reproducibility.md, with a STABILITY.md migration note, a CHANGELOG entry under Unreleased and regenerated llms-full.txt. Tests: new tests/test_preview_control_currency.py (14 tests; 13 fail on the prior main) and an extended tests/test_agent_control_reports_dir.py assertion. Review: one review cycle with no P0-P2 findings at the merged head; CI green. Closes #807 Route an enabled in-repository plugin's hook at a non-registry path to review (#809) (#860) A hook selected by a plugin that the repository's own project settings enable from an in-repository marketplace was published as execute/high with an expansion signal, yet check and a manifest-backed verify routed only registry paths: the issue's reproduction (plugins/demo/cfg/hooks.json, command changed) gave a widened row beside check allow and verify passed/mergeable, while the same hook at .claude/hooks/hooks.json required review. Routing now needs two facts, decided in boundary_registry: a path predicate (is_claude_plugin_reference_path, the hook, plugin.json and marketplace.json files the plugin hook reader opens) and the reader's own loading evidence (a private, unpublished HostBoundarySnapshot.enabled_plugin_hook_sources from the same selection that publishes project_enabled_plugin). A routed file joins the existing SHIP-AGENT-BOUNDARY-PROTECTED-SURFACE-UNCLASSIFIED rule (require_review), and a blocking plugin-reference limit on it adds SHIP-AGENT-BOUNDARY-INPUT-INCOMPLETE. An enabled plugin's hook file the reader refuses to open (enabled_plugin_unread_hook_files) is incomplete input, not protected-surface review. plugin_selected hooks stay unrouted, so the noise #714 removed does not return. enabled_plugin_hook_evidence reads the merge base and head with the same scope and snapshot builder as the host comparison, so a hook file deleted with its reference is routed from the base; check and verify take the same evidence (verify through two run-private VerificationContext fields excluded from serialization). Rows, the host inventory, diff, a manifest-free verify and the trigger catalog are unchanged; a manifest-backed verify and its PR comment move to review_required / human_review_required with check. No new public surface: two existing check IDs fire in more cases, which STABILITY.md allows; the migration note enabled-plugin-hook-routing-809 records it. The documented remaining limits are a selector-only change that makes an enabled plugin load an unchanged hook file, and a marketplace made unparseable keeping its non-blocking limit. CHANGELOG, STABILITY.md, docs/host-boundary-support.md and the capability_diff row in docs/distribution-surfaces.md are updated; tests/test_enabled_plugin_hook_routing.py adds 36 tests. Review cycle 1 at 2b80370a found one P1 (check and a manifest-backed verify disagreed on a deleted enabled hook file, because verify read only the head) and three P2 (notes claimed verify was unchanged; an unread enabled hook file was allowed with complete input; an unreadable-head claim was false for a marketplace), plus P3 wording and test fixes. All were addressed; the re-review at 0b982f62 found no P0-P2 findings, with CI green. Name the changed inputs a host comparison did not read (#821) (#852) A pull request that added a Cursor plugin's `mcp.json`, removed a guard from `.cursor/hooks.json`, changed a host settings file nested below the repository root, or moved a marketplace plugin's pinned source printed "No static host-grant changes detected", the answer a docs-only change gets, and a manifest-free `verify` handed it to the setup route. On the 23-PR public corpus, 0 of 9 comparable zero-row results named the changed relevant file. `core/unread_inputs.py` now matches the comparison's own changed-file set (`base..head`, or the working tree's tracked and untracked changes) against a bounded, documented candidate rule set, and the comparator adds each match no inventory published to the existing coverage list as `changed_not_read`: - plugin manifests' `mcpServers`, and Codex/Cursor/Copilot manifests' `hooks`, when their text differs; a manifest or marketplace that does not parse; a marketplace entry's object `source`, compared as text and never fetched; `.cursor/hooks.json`; host settings below the repository root; `mcp.json` beside a plugin manifest; a hook-named file a Codex/Cursor/Copilot manifest names. - At most 32 candidates, one presence question and one read per side, only plugin manifests and marketplaces read, no link followed, nothing fetched or run (`cli/verify/changed_inputs.py`, reusing the bounded identity-question plumbing). Past the bound, or where a needed file was not read or did not parse, a candidate is counted in `unread_candidates_not_examined`. - The item ranks right after blocking limits inside the existing cap and is never a row, a widening or a `check` violation. `read_sources_only` is false while one is named and the block's first line says so; otherwise that line is unchanged. - A manifest-free `verify` (and `verify --preview`) whose only host-relevant change is such an input, or a not-examined count, publishes the comparison on the advisory host route instead of the setup route. Verifier 0.20 -> 0.21 (new `docs/verifier-schema.v0.21.json`, 0.20 frozen), capability diff 0.3 -> 0.4, runtime contract 40 -> 41. Rows, digests, baselines, host-grants 0.6, `audit --host`, `check` and the benchmark replays are unchanged; `minimum_control_contract_version` stays 21. Re-running the 9 comparable zero-row corpus PRs: 9 of 9 now name the changed unread input, rows byte-identical on all 11 PRs measured. Review cycle 1: re-ran the route readiness dry run against contract 41; `verify --preview` moves with `verify`; marketplace entry names in coverage sources are redacted, bounded and digested; the not-examined line names both of its causes; over-deep JSON and an unpinned url source no longer crash or mis-redact; rebased onto #853 with the 1.1.0 labels kept accurate. Review cycle 2: rebased onto #849 (#827) keeping both sides in STABILITY, the surface registry and the parity test; lone surrogates in marketplace entries are escaped instead of crashing `diff`/`verify`; a manifest-free `verify` keeps a comparison whose only finding is a not-examined count; docs say "added", "removed" or "changed". Cycle 2 closed with no P0-P2 findings and green CI. Closes #821 Refs #812 Make check and diff agree on prompt-disabling settings (#827) (#849) Every Claude Code setting that disables or narrows a prompt now gets one rating on every surface. Before, audit --host, diff and check each rated defaultMode, enableAllProjectMcpServers, skipDangerousModePermissionPrompt and enabledMcpjsonServers differently, and diff rows showed a bare value. - core/host_settings.py is the single table rating each setting the host inventory publishes as a permission_mode grant, with a recorded basis. host_grants takes access and risk from it, capability_diff_rows prints "setting: value" with the basis as the row's why, and host_boundary raises HOST-PERMISSION-WILDCARD-ALLOW for a critical value and HOST-PERMISSION-ALLOW-EXPANDED otherwise. - critical: bypassPermissions, skipDangerousModePermissionPrompt: true, enableAllProjectMcpServers: true. high: acceptEdits, auto, undocumented modes, each enabledMcpjsonServers entry. medium: dontAsk, plan, default and the remaining switches. dontAsk moves from a critical block to medium. - check reads a modelled setting where the inventory does (permissions first, then top level); only a value a change sets raises a violation. defaultMode evidence keeps its {kind, mode} shape so existing fingerprints hold. - enabledMcpjsonServers becomes one high external grant per approved server with per-entry precedence keys; disabledMcpjsonServers stays an unknown key and is documented as unread. - No new public surface or version bump; migration notes in CHANGELOG and STABILITY record the rating and decision moves for saved baselines. - New tests/test_prompt_disabling_settings.py pins grants, diff rows, check, verify rows and findings on real two-commit repositories. Review: cycle 1 found a P1 (check evidence carried a setting's raw value before sanitization), fixed by keeping raw values out of evidence. Cycle 2 found a P2 (a value moved between permissions and the top level raised nothing), fixed by keying comparisons on container and key. The final cycle-2 review at 3bc3ef25 reported no P0-P2 findings, with CI green. Closes #827 Move the published-release pins to v1.1.0 (#853) docs/release-runbook.md § Cutting the release, step 8, done after the v1.1.0 Release was public (precedent e2ab0007, #777). Moving these pins before the tag existed is the #506 failure. Constants. LATEST_PUBLISHED_VERSION = "1.1.0" and LATEST_PUBLISHED_CONTRACT_VERSION = "40" in published_release.py. Every other change follows from what the enumerating tests then required (test_public_surface_contract, test_adopter_pins_resolve, test_distribution_surface_parity), plus the measured surfaces below. Pins. The Action, pip, uvx and shipgate_version pins in the GitHub Actions, CircleCI and GitLab examples, incidents, samples, the bug-report template, .well-known, docs, skills, plugins and rendered prompts now name 1.1.0 / @v1.1.0. Rendered prompts and CI recipes were re-rendered with the package's own renderer (skills/, .agents/skills/, plugins/ mirrors and prompts/ stay byte-identical); the five render hashes move, and the outgoing renders the v1.1.0 tag carries are appended to prior_render_sha256 in both adoption-kit metadata files so an unmodified install still upgrades cleanly. Statements about the newest release. v1.1.0 is described as advisory, contract 40, no qualification claim; v1.0.0 becomes the previous release. llms.txt, docs/ai-search-summary.md, docs/agent-contract-current.md and the regenerated llms-full.txt name v1.1.0; the "unreleased, ahead of" qualifier is dropped now that source and published contracts are equal. Historical records (changelogs, STABILITY migration notes, frozen schemas, run-of-record evidence) are untouched. Measured surfaces re-taken on the published build. - The five documented diff answers in the README and quickstart were re-captured with agents-shipgate 1.1.0 installed from PyPI into a clean virtualenv outside any checkout, on the fixture tests/test_host_diff_entry_docs.py builds. "Not yet released" becomes "Released in 1.1.0". - The pilot ledger's route readiness dry run was re-run on 2026-09-22 against PyPI 1.1.0, the source tree and PyPI 1.0.0. Only factual rows were updated; the standing 2026-09-14 research decision is unchanged and gains a dated factual checkpoint whose channel question is left to the research owner. - The examples README "What runs" names the commit init --ci from 1.1.0 writes (e3c6cb0c...), measured from the installed wheel, and the v1.0.0 checkout-import known issue is noted as fixed in v1.1.0. - The runbook's step 8 now names both measured surfaces and the tests that hold each. Tests. test_host_diff_entry_docs.py holds each page to _PUBLISHED_ANSWERS, the sha256 of each normalized answer as published 1.1.0 prints it (taken from the PyPI wheel, never this tree), since version strings cannot tell the two builds apart. _PUBLISHED_ANSWERS_VERSION must equal LATEST_PUBLISHED_VERSION, so the next step 8 fails until the record is re-taken. Negative controls seeded from the pages' own quotes (the v1.1.0-tag labels, an unprinted line, an older release label, reference lines from another build, missing and doubled labels) each assert the specific rejection. The contract-statement guard now exercises its equal-contract branch with the exact qualifier llms.txt carried. No public surface is added. docs/distribution-surfaces.md changes only prose in the github_action row and the check-run declared-exception note, kept consistent with the matching comments in test_distribution_surface_parity.py; no claim or parity row moves. Deliberately not changed: .github/release-channels.json, tags and releases, the 1.1.0 CHANGELOG section (the entry is under a new "## Unreleased"), and the pre-commit rev: v1.0.0 pins left to #796. Post-publication checks were observed only: the v1.1.0 tag peels to e3c6cb0c; PyPI serves the single wheel (sha256 038bdb46...) that matches the Release asset; the Release is latest, immutable, with SBOM, advisory statement, provenance, candidate manifest and Sigstore bundles. The Marketplace listing body, the external threemoonslab.com .well-known and llms.txt, and the immutable 1.1.0 PyPI description still carry v1.0.0-era wording owned outside this repository. Review. Cycle 1 findings were addressed in a follow-up commit (strengthening the published-quote guard to the recorded _PUBLISHED_ANSWERS and adding its negative controls); cycle 2 found no P0-P2 issues. CI green on the reviewed head. Refs #778 Prepare the 1.1.0 advisory release (#847) Prepares 1.1.0 on the advisory channel, per docs/release-runbook.md § Cutting the release (steps 1-3) and § The advisory release channel (step 1). It stops at "ready for the owner to review": no tag was pushed, no GitHub Release was created, the pypi environment was not approved, and nothing was uploaded anywhere. Why MINOR, not patch. STABILITY.md § Versioning defines a minor as new features adhering to the contract. This cycle adds new readers (#771 workflow step action references, #693 reusable-workflow secret mappings, #802 label redaction), new row kinds and semantics (#795 moved/widened/narrowed, rows[].disposition; #812 coverage) and four schema moves: host-grants inventory 0.5 -> 0.6, runtime contract 39 -> 40, capability diff 0.2 -> 0.3, verifier 0.19 -> 0.20. Channel: advisory, per docs/release-evidence-policy-decision.md § Amendment 5. Qualified is not an option here - .github/release-trust-roots.json's signer_identity is still CHANGE_ME and the four SAFETY_QUALIFICATION_* repository variables are unset, so the qualified path fails closed. A tagged version's channel is frozen forever, so declaring 1.1.0 advisory means 1.1.0 can never carry a qualified claim. The release note. CHANGELOG.md's "## Unreleased" is promoted to "## 1.1.0 - 2026-09-22", reshaped to one line per change, with a framing paragraph, a Highlights block and the two pointers 1.0.0 carries. The full reviewed prose for all 23 entries moves to docs/changelog/1.1.0.md (new), linked from docs/INDEX.md. The section extracts at 8,835 characters, within scripts/release_notes.py's 125,000-character limit. Release-note claim discipline, measured by re-running a 23-PR public corpus on main before and after this cycle. The notes say 1.1.0 is a legibility and presentation-correctness release with no change in what is analysed, and state the limits plainly: audit --host output byte-identical on 12 of 12 pairs diffed and zero new files read; rows byte-identical on 22 of 23 pull requests; comparable 18 of 23 in every measurement; automatic author-actionable yield 0 of 23 before and after; a zero-row result still does not name the relevant unread file (#821); hook script bodies, plugin-package files and marketplace inputs are still not read. The notes make no claim of increased coverage, improved findings, rows, severity or comparable rate, and no "catches issues" framing. Version-status reconciliation. Several entries described themselves as unreleased or unpublished and would have shipped a self-contradicting release body; reconciled in the tree, in both CHANGELOG.md and the record. STABILITY.md. Title stamped · 1.1.0. Nine "Migration Note: Unreleased" and three "Migration Note: 1.0.x" headings become "Migration Note: 1.1.0", with every anchor above them untouched so no inbound link breaks. Six preamble sentences and four note bodies are restated so each stays historically true after publication. Deprecation clock. SHIP-VERIFY-AGENT-INSTRUCTIONS-WEAKENED, SHIP-CODEX-BOUNDARY-AGENTS-SHIPGATE-REQUIREMENT-REMOVED and SHIP-CODEX-BOUNDARY-SKILL-COMMAND-CHANGED move from "deprecated in the unreleased minor cycle" to "deprecated in 1.1.0". § Versioning counts cycles in shipped releases, so this starts the clock: the earliest hard removal becomes 1.2.0. Version sites. pyproject.toml, src/agents_shipgate/__init__.py, .well-known/agents-shipgate.json, docs/agent-contract-current.md, the three plugin/marketplace manifests, docs/design-partner-pilot-results.md, llms.txt and docs/ai-search-summary.md (keeping their "unreleased" qualifier), README.md and docs/quickstart.md, the regenerated schemas and llms-full.txt, docs/examples/capability-lock.v0.8.example.json, eight tests/golden/codex_boundary_result/*.json and tests/test_v07_metadata_roundtrip.py. Channel declaration. .github/release-channels.json gains "1.1.0": "advisory", without which release-advisory-verify.yml's "Refuse a version that is not declared advisory" step stops the rehearsal; tests/test_release_channel.py's committed-declaration assertion moves with it. Deliberately unchanged. No tag, no GitHub Release, no pypi approval, no upload. No published-release pins: published_release.py's LATEST_PUBLISHED_VERSION and LATEST_PUBLISHED_CONTRACT_VERSION, .well-known's release_status.latest_release and package.github_action, docs/agent-contract-current.md's "Latest release", the four VERSION_LITERAL_TARGETS and the ~50 files carrying @v1.0.0, ==1.0.0, agents-shipgate@1.0.0 or shipgate_version: "1.0.0" all stay at 1.0.0. Those are runbook step 8, after publication; moving them before the tag exists reproduces #506. The two-phase "unreleased"/"not yet released" labels stay, as tests/test_public_surface_contract.py requires while contract 40 != 39. No published run-of-record evidence was rewritten - docs/changelog/1.0.0.md, docs/changelog/0.16.0.md, the 1.0.0 benchmark run-of-record directories and the published verifier-schema.v0.19.json / report-schema.v1.0.json bytes are untouched. No GitHub issue text was edited. Still to do, and owner-only: Release Engine Smoke then Advisory Release Rehearsal dispatched on this merged main commit (the rehearsal locates the smoke by head_sha and a byte-identical candidate manifest, so order matters and re-running the smoke afterwards invalidates the rehearsal), then the tag, the pypi environment approval, and the step-8 pin move. Rank and bound the coverage block, and stop it reading as a complete account (#812 follow-up) (#846) Refs #812 The `What this run established` block now states its own boundary as its first line, so a true list cannot be read as the account of the change; `coverage.read_sources_only` says the same in JSON. Enumerating changed-but-unread files stays #821. Blocking limits are ordered by kind before the cap — `unreadable`, `parse_failed`, `unresolved_precedence`, then `unsupported`, `dynamic_source_excluded`, `remote_source_excluded` — then by source name and every remaining field an item is keyed by, so no two items tie and one comparison always publishes one list. That is order, not severity, and it ranks kinds rather than items. Truncation now reads `N more items not listed, each ranked below those above` in the text as well as in `omitted_items`. A zero-row result and a refused comparison now end with the compared commits, the tool version and the reproduction command, labelled `Inputs:` rather than `Compared:` where the comparison was refused; a run that names no base commit still prints neither. The review question stays only where there is a change to ask about. An instruction file whose declared structure could not be established is no longer told to repair a file whose own text parses. Only `frontmatter_invalid`, `frontmatter_unterminated` and `instruction_text_invalid` keep the repair wording; `frontmatter_not_mapping` and `structure_value_unencodable` are split out of `frontmatter_invalid` so that list is true, and a `~` or `null` header keeps the digest an empty header has always produced. Schema extended in place: capability diff `0.3` and verifier `0.20` are unpublished. No row, digest, baseline, control state, permission, next action or exit code moves. Make diff JSON carry the row semantics the text now shows (#795 follow-up) (#845) Refs #795 On the 23-PR public corpus one run was described two ways: `diff` printed one widening from 20 rows while `diff --json` published two rows carrying `expands: true` and no disposition, `moved` or `narrowed`. Since agents are told to parse `--json` and the JSON shape is the contract, a machine consumer read a widening count the text contradicted, and could not reproduce the sentence a human was reading. The presentation is now published beside the rows, additively, from the one place that computes it: - `rows[].disposition` — the `allow`, `ask` or `deny` list a permission rule is declared under, `null` for every other grant kind, on every route that publishes rows including `check`'s boundary result, where a rule's arguments stay redacted and the disposition does not. - `review` — top-level in `diff --json`, `host_comparison.review` in `verifier.json`: `changes[]` in printed order, each naming its `row_indexes`, the `direction` the text uses (including `widened`, `narrowed` and `moved`), the `before`/`after` cells, the field difference, its `why` and one `expands`; `summary` with the three numbers the summary line prints; and `question` plus `reproduce_command` verbatim. The block is one projection of `review_changes`, the function the text already reads, built in `compare_host_inventories` where the rows are built, so it cannot become a second answer. `reproduce_command` moved into the same module and `comparison_reference_lines` calls it, making the printed and published command one string. The contract refuses a block whose changes do not stand for every published row exactly once, whose counters disagree, or that joins two rows reading alike — which keeps the redacting routes safe. Every row keeps the values, count and order it published on `1.0.0`, and the text is unchanged line for line. Capability diff `0.3` and verifier `0.20` are unreleased in this tree, so `review` is extended in place rather than bumping again; `minimum_control_contract_version` stays `21`. Rows are not closed objects in any published schema, so a reader pinned to `0.19` still validates one carrying `disposition`; `host_comparison` is closed, so a strict `0.19` reader rejects a `0.20` artifact for `review` exactly as it already does for `coverage`. Per docs/distribution-surfaces.md the `capability_diff` row and its parity-test comment are updated in the same change; the surface still makes no new claim. Migration note in STABILITY.md#host-diff-review-json-795. Tests: `tests/test_host_diff_review_changes.py` adds seven tests covering the corpus-shaped fixture, the parity of the block across `diff --json`, verify JSON and the PR comment, the redacting routes, the contract refusals, and the `1.0.0` row schema still validating a row with `disposition`. Two legacy-reader tests drop `review` and `disposition` alongside `coverage`/`unchanged_limits` before reading a payload back as an older artifact. Reviewed in one cycle with no P0-P2 findings; full CI green at the reviewed head. Say what each host comparison established, including zero-row and incomparable results (#812 slice 1) (#838) A zero-row or incomparable host comparison could not be read: a docs-only PR, an `env`/`apiKeyHelper` edit and a deleted env-only settings file all printed the same `No static host-grant changes detected`, and an unreadable or unsupported input printed only `base_inventory_incomplete; head_inventory_incomplete` with no source. The facts were already in the drift payload; the projection dropped them, and recovering them took a separate head-only `audit --host`. `core/host_comparison.compare_host_inventories` now builds one capped coverage list beside the rows, derived only from facts the comparator already computes, so `diff`, `verify` and the PR comment share one implementation. - Statuses: `compared` (with row counts), `changed_without_grant_change`, `changed_without_rows`, `unchanged_not_proven` and `blocking_limit`, each with the `side` that read or published the source. Grouped per source, status, side and limit; at most ten items plus `omitted_items`, ordered so quiet sources drop first. - Identity is answered per run by a three-way `blob_path_identities` question (identical / differs / neither) at a fixed number of Git processes, reading `CRLF` checkouts as unchanged and never following links; `unchanged_limits` keeps its two-way proof. - Paths are the already-redacted `public_host_path`, and details are the issue's sanitized message. Coverage stays out of inventory digests, baselines and drift payloads, and `check` and the provided-diff route record none. - Surfaces: `diff --json` `coverage` (capability diff `0.2` -> `0.3`), `verifier.json` `host_comparison.coverage` (verifier `0.19` -> `0.20`, with the generated `docs/verifier-schema.v0.20.json`), and a `What this run established:` block in `diff` text, `verify --format text` and the PR comment, bounded so it never pushes out a line the comment already prints. Legacy `0.19`/`0.18` artifacts read as `0.20` with `coverage: null`; one claiming coverage is refused. - Headlines, refusals, reasons, rows and the incomparable next action are unchanged. Runtime contract v40 is unreleased and extended in place; `minimum_control_contract_version` stays `21`. - Docs: AGENTS.md, `docs/INDEX.md`, `docs/agent-contract-current.md`, `.well-known`, regenerated `llms-full.txt`, the `capability_diff` row in `docs/distribution-surfaces.md` with its parity-test claims, a STABILITY migration note and a CHANGELOG entry. - Tests: new `tests/test_host_comparison_coverage.py` drives real repositories through `diff` text and JSON, `verifier.json` and `pr-comment.md`, asserting the routes agree; updated route-parity, advisory-recipe, entry-docs, review-changes and unchanged-limit suites. Review: six cycles. Cycle 1 tightened attribution for plugin manifests, marketplaces and retargeted links; cycle 2 replaced the "unread field" wording with `changed_without_grant_change` and added the identity question; cycle 3 stopped a checkout line-ending conversion reading as a change and held the Git process count constant; cycle 4 repinned the adapter static-only line numbers; cycle 5 bounded the block inside the PR comment so it cannot displace the review question, reproduction, advisory or evidence lines; cycle 6 found no P0-P2 issues. CI was green at the reviewed head. Slice 1 of #812. Unread candidate discovery is #821 and per-scope retention is Refs #812 Show concrete permission before/after and a review question in host diffs (#795 slice 1) (#837) Slice 1 of #795: host-diff text now names what changed concretely, so a non-author reviewer can read the change without translation. Text only; every JSON surface (diff --json rows, verifier.json, check rows, control envelope capability_rows) is unchanged. - core/capability_diff_rows.py builds an in-memory review view per row while grants are in hand; review_changes(rows) produces the entries printed by diff, verify --format text, pr-comment.md (both styles) and check --format text. The view is not a row field and is excluded from equality. - core/host_grants.py exposes permission_rule_replacements, the same one-out/one-in allow pairs the lattice decided (#816), and builds the expansion signals from it. Pairs print as one widened/narrowed entry; identical rule text leaving one disposition for exactly one other prints as moved; replacements take precedence over moves; nothing is paired by likeness or case; redacting routes (check, provided diffs) never join. - MCP entries show only published facts (transport, command name or sanitized URL, env and header key names, capped and label-redacted); unprintable URLs read "url not shown", and changes in unpublished details are named as such instead of "gh -> gh". - Comparable results with entries add a review question; commit and worktree heads with a base commit add Compared (short SHAs, tool version) and Reproduce lines. The widening mark now appears in verify text, the PR comment and check text. diff counts entries, noting "from N rows" when rows were joined. - Docs: README/quickstart quotes re-captured and marked not yet released, STABILITY.md migration note (#host-diff-review-text-795), distribution-surfaces capability_diff row and parity-test comment, host-only recipe notes, CHANGELOG Unreleased. - Tests: new tests/test_host_diff_review_changes.py exercising real repositories through diff, verify (commit and worktree heads), the PR comment and check (range and --diff), asserting JSON rows unchanged in every case; updated permission-direction, entry-docs, manifest-free PR rows and workflow step reference tests. Review: three cycles. Cycle 1 fixed a command whose path changed but name did not being described as unchanged (now labelled "command name" with an explicit not-shown fallback) and added move/replacement edge tests. Cycle 2 stopped raw paths of URLs the sanitizer leaves as written from reaching any text output ("url not shown") and made the review question count joined rows. Cycle 3 found no P0-P2 issues at 6067c2ed with CI green. Hook matcher/command/timeout and MCP arguments or version pins remain for Refs #795 Make relative --out resolution explicit and printed artifact paths resolvable (#818) (#836) Commands resolved a relative --out against three different bases (verify: Git root of --workspace; scan: the manifest's directory; audit --host and agent control --reports-dir: the current directory), and no help text said which. From a sibling directory, verify wrote an untracked reports directory into the scanned checkout (which the next run counted as changed files), printed artifact paths that did not exist from the caller, and emitted reproduction commands carrying the relative --out. fixture run --out rel wrote into a temporary copy it then removed, and audit --host --out failed with other_error/4 on a temp-file rename. One rule, current_workspace.explicit_output_path, now governs an output location typed on the command line: absolute as given, relative against the current directory. Omitted locations keep their defaults (default_reports_dir under --workspace for verify and agent control; the manifest's output.directory for scan). verify, agent control, scan and fixture run share it. - Where a relative --out now lands somewhere 1.0.0 did not (verify outside the Git root, or outside --workspace for --preview outside Git; scan outside the manifest directory), one stderr note names both directories. Stdout, verdict and exit code are unchanged, and nothing new prints from the Git root. - verifier.json on disk keeps workspace-relative artifact paths; verify --format json stdout and scan's Reports: list spell them relative to the current directory when beneath it, absolute otherwise, via the shared caller_path_spelling rule. - fix_task.verification_command, the commands derived from it, and preview verify commands name --out absolutely. - audit --host --out is refused before the inventory is read or a baseline is saved: config_error, exit 2, naming /host-grants.json, with a replayable next action only while that file does not exist. - --help for verify, scan, fixture run and audit states the resolution base, file-vs-directory, and the default. - Resolution happens before the #804 output-directory classifier, so --out . typed inside a tracked or .claude/ directory is still refused. - STABILITY.md carries a migration note listing the affected callers. No new command, flag, schema, JSON field, error kind or exit code. Tests: new tests/test_out_path_resolution.py (34 CLI tests against real Git repositories) plus updates to the tests that pinned or relied on the Git-root rule. Review: two review cycles; the second found no P0-P2 findings at the merged head, and CI was green. Closes #818 Read Claude Code :* rules and moved rules correctly in host diff direction (#816) (#834) The host diff read three ordinary permission edits in the wrong direction, putting a widening mark on changes that widen nothing; two of them also changed `check`'s decision. The fix corrects the one permission lattice (`core/permission_lattice.py`) and signal builder (`core/host_grants.py`) every route already reads, adding no command, schema, check ID or JSON member. - `:*` suffix: for `Bash` rules only, `subsumes` reads a trailing `:*` as the documented trailing ` *`, so `Bash(npm:*)` -> `Bash(npm test:*)` is a decided narrowing like its space spelling. A trailing ` *` also covers the bare command. A `:*` with nothing or whitespace before it stays undecided; other tools (e.g. `WebFetch(domain:*)`) keep the colon as text. Because the space before the star is part of the rule, `Bash(npm run test:*)` no longer covers `npm run test:unit`, so that replacement now requires review. - Moved rules: a rule whose exact text moved between dispositions in the same host and source is set aside before the one-left/one-arrived pairing (moves into `allow` always; moves out of `allow` only when another allow rule also left), so a deny-to-allow move no longer blocks pairing an adjacent narrowing. - One MCP tool: `_is_wildcard_allow` uses the new lattice predicate `names_tools_within_one_mcp_server`, so `mcp____` is scoped while `mcp__` and `mcp____*` stay whole-server grants and unclear tokens keep blocking. `mcp__` compares as `mcp____*`, and a saved baseline that rated a one-tool rule as wildcard is re-read without an expansion signal. - Docs: STABILITY.md extends the unshipped host-grants 0.6 / contract 40 notes in place with the changed signals, row values and decisions; CHANGELOG entry; `SHIP-HOST-BOUNDARY-PERMISSION-WILDCARD-ALLOW` `fires_when`, `docs/checks.md`, regenerated `docs/checks.json` and `llms-full.txt`. - Tests: new `tests/test_host_diff_permission_direction.py` pins `diff`, `check`, `verify --preview` host_comparison and drift signals on real repositories for each issue case and controls; `tests/test_permission_lattice.py` and `tests/test_host_settings_narrowing_review.py` extended. Host-config and cold-start replay benchmarks replay unchanged. Deferred: rewriting one spelling into the other still requires review (direction, not equivalence), and `PowerShell` `:*` rules are still read as text. Review: cycle 1 findings on moved-rule pairing and cycle 2 findings on the colon-continuation record were addressed in follow-up commits; cycle 3 found no P0-P2 issues. CI was green at the reviewed head. Closes #816 Refuse diff cleanly in a partial clone missing base blobs (#817) (#835) In a partial clone whose base objects were never fetched (`--filter=blob:none` or `--filter=tree:0`), `agents-shipgate diff` crashed with exit 1 and an uncaught traceback, and agent mode printed no `next_action`. `verify --preview` already reported `objects_missing`, but its hydration example (`git fetch --refetch origin`) reapplies the clone's configured filter and fetches nothing. `diff` now refuses the way the shallow case does (#683): exit 2, one stderr line naming the side it could not read, and in agent mode an `objects_missing` error whose `next_actions` name `git -C fetch --refetch --no-filter ` for each qualifying promisor remote in configuration order (never an `origin` fallback; a `review` action with a `` placeholder when none qualifies). - `promised_objects_missing` (`cli/verify/git.py`) detects the case from the repository rather than Git's error text: a `rev-list --objects --no-walk` walk that fails normally but succeeds with `--missing=allow-promisor`, run under `GIT_NO_LAZY_FETCH=1`, so a corrupt repository does not count and nothing is fetched. - Only `diff` uses the new `archive_fetched_tree` wrapper; `archive_tree` and `verify`'s published `comparison_input_unavailable:ConfigError` reason are unchanged. - `objects_missing_remediation()` is the single wording shared by `diff` and `verify`, whose example is corrected to `git fetch --refetch --no-filter origin`. - Docs: new `objects_missing` entry in `docs/errors.json` (exit 2, schema version unchanged), AGENTS.md and rebuilt `llms-full.txt`, STABILITY.md migration note, CHANGELOG entry, quickstart recovery held to real output, and a sentence on the `capability_diff` row in `docs/distribution-surfaces.md` (claims unchanged). - Tests: `tests/test_capability_diff_partial_clone.py` exercises real `file://` partial clones for the refusal, no implicit fetch, remote naming, a corruption control, shared wording, shallow-then-partial sequencing, and `check`/`verify` behaviour in the same clone. Deferred: `verify` without `--preview` in a treeless clone still exits 4 with `internal_error`; that changes `verify`'s answer and needs its own issue. Review: cycle 1 findings were addressed in a follow-up commit; cycle 2 found no P0-P2 issues. CI was green at the reviewed head. Closes #817 Describe the host-review task and manifest-free diff in discovery copy (#792) (#833) Discovery copy read first by coding agents and answer engines described only manifest-first tool-source scanning, although the README and quickstart already lead with the manifest-free `agents-shipgate diff` (#779/#784). This is the bounded in-repo slice of #792; the external surfaces in the PR's audit stay open. - llms.txt: the first "What it is", "When to recommend" and a Quickstart entry describe reviewing a change to declared coding-agent configuration (.claude/settings.json, .mcp.json, hooks, Codex, Cursor, VS Code MCP, GitHub workflow permissions) through `agents-shipgate diff`, with its limits. The detect continue rule now names `host_boundary_candidates`, so an agent no longer stops on a host-only repository that detect routes forward. The latest public release (v1.0.0, contract 39) and the source-tree runtime (contract 40, unreleased) are stated on separate lines, and the agent_boundary_result schema is corrected from v2 to v3. - .well-known/agents-shipgate.json: one when_to_use entry naming the route and four task keywords; tagline, positioning, primary_commands and quickstart are unchanged. - docs/ai-search-summary.md: a short host-configuration review section, and the source and published contracts stated apart. - docs/quickstart.md: the detector continue rule names `host_boundary_candidates`. - AGENTS.md: one bullet each in "What this project is" and "When to use this tool"; llms-full.txt regenerated. - pyproject.toml (release-bound): a description clause and four keywords; PyPI changes only with the next published build. - docs/distribution-surfaces.md and the parity test: the agent_instructions row records that the new AGENTS.md bullets add no claim, and the llms.txt classifier reason says hand-maintained rather than generated. No new public surface is added. New tests pin the source/published contract statements in llms.txt and docs/ai-search-summary.md (both spellings), require every detect continue rule to name all routing fields against a real host-only detect_workspace run, and require a .well-known when_to_use entry for `agents-shipgate diff`. Review cycle 1 extended the continue-rule fix to docs/quickstart.md, added the ai-search-summary contract guard and a guard for the "contract v40" spelling. The PR then passed review with no P0-P2 findings and green CI. Refs #792 Say why the worktree could not be read, and never skip the currency check (#813) (#832) live_workspace caught every failure to read the repository and discarded the exception. A worktree Git refused to read statically (a filter=lfs attribute, git-crypt filter/diff config, a local diff textconv, or a non-empty .git/info/attributes) produced a workspace_unverifiable refusal with no stated cause, and agent control pointed back to the producing verify, which refused the same way, so an agent following next_actions looped. When the workspace could not be observed at all, currency was withheld only for complete; review_publishable and every other pointer read with no currency check. - cli/verify/git.py raises UnboundGitConfigurationError (a ConfigError subclass, same message) for the four repository-configuration refusals, carrying a finding without the "commit and verify refs" remediation; the diff.* refusal names keys, never values, and key and path lists go through privacy.redact_text. - core/current_control.py adds LiveWorkspaceCause (kind plus redacted text capped at 280 bytes), LiveWorkspaceUnavailable and LiveWorkspace.changed_paths_cause; workspace_read_cause classifies failures as repository_configuration, resource_limit or other (foreign exceptions reduced to their type name), and #804's refusal gets a reports_directory kind. - The cause leads every refusal the missing view produces and is exposed as CurrentControlUnavailable.cause, so verify --format control/text pick it up through the existing denied() reason. - next_actions[0] depends on the cause: repository configuration gives a review step naming the key or path, stating that re-running verify (with or without --head) does not change the answer, and naming the read-only diff --workspace route; a read bound or timeout keeps the producing verify with "commit or shrink" first; other causes are unchanged. - LiveWorkspaceUnavailable refuses every pointer that binds Git identity (complete as workspace_unverified, others as workspace_unverifiable); pointers binding no Git identity (scan, non-Git verify --preview) and read_current_control(live=None) are unchanged. - No new command, schema, contract version, JSON field, error kind, refusal code or exit code. The behavioural changes (next action kind command -> review for configuration causes; exit 0 -> 4 for non-complete Git-bound pointers when the workspace cannot be observed) are documented in a STABILITY migration note and CHANGELOG. Deferred: metadata-only live_workspace, LFS/filter support in worktree verification, the hooks' filter handling and the orchestrator _safe_worktree_overlay swallow. Tests: new tests/test_live_workspace_cause.py (22 tests, real Git repositories: filter config, filter attribute and diff textconv through the worktree, preview and committed-tree routes; redaction of a token-shaped path across agent control, verify --format control, verifier.json and pr-comment.md; classification and cap; every pointer state refused through the real live_workspace; the review_publishable fail-open end to end; regression pins for live=None and non-Git preview); updated line pins in tests/test_adapter_static_only.py. Review: cycle 1 found no P0-P2 issues. Closes #813 Refuse an output directory that holds repository content (#804) (#831) verify leaves its output directory out of every working-tree read (the Git change set, the worktree overlay, the manifest-free host comparison and the static input census), and agent control does the same to the reports directory when it checks currency. When --out named an in-repo directory holding other content, that content dropped out of the decision: --out naming the agent host settings directory turned a shell-permission widening, or a deny-rule removal, into complete/mergeable, and a later edit under --out docs or --out tools refreshed as complete. - Add classify_output_directory in cli/current_workspace.py, one classifier shared by the writer and every reader. It lists the directory through Git (verify/git.py output_directory_inventory: ls-tree of HEAD and compared commits, ls-files cached and others, check-ignore over the artifact names a run writes) and refuses anything committed there, anything a compared commit held there that the change removes, staged or untracked unignored non-artifacts, any trust-root path whatever its name, an unignored directory inside a trust root, and the repository root or an ancestor. Outside the repository, gitignored, empty/absent, or artifacts-only directories stay accepted. - Containment reuses #801's physical-identity walk; below the root each component is respelled as its stored entry so a case variant of a directory classifies (and, when safe, works) as the real directory. - Writer: verify (worktree, --head, manifest-free) and preview refuse before any write with OutputDirectoryHoldsRepositoryContent, a ConfigError subclass (config_error, exit 2). Worktree runs compare HEAD and the merge base (including a manifest-free local base); archived --head runs compare only the evaluated head. - Readers: live_workspace classifies the reports directory before any other Git read and refuses as workspace_unverifiable for every pointer state, outside the live=None path; worktree pointers are also checked against their recorded merge base. agent control's recovery replays the producing run's verification command with only --out removed. - output_directory_remedy writes one remedy for the message, verify's next action and agent control's review step; the default agents-shipgate-reports is never told to omit --out. - No new command, flag, schema, contract version, error kind, refusal code or exit code. STABILITY migration note, CHANGELOG, agent-contract-current (llms-full.txt regenerated), troubleshooting, integrations, verify --out help and action.yml output_dir document the narrowed contract; pointers 1.0.0 published into such a directory stop reading as current. Tests: new tests/test_output_directory_content.py (real Git repositories: every writer route, committed/uncommitted/untracked widenings, staged and committed deny-rule deletions, trust-root files named like artifacts, symlink and case-variant aliases, legacy pointers, recovery run as printed, the default directory with committed or stray content, all 16 shipped samples and the skill reports leaving only recognized artifacts, and an allowlist test tied to every artifact registry); adjusted tests/test_current_control_closure.py and tests/test_install_hooks.py. Review: cycle 1 stopped an artifact name from granting a trust-root path the allowance, kept --base/--config in agent control's recovery, added the missing suggestion and Action payload names, gave the default directory a non-looping remedy, let an archived --head run remove files its own diff shows, and made the manifest-free writer compare the base its host comparison uses. Cycle 2 added the skill lint/security/review report names and documented the deferred gitignored-census gap and the advisory Stop hook. Cycle 3 found no P0-P2 issues. Rebased onto main after #802 with only CHANGELOG and STABILITY note conflicts, both kept. Closes #804 Redact token-shaped workflow job ids and step labels in host grants (#802) (#815) Host grants published credential-shaped workflow labels verbatim. GitHub accepts a job id, trigger or permission scope name shaped like a token (ghp_..., AKIA..., xoxb-...), and a step name can carry registry userinfo (docker://ci:@gcr.io/...). Those values reached the inventory, saved baselines, drift, diff, check rows and evidence, the control envelope, verify, verifier.json and pr-comment.md. - Add published_workflow_label in core/host_grants.py: redact_text, the host step sanitizer, then a linear label-level scheme://userinfo@ rule that keeps the host and any trailing @algorithm:hex digest. Every job id, trigger, scope name and step label is published through it once, and every grant field, row, why, drift, diff, check, verify and PR-comment output reads that same label. - check evidence.job and evidence.scope in core/host_boundary.py use the same helper; evaluation still reads raw declarations. evidence.old on PERMISSIONS-EXPANDED publishes only read/write/none, else null. - config_sha256 is computed over the redacted projection, so no grant digest is taken over a raw token-shaped label. - Extend the #767 collision rule: two distinct raw job ids, triggers or scope names in one namespace that publish alike block comparison, with a rename hint. Scope names that publish alike merge at the widest level. A lone redacted label is not a collision and refuses nothing. - Host-grants 0.6 and contract 40 are extended in place (unshipped): the regenerated schemas gain descriptions only; STABILITY, the support page, the agent contract and the capability_diff surface row document the behavior and its limits (display-only over-redaction, a rename between labels that redact alike is no row, scheme-less userinfo is not read). Tests: new tests/test_workflow_label_redaction.py (redaction and negative controls, a linear-time scan checked against the old backtracking pattern as oracle, an end-to-end canary across every published output, collision and unchanged-limit behavior through the CLI) and redacted-evidence cases in tests/test_host_boundary_check.py. Review: cycle 1 documented and pinned what an unchanged collision costs (check incomparable, --save-baseline exit 2, drift incomparable), corrected the digest claim (the artifact redacted_sha256 still covers the parsed file), narrowed evidence.old and kept the widest level for colliding scope names. Cycle 2 replaced a quadratic backtracking userinfo pattern (73s on a 300k-character label) with a linear scan checked against the old pattern, added the rename hint to the collision message and documented check reading a renamed lone job against top-level permissions. Cycle 3 found no P0-P2 issues. Rebased onto main after #814 without conflicts. Closes #802 Keep check coherent when settings add keys outside the host-boundary rule allow-list (#810) (#814) `shipgate check` exited 1 with an unhandled pydantic ValidationError ("publication authority requires complete boundary input coverage") whenever a changed host settings file (project or local settings.json, or Cursor cli.json) set a top-level key outside the host-boundary rule allow-list (enabledPlugins, extraKnownMarketplaces, outputStyle, enableAllProjectMcpServers, enabledMcpjsonServers, unknown Cursor keys), in every maintained format and in both worktree and --base/--head modes. Root cause: core/codex_boundary.py treated the resulting `unknown_host_config_key` parse-failure row as read input (publishable), while _coverage_for in core/agent_boundary.py treated it as unread, so input_coverage became partial and the v3 schema invariant rejected the projection. The BOUNDARY-INPUT-INCOMPLETE suppressor had the same disagreement, which crashed an unreadable legacy host policy beside such a key. A single predicate, parse_failure_kind_was_read(), backed by the existing _PARSEABLE_EVIDENCE_KINDS, now answers the question for the band and publication predicates, for _coverage_for, and for the BOUNDARY-INPUT-INCOMPLETE suppressor. Kinds that really were not read (json_parse_failed, toml_parse_failed, host_config_content_unresolved, ...) stay partial and unpublished; the schema invariant is untouched. Resulting authority: a key alone is require_review / agent_action_required; beside a scoped shell rule it is review_publishable; beside Bash(*) it still blocks; with an unreadable legacy host policy it adds BOUNDARY-INPUT-INCOMPLETE and stops for a human. None of these reaches allow, merge or report_complete. Audit ids and codex-boundary-json output are unchanged for input with no input issue; check --diff and MCP shipgate.check become comparable there. No new public surface. CHANGELOG and docs/host-boundary-support.md note that complete coverage means the file was read, not that every key's meaning is modelled. New tests in tests/test_check_unmodelled_host_config_keys.py cover the variant matrix across formats, modes and callers, audit-id parity, the unread-input and input-issue cases, and the shared predicate. Review: cycle 1 findings addressed in a follow-up commit (shared read/unread predicate, BOUNDARY-INPUT-INCOMPLETE suppressor path and its tests, documented intended changes); cycle 2 found no P0-P2 issues. Rewording the parse-failure check text, plugin-specific rules and a generic ValidationError path in check.py are deferred. Closes #810 Show named reusable-workflow secret remapping in host diffs (#693) (#806) * Show named reusable-workflow secret remapping in host diffs (#693) A job changed from `secrets: {credential: ${{ secrets.STAGING_TOKEN }}}` to `${{ secrets.PRODUCTION_TOKEN }}` produced no row, no limit and a comparable result on published 1.0.0 and on main: the reusable-workflow grant never modelled named secrets. Each reusable call on the workflow grant now lists `secret_mappings[]`: the called workflow's secret input and, for a whole-value `${{ secrets.NAME }}`, the source name. Values are never read. Adding or removing a destination, or pointing one at a different source, is one `changed` row naming `job/destination` with `expands: false`; its `why` says a name does not establish privilege, caller availability or downstream use. The same row reaches `diff`, `check`, manifest-free `verify`, `verifier.json` and the PR comment. Reordering entries or re-spacing/re-quoting the expression is quiet, and `secrets: inherit` keeps its own widening signal. A literal value, any other expression, a non-string, or a `secrets:` that is neither `inherit` nor a mapping publishes nothing of its value (no text, no digest) and records a blocking coverage issue, so a changed workflow refuses and an unchanged one is a named limit. The #767 collision rule is now one shared helper for every compared workflow text. A job's reusable `uses:` passes through the same redaction as a step reference (`uses_redacted`), as do secret destination and source names. A rewritten value records the blocking issue, so `org/repo/.github/workflows/x.yml@token=aaaaaaaa` -> `@token=bbbbbbbb` no longer compares as unchanged, and a token in a reusable target is no longer published verbatim. Host-grants 0.6 and contract 40 are extended in place: neither has shipped in a tagged release (published 1.0.0 is contract 39 / 0.5). A 0.4/0.5 baseline holding a reusable call also names `baseline_reusable_workflow_secret_mappings_unavailable`. The support page drops the #693 unread bullet and names the read and its limits; STABILITY gains a migration note. Co-Authored-By: Claude Opus 5 * Address review cycle 1 on reusable-workflow secret mappings (#693) An unreadable secret value recorded a blocking coverage issue, so one unchanged reusable call passing `${{ github.token }}`, a literal, or any other expression made `check` require review, GitHub coverage partial, `audit --host --save-baseline` exit 2 and drift incomplete — and any edit to that workflow made diff, verify, the PR comment and `check` incomparable with zero rows, hiding real widenings elsewhere in the pull request. An unread value now records a NON-blocking `unsupported` issue naming its `job/destination`, the way an unread `envFile` does. Coverage stays complete, so every route reads exactly as it does on published 1.0.0. The entries are still compared on (destination, form, unresolved_reason) with no value, so adding, removing or re-forming one is still a row; only an edit between two values of the same unreadable form is unreported, and the named limit says where. A redacting secret name or reusable target stays blocking: the #767 collision rule is unchanged. `${{ github.token }}` is read as the `GITHUB_TOKEN` source, since GitHub documents it as functionally equivalent to that secret, so migrating between the spellings is quiet. Source names compare case-insensitively, as GitHub references them; the callee's secret id is compared as written, because GitHub does not document it as case-insensitive, and the uppercase context spelling `${{ SECRETS.X }}` stays an unread expression for the same reason. The mapping row's closing sentence is scoped to the mapping, so it no longer appears to deny a real widening on the same row, and a row whose entry moved between named and unread says so rather than claiming a source name that was never seen. STABILITY, the contract page, `llms-full.txt`, the support page, the distribution-surfaces row and the CHANGELOG bullet state the final behaviour. No schema, version or discovery file moved. Co-Authored-By: Claude Opus 5 * Say where an uncompared secret mapping is named (#693) Review cycle 2 found two claims that no longer matched the behaviour this PR ships, both on surfaces someone reads to decide what a run proves. The `HostReusableWorkflowSecretV6` docstring still said every unresolved mapping records a blocking coverage issue — the pre-fix rule that flipped untouched repositories to require_review. Pydantic publishes that docstring verbatim as the JSON-Schema description, so both 0.6 schema files asserted it too. It now states the shipped rule: a redacted name blocks, because two values that redact alike must never compare as unchanged, and every other unresolved mapping records a non-blocking issue naming its job/destination. The distribution-surfaces `capability_diff` row claimed this surface leaves an unreadable value "as a named non-blocking limit". It names it nowhere: `unchanged_limits` is built from blocking issues only, so `diff`, `verify` and `check` carry no limit for it, exactly as on 1.0.0. The row and the support page now say that, and where the limit is named instead. `audit --host` also printed it as a "declared exclusion", telling the reader the repository had asked for it. An unsupported or unreadable surface now prints "not compared". Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Keep base-cache access inside its Git metadata namespace (#638) (#805) * Keep base-cache access inside its Git metadata namespace (#638) A directory link at the base-cache entry, base-scans or agents-shipgate below Git's metadata directory redirected report, checksum and capability-lock publication and pruning outside the cache: an external entry was replaced, a cold run created files there and pruned user data, a warm run reused an external entry, and verify then stopped with an internal error at artifact export. Every cache read, publication and prune now goes through one boundary, _base_cache_directory, anchored at the metadata directory Git selected. On POSIX it walks agents-shipgate/base-scans/ with O_NOFOLLOW | O_DIRECTORY relative to parent descriptors (creating inside the same descriptor), reads singly-linked regular files relative to the entry descriptor, publishes through an exclusive temporary renamed with src_dir_fd/dst_dir_fd, and prunes with scandir(fd) and rmtree(dir_fd=...). The walker and reader are factored out of read_regular_file_beneath rather than reimplemented. Where descriptor operations are missing (Windows) the boundary inspects each component lexically before pathname operations. A linked, non-directory or unopenable namespace component makes the cache unavailable for that run: the base is regenerated from Git into a run-scoped directory, nothing is read, written or pruned through the component, and a base note names the path to repair. Links at or above the Git-selected metadata directory (linked worktrees, a linked .git) keep caching. The remaining same-run pathname reopen and the Windows check-then-use window are tracked in #803. Co-Authored-By: Claude Opus 5 * Address review cycle 1 on the base-cache namespace (#638) The cache is now settled by one no-follow open per run, and that open returns the validated `report.json` bytes and the capability lock read through the same directory handle instead of a pathname to reopen. Those bytes are written into a directory private to the run, and the head scan's diff reference, the `verification-base-report.json` export, gap provenance and capability review open only that copy — on a warm hit, on a cold run that has just published the entry, and when the cache was unavailable. A namespace component replaced during the run can no longer be read back, in an ordinary checkout or a `git worktree add` layout. The published diff reference is the entry's cache location, spelled without resolving it, so it is the same string for warm, cold and unavailable runs; `report.json` stays byte-identical to main for ordinary warm and cold runs, and two cache-unavailable runs agree byte for byte instead of naming a deleted temporary directory. `replace_file_at` takes the creation mode, restoring `report.sha256` and `capabilities.lock.json` to the 0600 they have on main; `report.json` stays 0644. `_cache_report_valid` is gone: its callers use the one boundary that returns the bytes. Co-Authored-By: Claude Opus 5 * Publish the run's base report without following a link (#638) Review cycle 2 found the one write left in the new path that resolved a pathname: `keep_for_this_run` wrote the validated bytes with `write_bytes`, so a link planted at that name in the run's own directory was followed and its target clobbered. The directory is a fresh 0700 mkdtemp, so this needed a same-uid process racing an unguessable path, but "nothing is written through a link" is the whole claim of this change. Where descriptors allow it, the copy now goes through `replace_file_at` on an open handle for the run directory, at the same 0600 the cache uses. The lexical fallback keeps its pathname write, as on Windows today. A regression plants that link and asserts the target survives; it fails without the fix, clobbering the victim with the base report. Also corrects two statements review cycle 2 found untrue: the lexical replace cannot reproduce the descriptor boundary's mode (fchmod ignores the umask that masks the descriptor creation), and a cold run opens the cache three times, not once — the guarantee is that no namespace pathname is reopened after the cache step. The docs now say the published base reference is a location rather than a source, which is what lets two cache-unavailable runs agree byte for byte. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Resolve agent control's default reports directory under --workspace (#575) (#801) * Resolve agent control's default reports directory under --workspace (#575) `agents-shipgate agent control --workspace ` validated the workspace but read its default `--reports-dir` relative to the caller's current directory, while `verify --workspace ` publishes its default output under the workspace. From anywhere else a valid run read as `missing`, and a caller standing in another verified repository had that repository's pointer read and refused against the requested one. - One rule: `current_workspace.default_reports_dir(workspace)` is `agents-shipgate-reports` under the resolved workspace. `verify`'s `_resolve_out_dir` now delegates its default to it (behaviour unchanged), and `agent control` reads by it when `--reports-dir` is omitted. The leaf component is not resolved, so the reader's refusal of a symlinked reports directory still applies. - An explicit `--reports-dir` keeps its meaning: absolute as given, relative against the current directory. Nothing else is searched when it is absent. - The searched directory is made absolute before the read, so the `missing` refusal, the guidance, and the `verify --out` recovery all name a directory that resolves from any cwd (verify resolves a relative `--out` against the Git root, so the old relative echo named somewhere else). - Envelope artifact paths keep the caller's spelling: explicit paths as given; the default relative to the cwd when beneath it (so `--workspace .` output is byte-identical) and absolute otherwise. - `live_workspace` only passes the reports directory to Git as a change-set exclusion when it lies inside the repository. An outside directory (the #627 case recorded on #575) was refused by the Git helper, swallowed into `workspace_unverifiable`, and could never refresh. Docs: `--reports-dir` help, docs/agent-contract-current.md (and the regenerated llms-full.txt), CHANGELOG. Co-Authored-By: Claude Opus 5 * Pin the shared reports default and verify's outside-output control read (#575) Add a CLI test that verify --format control on a run published beside the repository returns its complete envelope (it withheld authority on main), pin DEFAULT_REPORTS_DIR to the published DEFAULT_PATHS spelling, and state the compatibility change precisely in the CHANGELOG: refusals under --workspace . now spell the searched directory absolutely. Co-Authored-By: Claude Opus 5 * Address review cycle 1 on the workspace reports directory (#575) An output directory outside the repository was still handed to Git as a change-set exclusion by the writing run: plain `verify` and the preview's worktree read, the pointer's worktree overlay, and the manifest-free host comparison. The Git helper refuses such an exclusion and every caller swallowed the refusal, so plain `verify --out ` exited 2 with "input directory capture is unavailable", `--preview` read an untracked file it had been run on as a change it never saw, and the host comparison went `incomparable` with no rows (#785). `cli/current_workspace.py::worktree_exclusion` is now the one rule for every `exclude=` built from an output directory. It drops the exclusion only when the directory is disjoint from the repository under both its given and its resolved spelling; inside, the root, or an ancestor is passed through unchanged, so in-repository exclusion is exactly as before and the root or a parent still reaches the helper's refusal. Dropping an exclusion can only surface more changed paths. Nits: the refusal guidance spells `verify` through the invocation policy; the host comparison and the rerun-command default use `default_reports_dir`; bootstrap uses the shared name. Co-Authored-By: Claude Opus 5 * Address review cycle 2 on the workspace reports directory (#575) One directory could be inside the repository for the writing run and outside it for the refresh that checks it. `verify` resolved `--out` before asking, `agent control` asked about the spelling it was given, and `worktree_exclusion` kept the exclusion when *either* spelling looked inside. An in-repository symlink pointing outside was therefore excluded for the reader and not for the writer; the Git helper refused the reader's exclusion ("must remain inside workspace"), `live_workspace` swallowed that into `changed_paths=None`, and the unseen-change test was skipped for every pointer that does not authorize completion. A `review_publishable` answer — commit, push and update_pr granted — stayed current across a tracked edit and a new untracked file. Both layers are fixed. Containment is now decided by physical identity. The reports path is resolved and its ancestors are compared with the repository root's `(st_dev, st_ino)`, so the answer is a property of the directory rather than of its spelling, and the writer and the reader cannot reach different ones. Inside the repository the exclusion is returned respelled beneath the root, which is the spelling the Git helpers can turn into a pathspec; that also makes a case-variant `--out` (`/REPO/rpt` for `repo`, where the filesystem folds case) the in-repository directory it physically is, instead of exit 3 "path changed identity while it was read". The root itself and its ancestors are still handed over, still refused, and still withhold authority, and a directory genuinely outside still loses its exclusion. Defence in depth: a currency test that could not be run now refuses for every pointer state. `changed_paths=None` raises `workspace_unverifiable` on the plan-bound worktree path, the plan-less preview path and the committed-tree path, and an overlay that cannot be recomputed no longer returns silently from the preview path. A live workspace carrying no ref resolver refuses the base check the same way. Only completion-authorizing pointers failed closed before. Tests: the symlink rule assertion is flipped — an in-repository spelling that resolves outside is disjoint and drops its exclusion — and four regressions are added: a refresh spelled through an in-repository link after edits, one string used for both `--out` and `--reports-dir` through a link out of the repository, a case-variant `--out` end to end (skipped on a case-sensitive filesystem), and a unit test that `changed_paths=None` refuses for both a `review_publishable` and a plan-less preview pointer. All seven fail on `c685856a` from a separate worktree; on `origin/main` every one of these setups exits 2 at the writer. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Distinguish discovered Claude hook files from proven host loading (#714) (#789) * Distinguish discovered Claude hook files from proven host loading (#714) A parsed `.claude/hooks/hooks.json` proves the file exists, not that Claude Code loads it; Claude Code documents no such project location. #704 published every such file, nested copies included, as an `execute`/`high` grant and an expansion, from its path alone. Hook grants now carry a loading basis, read back from existing published fields (no schema or contract change): - host configuration (Claude Code settings layers, Codex hooks.json): `execute`/`high` and an expansion signal, unchanged; - selected by a `.claude-plugin/plugin.json` in the repository (default `hooks/hooks.json`, a `./` path or array in `hooks`, or inline hooks): `execute`/`high`, event and command evidence kept, no expansion signal, because plugin installation/enablement is not in the repository; - selected by nothing: `access`/`risk` `unknown`, still a row, never an expansion. Plugin manifests and hook-named files are materialized in the scoped base tree so a selected plugin hook is compared on both sides, without joining the adapter registry, so `check` and the triggers route nothing new. Malformed references are blocking limits; a missing file is non-blocking. A hook file is no longer read as a settings file. Rows state the basis in `diff`, check rows, the Stop hook and PR output. The #704 oracle is corrected with dated provenance pinned to loopkit b3e55147/5ae033e6, where nothing selected the file. Co-Authored-By: Claude Opus 5 * Address review cycle 1 on the hook loading basis (#714) P1: `check` cannot name a limit and routes no plugin manifest, marketplace or plugin-selected hook file, so the limits those raise no longer enter its input completeness or its host comparison. The reader records their issue ids on the snapshot; `agent_boundary` skips them and `check` compares with them stripped. `diff`/`verify` unchanged: an untouched malformed manifest is a named unchanged limit, a newly malformed one refuses. An untouched malformed nested manifest + README-only head is `allow` again, and a deny-rule change keeps its row, as in 1.0.0 (text, agent-boundary-json and agent-control-json controls). P2-1: the basis is read only from the published pair. Plugin-selected hooks are `execute`/`medium`; a hook-file grant with neither pair (a 1.0.0 baseline's `execute`/`high`) is `unestablished` and claims no selection; every non-settings removal row is neutral. P2-2: `./`-sourced and `pluginRoot` marketplace entries select hooks through the same resolver (entry `hooks` path or inline object, `strict: false`, a manifest-less root's default `hooks/hooks.json`); marketplace limits never block. Loopkit's marketplace entry (no `hooks`/`strict`) is recorded. P2-3: a reference into a walk-skipped directory is named as unread, not missing, without blocking. P2-4: one-time drift note in the drift example and docs/baseline.md, including #771/PR #788. Nits: `hooks.json` matched at a name boundary (not `webhooks.json`); a selected file without a `hooks` object is named; subagent frontmatter hooks listed as unread; case-insensitive matching documented. Co-Authored-By: Claude Opus 5 * Address review cycle 2 on hook loading evidence (#714) - A hook a plugin selects is execute/high with an expansion signal when the repository's project settings enable @, register that marketplace as a relative directory or file source inside the repository, and the marketplace lists the plugin with an in-repository source. Remote, settings, absolute and escaping sources, unlisted plugins and non-true values keep it execute/medium with no expansion. - check drops only plugin-reference limits both sides share on an untouched source; a limit on one side only, or on a changed source, makes the host comparison incomparable instead of building a removed or added row. - A marketplace source containing a slash is not a bare name, an unusable metadata.pluginRoot is named without blocking, and a reference beneath a link leaving the workspace keeps only the link's blocking limit. - STABILITY.md, the support page and CHANGELOG state that read limits make diff and verify incomparable, as an oversize settings file does. Co-Authored-By: Claude Opus 5 * Address review cycle 3 on hook loading evidence (#714) P2: the docs claimed a 1.0.0 baseline of an enabled plugin's hook does not drift. Against a baseline published 1.0.0 actually saved, it does. The hook grant is unchanged — 1.0.0's execute/high, same grant_id, no typed change and no expansion signal — but the plugin manifest that selects it is a newly read artifact and a newly observed source, so `audit --host --drift --fail-on-drift` exits 20 once reporting "0 typed grant change(s)". STABILITY.md, docs/baseline.md, the drift example workflow and the CHANGELOG now say that, and point at `artifact_changes`/`coverage_changes` in `--json` and at re-saving from the reviewed default branch. `_recorded_by_1_0_0` rebuilt its "1.0.0" baseline from this reader's own inventory, so it held artifacts 1.0.0 never read and could never fail. It now refuses an inventory holding a plugin manifest or marketplace. The enabled-hook case uses tests/fixtures/host-grants-1.0.0-enabled-plugin.json, a baseline published agents-shipgate==1.0.0 saved for that workspace; a test pins what makes it a 1.0.0 artifact. Three tests replace the one that could not fail: the fixture's shape, the drift payload, and the gate exiting 20 through the CLI. The two other helper callers build workspaces without plugin files, where the synthesis is faithful. Nit: `enabledPlugins` matched a marketplace only by its `extraKnownMarketplaces` key. Claude Code's marketplace schema calls the `name` in `marketplace.json` the identifier users see after the `@`, and the settings documentation only ever shows the two equal, so neither is documented as the one the host matches. Both are now accepted, as case-insensitive reference matching already errs toward showing a hook. A `name` alone still registers nothing: with key `market`, name `real-market` and `demo@real-market: true`, the hook is again `⚠ high widened`, as on 1.0.0. Follow-ups filed, with the check limits recorded on the support page: #808 (a newly broken plugin manifest withholds unrelated rows), #809 (an enabled plugin's hook at a non-registry path gets an expanding row under `allow`) and Refs #808, #809, #810. Co-Authored-By: Claude Opus 5 * Say how to re-save a pre-0.6 host-grants baseline (#714) Review cycle 4 found that the #714 upgrade paragraphs end at "re-save the baseline", but `--save-baseline` refuses to overwrite any baseline older than 0.6 — which every baseline 1.0.0 wrote is. The reader of those paragraphs was never told to move the old file aside, so the documented recovery exited 2 and the gate kept failing. STABILITY.md, docs/baseline.md and the scheduled drift example now name the move-aside step and the refusal. Verified end to end against the checked-in 1.0.0 baseline fixture: save over it exits 2 with unsupported_baseline_schema, and after `git mv` the re-save exits 0 and the drift gate reports no drift. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Read workflow step action references (#771) (#788) * Read workflow step action references (#771) A workflow step's `uses:` never entered the workflow grant, so moving `actions/checkout` from a pinned SHA to `@main` produced no row, no limit and a comparable result on every route. The workflow grant now lists each step's remote (`owner/repo[/path]@ref`) and `docker://` reference with its job and step (`id`, else `name`, else `steps[N]`). Each job's multiset of references is compared, so an added, removed or changed reference is one `changed` row naming `job/step`, while a reorder or rename is quiet. A reference names code, not scopes: the row never widens, and permission direction is still decided by the permission contexts. The same row reaches `diff`, `check`, manifest-free `verify`, the PR comment and the control envelope; the Stop hook stays quiet because nothing widens. Local `./` actions stay unread (#701). Expressions, other strings and non-string values are listed as `unresolved` with a reason. A credential-shaped value is published redacted and records a blocking coverage issue, so two such values cannot compare as equal (#767). Host-grants schemas move to 0.6 and the runtime contract to 40. A 0.4/0.5 baseline holding a workflow grant is incomparable (`baseline_workflow_step_actions_unavailable`); one without a workflow stays comparable. The support page drops only the remote-reference limit, and the pilot ledger's source-tree column was re-measured. Co-Authored-By: Claude Opus 5 * Address review cycle 1 on step action references (#771) - P2-1: a job whose `steps` is not a list, or a step that is not a mapping, is listed as `unresolved` (`steps_not_a_list`, `step_not_a_mapping`) with no `uses` text, so an absent `step_actions` still means "read, none declared". - P2-2: step `uses` text and step labels pass through `privacy.redact_text` before the host sanitizer, so `ghp_…`, `AKIA…` and `xoxb-…` take the `redacted` + blocking-issue path. `docker://user:password@registry/…` userinfo is replaced whole; `docker://image@sha256:`, SHAs, tags and `owner/repo/path@ref` are never marked redacted. A CLI sweep proves no canary reaches inventory JSON, Markdown, drift, refused saves, `diff` or the baseline file. - P2-3: `llms-full.txt` rebuilt after the AGENTS.md schema-table label fix. - P2-4: the STABILITY migration note now states the incomparable drift, the preflight `high` human signal, the `--save-baseline` refusal and the move-aside/re-save steps from the reviewed default branch; a test drives every step through the CLI. - nit 1: a reference that moved between jobs says so instead of "names different code to run". Co-Authored-By: Claude Opus 5 * Address review cycle 2 on workflow step action references (#771) - P2-1: step-reference userinfo is everything before the last `@` once a trailing `@algorithm:hex` digest is set aside, so a `docker://` password holding `/`, `:` or `@` (a base64 key file) is redacted and blocks instead of being published verbatim. The same rule covers `DOCKER://`, other schemes, and scheme-less values whose text before the last `@` holds `:` or `@`. `owner/repo/path@ref`, tags, ports and digests stay as written (16 negative controls). Unit cases cover a `/`, `@` or `:` in the password and userinfo with a digest. The CLI canary sweep now commits the change and also checks manifest-free `verify` text, `pr-comment.md` and `verifier.json`. - nit-1: STABILITY says `--save-baseline` refuses any pre-0.6 baseline with exit 2, with or without a workflow grant, and to move it aside and re-save. A CLI test pins it for 0.4 and 0.5. - nit-3: `diff` text renders every field through `single_line_text`, which now also escapes U+2028/U+2029 and the bidi controls, so a newline, ESC or U+202E in a step name cannot forge a row. JSON keeps the exact text. - nit-4: token-shaped job names are filed as #802. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Refuse FIFOs in evidence readers before the open can block (#577) (#790) * Refuse FIFOs in evidence readers before the open can block (#577) read_regular_file_beneath opened the leaf without O_NONBLOCK and only then checked S_ISREG, so a FIFO with no writer hung the reader before its regular-file refusal. Open the leaf non-blocking (with O_BINARY, keeping O_NOFOLLOW and descriptor-relative containment) so the existing "is not a regular file" failure is reached; size limit, stability checks and bytes are unchanged for regular files. The sweep found the same open-before-type-check shape in the trust-policy reader and the local-review exclude reader, and two verify reads with no type check at all on paths an agent can place before the run (the declaration-continuation receipt and the cached base capability lock). Those now go through read_regular_file_beneath and still fail closed. Tests run every probe that could block in a child with a hard timeout. Co-Authored-By: Claude Opus 5 * Address review cycle 1 on non-regular evidence readers (#577) - Test the two behaviour changes the reroute brought with it: a symlinked cached base capability lock falls back with the invalid-lock note, a symlinked declaration-continuation receipt makes the continuation check return False, and a cached lock that is not UTF-8 falls back instead of raising UnicodeDecodeError. Each test first shows the same bytes accepted as a regular file. All three fail against origin/main's orchestrator. - Skip the directory refusal tests where their reader cannot run: dir_fd for the cached lock and continuation receipt, POSIX for the local-review exclude file. - Note at the 1 MiB receipt cap that exceeding it fails closed. - Read local_review's exclude file with an os.read loop, like the other two readers, instead of a buffered stream over an O_NONBLOCK descriptor. It still reads to EOF with no new size limit, and every error message is unchanged. - CHANGELOG: mention the symlink and non-UTF-8 outcomes in this PR's bullet. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Add a host-only advisory PR recipe with no manifest or baseline (#780) (#786) * Add a host-only advisory PR recipe with no manifest or baseline (#780) examples/github-actions/14-host-only-advisory-pr.yml runs the pinned Action on pull_request with contents: read, pull-requests: write and full history, no shipgate.yaml, failure policy or check. With no manifest the Action's verify takes its manifest-free host route and comments the same comparison diff makes. The examples README, integrations and quickstart send host-only repositories to it and explain comments, permissions, forks and limits. A test feeds the recipe's inputs through action.yml's own run step on manifest-free PR branches. Co-Authored-By: Claude Opus 5 * Address review cycle 1 on the host-only advisory recipe (#780) P1: the Action imported Python from the pull request's checkout. Every composite step runs in GITHUB_WORKSPACE; `python -m pip` put it first on sys.path, so a PR adding pip/__main__.py ran before the engine was installed, the local-wheel installer's own `sys.executable -m pip` did the same, and the merge-verdict step's `python - < * Address review cycle 2 on the host-only advisory recipe (#780) - State the trust boundary: on pull_request GitHub runs the workflow file from the PR's merge commit, fork PRs included, so the job's output is advisory output of a job the PR controls. -P closes the in-Action import route; it does not make the result independent of the PR. - Make the failure list explicitly non-exhaustive and add setup-python, artifact upload, timeout and cancellation; describe concurrency accurately (always() steps on a cancelled run, manual re-runs of an older run). - The quickstart's manifest recipe no longer says it never fails the job. - The static Python rule checks every interpreter on a line, including absolute and ${pythonLocation} paths and chained commands, with positive and negative controls. - Fix the narrowing test's comment; restore the recipe CHANGELOG bullet the rebase onto main dropped. Co-Authored-By: Claude Opus 5 * Address review cycle 3 on the host-only advisory recipe (#780) - README's first task and the quickstart's host-change Next list now hand off to examples/github-actions/14-host-only-advisory-pr.yml after local diff value; a test asserts both links. - The static Python rule also catches versioned interpreters (python3.12, toolcache paths). - The examples README says which recipes need a manifest (12 needs a host-grant baseline, 14 neither) and that fork approval is the default. - The recipe header names only the install step as exposed in this recipe. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Lead the entry pages with the published host diff (#779) (#784) * Lead the entry pages with the published host diff (#779) The README's first task was a constructed tool-surface fixture and the quickstart never showed `shipgate diff`, the manifest-free comparison 1.0.0 ships. The README and distribution plan also described the advisory line as preview-only, every `v*` tag as qualified, and 1.0.0 as predating contract-33 discovery. README and quickstart now start from a real permission/MCP PR and `agents-shipgate diff`, quoting its three answers and the `--base` recovery from the published 1.0.0 installed outside a checkout; a test holds each quoted block to this tree's output. Current release statements follow .github/release-channels.json. The pilot runbook and ledger record #571's 2026-09-14 advisory-channel decision and freeze the Git-backed Route H artifact mapping, keeping the earlier decision and definitions. Co-Authored-By: Claude Opus 5 * Address review cycle 1 on the host-diff entry pages (#779) - Document the fourth answer shape: a no-change or changed answer that lists `Not compared:` sources neither side could read. README and quickstart no longer call such an answer fully covered, the quickstart quotes it, and the pilot records `unchanged_limits` as the result's coverage limit. - Re-capture every quoted answer from the published 1.0.0 in a clone with a remote, state which refs base detection reads, tell users to pass --base for a non-default PR base, give the single-branch refspec recovery, and narrow the shallow-clone statement to an unreachable merge base. - Move the pilot runbook's executable Route H sections (commands, read order, partner prompt, evidence, first touch, tracker) to Git-backed `diff`, keeping the baseline recipe as the optional/earlier route; fix the stale envelope sentence; add separate installation, comparison and reading times. - The 2026-09-14 decision takes its terminal decision when the cohort closes. - README line 16 no longer promises a merge answer; the status block adds comparable coverage 41/50. - The entry-docs test rebuilds a bare remote and clone, checks every quoted shape and the documented recovery, and reads the channel table rows. Co-Authored-By: Claude Opus 5 * Address review cycle 2 on the host-diff entry pages (#779) - Make the entry-docs test independent of CI terminal state and the caller's git: CLI output is unstyled with color and width pinned, git runs with no inherited GIT_* variables, global config or hooks. It failed on CI because Rich colored the error panel under GITHUB_ACTIONS. - The channel guard now rejects the exact pre-#779 qualified cell and checks the cadence sentence. - Quickstart: base detection fails only for a single-branch clone or a checkout without origin/HEAD whose default branch is neither main nor master; a local main or master is used without a remote; v0.15.0 has no `diff`, so upgrade first. - Pilot runbook: a Git-backed first valid result must touch a surface diff reads; the pip fallback floor is >=1.0; the partner prompt's baseline steps are nested under an explicit opt-in (3a-3c); the tracker separates Git-backed H, baseline H and A; the check pointer, step reference and Route A follow-up questions are corrected; the ledger's reproduction step names diff. Co-Authored-By: Claude Opus 5 * Address review cycle 3 on the host-diff entry pages (#779) - The entry-docs test no longer imports click, which the locked CI environment does not install; ANSI escapes are stripped with the regex tests/test_verify_auto_base.py already uses. The channel declaration is read by absolute path. - README, quickstart and the pilot runbook tell a fork clone to fetch upstream and pass --base upstream/, because origin is the fork. - The runbook's build section says v0.15.0 has no `diff`. Co-Authored-By: Claude Opus 5 * Drop the duplicated #779 CHANGELOG bullet left by the rebase Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Render the release engine into the optional adoption kits (#781) (#783) * Render the release engine into the optional adoption kits (#781) A final wheel is built before the published-release constants move, so it carries the previous release in them. The skill kits' CI recipes and two prompts held hand-written literal pins, and the runner pins and contract-floor sentence rendered from those constants, so the published 1.0.0 told adopters to install 0.15.0 and that no release reports contract 21. release_source.release_engine() now selects one engine for every pin init writes: a valid release-source record gives the wheel's own version, Action source SHA and emitted contract; no record keeps the published fallback; a malformed record raises. init --ci and the kit renderer both use it. The kit templates carry no literal pins. AGENTS_SHIPGATE_WORKFLOW_REF keeps its scope to the generated workflow. The distribution smoke checks the installed candidate's kit pins, and an installed-wheel test exercises the real init route with pre-publication constants. Co-Authored-By: Claude Opus 5 * Address review cycle 1 on the adoption-kit release engine (#781) - add-shipgate-to-repo.md told agents to expect `@v…` in the generated workflow, which a stamped wheel writes as a SHA; the line is now rendered from the release engine, and `v…` is no longer an exempt reader blank in the pin sweep, the kit tests or the smoke check. - init refuses a malformed release-source record before writing any file, as an agent-mode config_error, whenever --ci or a skill kit is requested. - The kit renderer resolves the engine once per render. - The smoke's kit check has a unit test for acceptance and each refusal. - The distribution-surface registry registers the stamped-wheel proofs on the adoption_kits row, and the release-channels row names the kit. Co-Authored-By: Claude Opus 5 * Address review cycle 2 on the adoption-kit release engine (#781) - The add-shipgate prompt's workflow check no longer names one concrete ref. Its checked-in copies render from source but pin a release wheel whose init --ci writes a source SHA, so the line now accepts a release tag or a release wheel's full source commit; a test writes both and checks the line. - The malformed-record refusal test asserts the agent-mode config_error line. - The smoke kit check refuses an unrendered template placeholder. - The refusal text no longer repeats the reinstall instruction. Co-Authored-By: Claude Opus 5 * Name the stamped-wheel pins in the #506 gap rows (#781) Review cycle 3 nit: the closed emitted-workflow and rendered-prompt gap rows described only ordinary builds. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 docs: sequence post-1.0 adoption around real reviewer value (#782) Move the published-release pins to v1.0.0 (#777) v1.0.0 is published (PyPI wheel d34012ca, GitHub Release public), but the repository still named v0.15.0 (contract 10) as the newest release: Action refs, pip/uvx pins and shipgate_version inputs in examples, docs, skills, plugins and the rendered adoption prompts; .well-known, llms.txt, the README, the quickstart channel table and the FAQ. - LATEST_PUBLISHED_VERSION = "1.0.0", LATEST_PUBLISHED_CONTRACT_VERSION = "39". - Every pin the tests enumerate names 1.0.0; adoption prompts re-rendered by the package's renderer, with render hashes updated and outgoing hashes kept. - Statements about the newest release describe v1.0.0 (advisory, contract 39, no qualification claim); v0.15.0 comparisons are relabelled "the previous release", and historical records are untouched. - The check-run example pins v1.0.0, whose action.yml defines check_run_policy. - Quickstart step 3 drops its v0.15.0-only PR comment excerpt. - The pilot ledger's Route H dry run was re-measured on PyPI 1.0.0 and the source tree (identical cells); the standing decision stays narrow with a dated factual checkpoint. - The rendering-rule test skips only while source and published versions are equal; the format-reader test proves non-vacuity on the v0.15.0 tag. Co-authored-by: Claude Opus 5 Fold post-cut changes into 1.0.0 and record the final engine's runs of record (#776) * Fold the post-cut changes into the 1.0.0 release section Five changes merged after the `## 1.0.0` section was cut, and `## Unreleased` is not part of the release body, so a tag cut now would publish notes that omit them: #767, #768, #713, #581, and the support-page entries for #693, `### Changes`, and their reviewed prose goes into `docs/changelog/1.0.0.md`. `scripts/release_notes.py --tag v1.0.0` extracts the section at 27,300 characters, against the 125,000 limit. Co-Authored-By: Claude Opus 5 * Record the final 1.0.0 engine's live measurements as the runs of record record, so the three harnesses ran again on the wheel Release Engine Smoke built and exercised on 747d6080 (run 34898276101, sha256 1b846258). No outcome moved: host-config and the ten MCP servers reproduce e5ec2311's scores.csv byte for byte, and cold start matches every outcome column at 26 of 30. Each runs.json was masked with mask_local_paths and re-scored byte for byte. The replay and findings-table tests now point at these directories, and the 1.0.0 release note names the new wheel. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Build the MCP server on the SDK 2.x MCPServer (#713) (#773) * Build the MCP server on the SDK 2.x MCPServer (#713) The [mcp] extra has installed MCP Python SDK 2.x since #354, but the server still imported mcp.server.fastmcp.FastMCP. SDK 2.0 renamed it to MCPServer and left mcp.server.fastmcp as a module that raises, so `mcp-serve` could not start on any install of the declared extra, and the error told the user to install the extra they already had. - Import mcp.server.mcpserver.MCPServer and pass snake_case ToolAnnotations. - Tell a missing SDK apart from an SDK without MCPServer, and name the installed version and required range for the second. - Declare the floor as 2.0.0, the first release with MCPServer, and hold the server's named range equal to pyproject's in a test. - Add a CI job that installs the extra at its floor and its newest admitted release, runs the MCP tests with the real SDK, and lists the five read-only tools over stdio from the installed command. Closes #713. Co-Authored-By: Claude Opus 5 * Import installed-version metadata the way the static-only lint allows (#713) `from importlib import metadata` is an import from the forbidden `importlib` module; `from importlib.metadata import ...`, which environment.py already uses, is the permitted metadata module. The trust-model invariant lint in CI rejected the first spelling. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Read Git pathnames that contain spaces (#581) (#774) Git never quotes a space in a pathname, but the shared unified-diff parser split every unquoted `diff --git` header at its first space, and read `---`/`+++`/`rename` values the same way. Any change touching a path such as `docs/new scope/notes.md` became a record whose paths were the `\0invalid-diff-path` sentinel. `check` published that sentinel as a changed file and asked for review, and a scoped manifest under a spaced directory crashed `verify` in `os.lstat` on the embedded NUL. - Resolve an unquoted header pair the way `git apply` does: by the one split whose halves name the same file. A single-space header keeps its only split; a rename or copy between spaced names is named by its rename/copy lines, and a record nothing names is refused, never guessed. - Take an unquoted `---`/`+++`/`rename`/`copy` value whole, up to the first TAB, which Git never leaves unquoted. - A path no filesystem entry can have is not the configured manifest, so `is_configured_manifest` no longer raises on it. Tests run real `git diff` output with both `core.quotePath` settings, plus `check` with and without a manifest and the issue's `verify` reproduction against its no-space control. Closes #581. Co-authored-by: Claude Opus 5 Name three more unread host surfaces and the link refusal rule (#775) The support page listed two host surfaces no adapter reads (#701, #702). An adversarial probe of `shipgate diff` during the 1.0 readiness review found three more that also change what runs, or what it can reach, with no row and no coverage limit, and the page did not say which symlinks refuse a comparison. - A workflow step's action reference, e.g. a pinned SHA moved to `@main` (#771). - A named secret mapping on a reusable-workflow call (#693). - The path of a remote MCP server's URL, excluded from the digest because it can carry a secret (#772). - A hook row describes the file, not proof that Claude Code loads it (#714). - A dangling link, a link leaving the repository, or a directory link outside the boundary paths refuses the whole comparison (#688, #659). The pin test now also runs the page's own examples: each named surface must still produce no row, and each fixture must produce one row when a field the reader covers changes. A reader fix therefore fails the test until its page entry is removed. Co-authored-by: Claude Opus 5 Expose malformed Claude permission shapes as incomplete coverage (#768) (#770) Preserve permission semantics through comparison redaction (#767) (#769) Exercise the #659 oracle with controls a bad engine fails (#766) Co-authored-by: Claude Opus 5 Re-derive the release verification suite budget for today's suite (#765) Co-authored-by: Claude Opus 5 Record the 1.0.0 candidate's live measurements as the runs of record (#764) Host-config, cold-start and the ten MCP servers, run on the wheel Release Engine Smoke exercised on e5ec2311 (sha256 071abe4f). runs.json is masked with mask_local_paths, and a test refuses a committed run with an unmasked path. Co-authored-by: Claude Opus 5 Name composite actions and hook-run scripts as unread host surfaces (#701, #702) (#763) Co-authored-by: Claude Opus 5 Remove the runbook's unverified manual undraft (#618) (#762) Co-authored-by: Claude Opus 5 Make the rehearsal provenance drill reach the payload (#615) (#761) Co-authored-by: Claude Opus 5 Name every kind of host expansion in preflight's explanation (#681) (#760) Co-authored-by: Claude Opus 5 Publish as 1.0.0 on the advisory channel (#759) * Publish as 1.0.0 on the advisory channel Bump the package, plugin manifests and discovery metadata from 0.16.0 to 1.0.0. 0.16.0 was prepared but never published, so its CHANGELOG section and migration notes ship in 1.0.0: the section is folded into ## 1.0.0 and the notes are restamped. The post-0.16.0 prose moves to docs/changelog/1.0.0.md. STABILITY.md states the 1.x line. Seven migration notes that carried no version, five of them filed under "Reporting a contract violation", are stamped 1.0.0 and gathered with the rest; their old anchors still resolve. codex-boundary-json and legacy policy discovery were documented to last "through 0.16.x", a line that never shipped; both now stay through 1.x. Co-Authored-By: Claude Opus 5 * Cite the committed runs in the 1.0.0 highlights Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Add the ten-server MCP false-finding table (#658) (#758) benchmark/mcp-servers/ pins ten public MCP servers chosen by a rule written from committed text (#658's five named servers, the survey table's other three, two reference servers in README order), runs a candidate over each with one bound mcp_server_source manifest, and records every published finding plus what the reader could not establish. labels.csv judges each finding against the server's source; score.py fails on an unlabelled or stale label; the new test re-scores the committed run. Candidate 9946f80a: 76 findings, 1 false (1.3%, under the 2% bar). Co-authored-by: Claude Opus 5 Leave eval-harness and mock tools out of MCP source catalogs (#658) (#757) mongodb-mcp-server keeps its eval runner's judge tools (Vercel ai SDK tool() wrappers with a static toolName) under packages/eval-tests and test fakes under integration-tests/src/mocks. The source reader enumerated all 11 as server tools, and five raised SHIP-DOC-MISSING-DESCRIPTION. Both readers now treat eval-tests, integration-tests and mocks as test directories. Co-authored-by: Claude Opus 5 Contradict readOnlyHint only with modifying, action evidence (#658) (#756) Bound scans of github-mcp-server and terraform-mcp-server raised 35 false SHIP-MCP-ANNOTATION-CONTRADICTION findings on read-only tools: topic nouns (secret, message, terraform) raised privileged, communication and infrastructure evidence the check read as a contradiction. The MCP hint promises no modification, so only a modifying effect contradicts it now, and a keyword inference counts only when the tool's name or description carries an action verb. No name-shaped reading is used (#419). A retrieval tool whose description names deletes still contradicts, pinned as a known limit. Co-authored-by: Claude Opus 5 Stop reporting an unreadable MCP source description as missing (#658) (#755) SHIP-DOC-MISSING-DESCRIPTION read an empty Tool.description, and on mcp_server_source a description is empty whenever the reader cannot resolve it: mcp-grafana's positional MustTool descriptions, literal concatenations, template literals with substitutions, f-strings and variables. 148 of 150 such findings on six public servers were false. Both readers now set RegistrationSite.description_unresolved when a description is written in a form they cannot resolve, mcp_server_source records it in the tool's extraction, and the check skips it. An absent or empty description is still reported. MustTool's literal second argument is now read as its description. Co-authored-by: Claude Opus 5 Read .vscode/mcp.json as JSON with comments (#659) (#754) VS Code documents and runs mcp.json with comments and trailing commas. The host readers parsed it as strict JSON, so a commented file in the host-config benchmark (amelioro/ameliorate#936) was parse_failed and refused the whole comparison. loads_jsonc blanks comments and a trailing comma outside strings, preserving offsets, for .vscode/mcp.json only, in the audit parse cache and the boundary check. The benchmark oracle reads the file the same way with its own implementation; the case now replays as a named widening. Co-authored-by: Claude Opus 5 Re-measure cold start and host-config precision on candidate 9d7d145d (#660, #659) (#753) Cold start: 26/30 against a bar of 24 (25 on 294d0443). Host-config: precision 69/69, widening recall 54/64 and benign zero-row 5/6, unchanged; comparable 40/50 (36). #748 made four more cases comparable, none holding an expected widening. Each remaining refusal is mapped: one .vscode/mcp.json with comments, and links whose created paths reach no registered host location. The host-config replay test re-scores the new run of record. Co-authored-by: Claude Opus 5 Load CLI commands on demand so the Stop hook's diff fits its budget (#661) (#752) Owner decision D4(b): meet the 1.5 s Stop fast path by loading subcommands on demand, not with a separate diff entry point. Importing agents_shipgate.cli.main imported all 34 root commands (0.67 s) before any ran. The root group now lists every command from one table and imports a command only when Click resolves it; help, --help-all, completion and typo suggestions walk the same table. Two package imports then loaded the same machinery back into diff: require_workspace imported cli._helpers on the success path, and importing cli.verify.git ran cli/verify/__init__.py, which imported the verify command and orchestrator. The guard now imports its reporting helpers only to refuse, and the package resolves verify and run_verify on first use. Stop hook diff route on a manifest-free repository: 1.57-1.73 s before, 0.90-1.00 s after; shipgate diff alone 1.15 s before, 0.55 s after. Tests that read app.registered_commands or a group's commands dict now use the Click group API that help and resolution use. Co-authored-by: Claude Opus 5 Stop repeating hook advisories within a session (#661) (#751) A scripted 50-event session of docs, test, README and source edits plus a supported settings narrowing drew 26 interrupts in a repository without a manifest. PostToolUse and Stop repeated the withheld-verdict advisory after every source edit, and a narrowing beside other edits fell back to the manifest advice instead of the host comparison. The hooks now record, per session and path, the last verdict and the verdicts already announced, and speak only when an edit brings a new one. PostToolUse names the edited paths; Stop names the changed paths that were not last found quiet or that no hook evaluated. A new session, base or manifest starts over, and a verdict withheld for unreadable input is always announced. Without a manifest, Stop compares host configuration with the host readers even beside other edits, evaluates the rest on a diff without the host sections, and speaks once. PostToolUse leaves host configuration to Stop in both modes. Co-authored-by: Claude Opus 5 Let a decided host settings narrowing finish without human review (#661) (#750) In an adopted repository, tightening .claude/settings.json from Bash(*) to Bash(git status) raised no expansion after #745, but check still added PROTECTED-SURFACE-UNCLASSIFIED for the protected change and verify raised TRUST-ROOT-TOUCHED, so the Stop hook ended the turn on a human review. Owner decision D1(a): a host settings change whose only effect is a narrowing the permission lattice decides has been evaluated completely. The host boundary evaluator records it as the host_settings_narrowed diagnostic, the boundary treats the path as evaluated, and the trust-root touch clears. A hook, any other key, a malformed rule list, or a rule the lattice cannot decide keeps the review. Co-authored-by: Claude Opus 5 docs: publish changed-file paths verbatim, by decision (#742) (#749) Co-authored-by: Claude Opus 5 Read through in-tree links at boundary paths (#700) (#748) Step two of #700, by owner decision. A boundary path that is a symlink resolving inside the repository is read at its target and published under its own path. That covers a file link a host adapter names, such as CLAUDE.md -> AGENTS.md, and a directory link at a fixed-prefix boundary location, such as .claude/skills -> ../.agents/skills. Each artifact records the hops in resolved_through, and the resolution is bound to the identity-bound read session, so a retargeted link fails the snapshot. A scoped base tree now holds those targets' bytes, so both sides of a comparison read the same file. External, escaping and dangling targets, links inside linked directories, targets under skipped directories, chains past eight hops, and directory links that could only hide a **/ match stay coverage limits. Host-grants schemas move to 0.5 and the runtime contract to 39. Drift still compares a 0.4 baseline, which could never hold a read-through artifact. Co-authored-by: Claude Opus 5 Read Go MCP servers' literal tool hints for the contradiction check (#658) (#747) Both readers now read the go-sdk struct field (Annotations: &mcp.ToolAnnotations{ReadOnlyHint: true}, as github-mcp-server writes it) and the mcp-go hint options (WithReadOnlyHintAnnotation / WithDestructiveHintAnnotation), in SDK order. Only a bare true or false counts. Pointer helpers, WithToolAnnotation, variables and positional fields refuse the whole value. The hints reach SHIP-MCP-ANNOTATION-CONTRADICTION as claims, and the existing gate keeps them out of every effect reading. There are 19 shared corpus cases, plus an end-to-end scan. Co-authored-by: Claude Opus 5 fix(boundary): route a blocked result to human review instead of crashing the envelope (#728) * fix(boundary): route a blocked result to human review instead of crashing the envelope * fix(boundary): keep a policy block's stop ahead of the undeclared-surface route (#694) The undeclared-surface branch in _control_for_result ran before the block branch, so a blocked result was routed to a coding-agent command that could publish. Every projection read that control: text, agent-boundary-json and agent-control-json crashed on the schema invariant, and codex-boundary-json told the agent to run a configuration command for a blocked result. Guard the route at its source and drop the envelope-only reconciliation, so all formats agree. Keep the regression test from #728 and add one through the check command in all four formats. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: wangzhengzhuo05 <175673456+wangzhengzhuo05@users.noreply.github.com> Co-authored-by: Pengfei Hu Co-authored-by: Claude Opus 5 Re-measure cold start and host-config precision on candidate 294d0443 (#660, #659) (#746) Cold start: 25/30 correct covered comparisons, against a bar of 24 (17 on 8d43106f). Host-config: precision 65/65, widening recall 54/64 and benign zero-row 5/6 (42/63 and 5/6 on 71ef771d). Every miss is a refused comparison, and each refusal is mapped to its cause. Nearly all are in-tree symlinks at boundary paths, which is #700 step two. The host-config replay test now re-scores the new run of record. Co-authored-by: Claude Opus 5 Stop calling a narrowed allow rule an expanded allowlist in check and verify (#661) (#745) The host boundary evaluator took the set difference of allow rules, so tightening Bash(*) to Bash(git status) raised PERMISSION-ALLOW-EXPANDED and, in an adopted repository, the Stop hook handed the narrowing to a human as an expansion. Skip an added rule an old rule already subsumes, read from the permission lattice the drift reader and audit table already use (#657), for Claude Code settings and the Cursor CLI config. An undecided pair stays an expansion. Co-authored-by: Claude Opus 5 Keep credential-shaped path bytes out of host inventory output (#590) (#744) A host inventory published each source's exact path in artifacts, grants, coverage, baseline, drift and verify's host comparison, so a token-shaped directory or file name reached all of them. public_host_path redacts each path component with the shared sanitizers and, when anything was redacted, stamps a digest of the exact source so two sources that redact alike stay distinct. Ids, reads and policy classification keep the exact path; a path with nothing to redact is unchanged. Issue sources use the same label, bounded without being re-sanitized, and bind their id to the exact source whenever the shown form is lossy. Co-authored-by: Claude Opus 5 Name the host capability change in the compact control envelope (#662) (#743) shipgate.agent_control/v1 gains an optional, bounded capability_rows block on check --format agent-control-json, verify --format control and agent control: up to five rows with widenings first, the count cut, the comparison status and reasons, and the unchanged-limit count. It is a copy of the rows the producer already published and moves no state, permission or route. It is omitted when no comparison ran. The envelope is a closed object pinned by hash, so this widens it in place by owner decision: a v37 validator rejects a row-bearing envelope. Contract 38; the minimum control contract stays 21. Co-authored-by: Claude Opus 5 Support .vscode/mcp.json as a first-class host surface (#731) (#738) Experimental coverage refused every host comparison whenever the file changed (0 of 6 PRs in #659). The server set is compared as for .mcp.json; sandbox and sandboxEnabled are sandbox grants; an ${input:...} reference contributes its name to the digest, never a value; envFile is a non-blocking limit; other top-level keys keep coverage partial. The trigger catalog's adapter projection follows the registry. Five of the six #659 replays now compare. Owner decision recorded on #731. Co-authored-by: Claude Opus 5 Digest undocumented skill frontmatter keys and read frontmatter-less skills (#730) (#739) Undocumented keys made a skill unresolved, which refused every host comparison in its repository (4 of 50 PRs in #659). They now enter the structure digest as written, so changing one is still a change and none is read as a permission; an unquoted YAML date is digested as its ISO text. A skill without frontmatter takes the documented defaults. Cursor rules still refuse an undocumented key. Owner decision recorded on #730. Co-authored-by: Claude Opus 5 Treat an in-tree link to a file as a readable host file, not a coverage limit (#700) (#741) A symlink whose target resolves lexically inside the repository to a regular file is no longer a symlink coverage limit. Its type comes from the identity-bound read session's enumeration and is revalidated when the session finishes, so swapping the target for a directory fails the snapshot. External and dangling links stay blocking. verify's scoped base tree materializes the link target's type the same way. Co-authored-by: Claude Opus 5 Compare host configuration in the Stop hook when no manifest exists (#661) (#740) A narrowing and a widening host-config edit got identical advice to initialize a manifest, naming neither change. When every changed file is host configuration and no manifest exists, the Stop hook now runs shipgate diff: quiet when no row expands, each widening row named once (bound to the change's signature, base and rows), never quiet when it cannot compare. The PostToolUse hook stops nudging on those edits. Mixed changes and manifest repositories keep their route. The path list is rendered from the boundary registry. Owner decisions recorded on #661. Co-authored-by: Claude Opus 5 Compare past unchanged partial and experimental limits (#721) (#727) * fix(host): compare past unchanged partial and experimental limits (#721) A host comparison refused whenever either inventory was incomplete, even when the surface that made it so was untouched by the change. In the #660 cold start, 11 of 30 public repositories stopped that way: one unresolved skill, one .vscode/mcp.json, or one symlink blocked every comparison. diff and verify now share one comparison. It names a limit instead of refusing when the limit is present on both sides, is a per-source unsupported or parse_failed issue or experimental coverage, and is byte-identical by Git object ID (or unfiltered hash for a working tree); git diff is not used, since an eol filter can hide a CRLF rewrite. Everything else still refuses: a changed limit, a one-sided limit, and unreadable sources. An unchanged symlink with a changed in-tree target would otherwise hide that change, which is #700's decision. The limits are published in host_comparison.unchanged_limits (verifier 0.19, runtime contract 37) and in diff --json (capability diff 0.2), and rendered in text and PR comments. check's boundary result v3 cannot carry a limit, so check keeps refusing with unchanged_limits_not_representable. A 0.18 verifier reads as 0.19 with no limits, and one claiming limits is refused. audit --host --save-baseline still refuses an incomplete inventory. Seventeen tests, and eight mutations that each fail one. The static-only allowlist is re-pointed at git.py's two moved subprocess call sites; the calls themselves are unchanged. Co-Authored-By: Claude Opus 5 * docs(pilot): re-measure the Route H dry run at contract 37 (#721) The source-tree column was rerun on 53b28313 rather than carried forward. Every cell reproduced unchanged except the contract number: - check: block/critical, 4 violations - init: not_applicable_host_review, no workflow - manifest-free verify: exit 0, six rows - drift: four expansion signals Co-Authored-By: Claude Opus 5 * Re-record host-config replays for the shared incomparable reasons (#721) diff now reports an incomplete inventory with the comparison's shared reasons: an incomplete head is head_inventory_incomplete, as verify names it, and both sides are named when both are incomplete. Six .vscode/mcp.json replays of the #659 harness change only those reasons; every stop point, row and score is unchanged. The run of record stays as measured. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Add 12 scripted route-parity cases for host-capability changes (#662) (#737) One widening fixture is driven through shipgate diff, check's boundary JSON, the verify PR route and the MCP shipgate.check tool; every route must name the same five changes. Controls cover covered no-change, narrowing, malformed input, a refused boundary link, a second change, authority, and the generated AGENTS.md command run as printed. Scripted engineering tests only, never evidence of unprompted provider behaviour. Co-authored-by: Claude Opus 5 Name the change first in the generated AGENTS.md and CLAUDE.md blocks (#662) (#736) The Cursor rule tells an agent to run shipgate diff and quote the rows before routing on control; the AGENTS.md and CLAUDE.md managed blocks routed only on control.state, so two of three maintained copies could end a turn with "a human must review" and no named change. Both blocks now share one diff-first paragraph, and a test pins every copy to name the change before the control contract. Co-authored-by: Claude Opus 5 Read a Cursor rule's bare globs the way Cursor writes them (#729) (#735) Cursor documents globs as unquoted, comma-separated patterns, and one that begins with * is YAML's alias indicator, so the rule was refused as invalid YAML and the repository's comparison with it. A bare top-level globs value is now read as its literal string before either parse; editing it is still a change and an alias anywhere else still refuses. Co-authored-by: Claude Opus 5 Read Claude Code extraKnownMarketplaces so a marketplace change produces a row (#720) (#734) * Read Claude Code extraKnownMarketplaces so a marketplace change produces a row (#720) enabledPlugins entries install from the marketplaces a project trusts, but the adapter read only the plugins: the #660 cold start saw myplanet add a marketplace beside its plugin and named only the plugin. Each marketplace is now a plugin_or_app grant named marketplace: with its source as config, so adding or re-pointing one is an expansion and removing one is not. No schema change. The cold-start oracle counts the key as supported and the myplanet case is re-recorded as a covered success. Co-Authored-By: Claude Opus 5 * Re-record the host-config replay for a marketplace now read (#720) The #659 harness scores Claude settings through the cold-start oracle, which now counts extraKnownMarketplaces as supported, so the replayed outcome of the case that changed one moves in the same change. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Read FastMCP literal MCP hints from source as claims, never evidence (#658) (#732) * Read FastMCP literal MCP hints from source as claims, never evidence (#658) The source route never read readOnlyHint/destructiveHint, so SHIP-MCP-ANNOTATION-CONTRADICTION had nothing to challenge on a FastMCP server. Both MCP readers (the package reader and the zero-install port) now keep those two hints when written as exact boolean literals in a dict literal or a ToolAnnotations(...) bound to mcp.types; anything else refuses the value as annotations_unresolved. Reading them as the export route does lowered transfer_funds from write to read with no finding and closed three of eight open effect questions, so core.domain.annotation_hints_are_effect_evidence keeps source-read hints out of risk_hints and semantic_assessment. The contradiction check still reads them. A scan-level A/B test pins that only contradiction findings change. Co-Authored-By: Claude Opus 5 * Keep a source-read MCP hint off a merged identity's primary (#658) A reviewed tool_identity binding copies member annotations onto the merged tool, which keeps the primary's source type. With an exported primary and an mcp_server_source member, a source-read readOnlyHint would have become the export's published hint and counted as effect evidence past the gate. The merge now leaves those hints on their own observation and still reports a disagreement. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Measure diff precision and recall on 50 merged host-config PRs (#659) (#733) Add benchmark/host-config/: a frozen population of 50 merged public PRs across six host-config file kinds, an engine-independent oracle, a driver that also records audit --host for each refusal, a scorer, and an offline replay pinned in tests/test_host_config_replay.py. Run of record, candidate 71ef771d: precision 52/52, widening recall 42/63 (bar 0.90), benign zero-row 5/6 (bar 95%), comparable 25/50. Every miss is a refused comparison, mapped to #700, #721, #729, #730 and #731. Co-authored-by: Claude Opus 5 fix(host): make an MCP server URL's query part of its change digest (#723) (#726) The change digest hashed only the redacted URL, which keeps the scheme and host and drops the path and query. So removing read_only=true from a Supabase MCP URL, adding features= or switching project_ref produced no row. The #660 cold start found the features case on BigSimmo/Database. The query now enters config_sha256 and redacted_sha256: - a non-secret parameter contributes its name and value; - a secret-named parameter contributes its name only, the way an env or header key does; - env and headers stay out; - published fields are unchanged. The path stays out on purpose. A webhook path is a secret whose rotation must stay quiet, and the existing host-audit invariant requires exactly that; a digest cannot tell a capability path from a secret one. That invariant's fixture removed the token parameter while calling it a rotation, so it now rotates the value, and it also asserts that adding a query parameter fires. Servers without a query keep their earlier digest. The BigSimmo replay now records no missing change, and six mutations each fail a test. Co-authored-by: Claude Opus 5 fix(instructions): resolve documented Claude skill and command frontmatter (#722) (#725) The instruction reader refused frontmatter that Claude Code documents as valid. An unresolved instruction is a blocking coverage issue, so one such skill or command made the whole host inventory partial. - Skills now accept when_to_use, arguments, disallowed-tools, effort, background, paths and shell. - Commands, documented as taking the same frontmatter, accept user-invocable, disallowed-tools, effort, arguments, name and paths. - Values are type-checked. String-or-list fields include argument-hint, whose documented [issue-number] form YAML reads as a list. effort and shell must be one of their documented values. Claude booleans accept their documented string spellings; Cursor's alwaysApply stays exact. - A skill without name takes its directory name, and the default enters the digest. - description stays required, and undocumented keys stay refused. Seven mutations, one per rule, each fail a test. The #660 results README now names winze's real cause: its bracket-form argument-hint. Co-authored-by: Claude Opus 5 Measure correct covered comparisons on 30 frozen public repositories (#660) (#724) * bench(cold-start): measure correct covered comparisons on 30 frozen public repositories (#660) Adds the harness #660 asks for, under benchmark/cold-start/: - select.py freezes the population into selection.json. It takes 10 public repositories each for root .claude/settings.json, .mcp.json and .cursor/mcp.json, each with at least two commits touching the file in 90 days. Head is the newest non-merge commit; base is its parent. Every rejected candidate is recorded with its reason. - expected.py derives each case's expected changes and scope from the two file versions and host documentation, never from the engine. - run.py is the live driver: fresh clone, a fresh venv with pip install ./, then shipgate diff. It adds nothing that would hide a broken route. - score.py scores a case a success only when it is comparable, in supported scope, names every expected change and nothing else, and uses at most 2 commands within 5 minutes. - vendor.py and replay.py keep each case's two file versions, redacted and with local paths masked, and replay them offline. The new test pins every case to its recorded outcome. The first run against candidate 8d43106f gave 17/30, below the 24 bar. All 19 comparable cases were correct. All 11 incomparable cases stopped on a surface the commit did not change, and each stop point is mapped to an issue: #700, #720, #721, #722, #723. Co-Authored-By: Claude Opus 5 * docs(changelog): record the cold-start harness (#660) Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 fix(host): resolve local Claude Code settings layers by documented precedence (#657) (#719) audit --host --scope local-static raised a blocking unresolved_precedence issue whenever a host had two config layers. Coverage was then partial and a baseline was refused. The ordinary developer setup, user settings plus a project .claude/settings.json, could never save a local baseline. Claude Code layers now follow the documented stack. Permission rules, additional directories and hooks merge across layers. A scalar setting keeps the highest layer (managed > project local > shared project > user), and the kept grant's source names that layer. A same-named MCP server keeps the local-scope entry. The managed-only rule and hook restrictions apply only from managed settings. A project-file defaultMode of auto or bypassPermissions is still reported beside the lower mode, because older clients honored it. Disagreeing sandbox or enabledPlugins values, unranked layers, and Codex and Cursor layers still fail closed, now naming the key and each layer. The repository scope, used by the PR route, is not projected. Co-authored-by: Claude Opus 5 Publish a version declared advisory through the release pipeline (#648) (#718) * feat(release): declare each version's release channel in reviewed code (#648) Stage 1 of the advisory v1.0.0 publication path. A v* tag no longer identifies its release line, so the line comes from a committed, reviewed declaration, never from which artifacts a run happens to find: choosing by artifact presence fails open, and PyPI never gives a version back. scripts/release_channel.py resolves a v tag against .github/release-channels.json and fails closed on a missing or malformed declaration, an unknown channel, a duplicated key (json would otherwise keep the last one silently) or a tag that is not v. select() is the check release.yml's normalizing job will call: it accepts only the declared channel's verification succeeding with the other skipped. The committed declaration records the owner's advisory 1.0 decision. Mutations prove the guards fire: defaulting an undeclared version to qualified, letting both channels' verification run, and last-duplicate- wins each make the tests fail. Not wired into any workflow yet; that is the next stage, before a PR. Co-Authored-By: Claude Opus 5 * feat(release): keep advisory v* releases off the gate line's cadence (#648) An advisory v1.0.0 has the same tag shape as a qualified release, so release_cadence counted it as the gate line and would have reported that cadence as kept. That is the dishonesty the preview namespace exists to prevent (Amendment 4). Each tag's channel is now read from the declaration in that tag's own committed tree. No declaration means history, on the gate line. qualified counts on the gate line, advisory counts on the advisory line, and a declaration that does not name the version counts on neither, because the release workflow refuses such a tag. An advisory v* entry is dated by its tag, so the advisory block reports tagged_at rather than a preview's built_at, following the JSON's existing date-key convention. release_channel gains parse_declaration, so both readers share one validator, and resolve_tag now refuses a local segment explicitly: is_release_version accepts one and no public index does. A mutation that stops the gate-line filter makes the new tests fail. Co-Authored-By: Claude Opus 5 * feat(release): publish a version declared advisory through the release pipeline (#648) The owner chose advisory 1.0: the production-tier bar governs a qualified blocking claim, not a version number. Every v* tag used to need a signed qualification artifact that does not exist, so 1.0 could not publish. release.yml now resolves the channel from .github/release-channels.json at the tagged commit, before anything is built. It calls exactly one of release-verify.yml, unchanged, and the new release-advisory-verify.yml, which has no qualification step to skip. `release_channel.py candidate` forwards only the declared channel's values, and only when that job succeeded and the other stayed skipped and exported nothing. It then names the channel's rehearsal file, asset set and signed assets. Every later job names each dependency's success, because a skipped verification would otherwise skip everything downstream. One publisher still serves both lines. The advisory candidate is the wheel a Release Engine Smoke run exercised for the exact commit. The smoke evidence must name those bytes, and the wheel must be byte-identical to the tagged tree's build (--exercised). The release ships a signed advisory-statement.json with fixed claims and no tier. release_publication holds one closed asset set per channel; verify-manifest requires --expected-channel and derives the required bundles with --require-signatures. A tagged version's channel cannot change. Amendment 5 records the policy. Pinned tests keep their meaning under the new spelling. Twelve mutations, one per new control, each turn a test red. One of them first passed on a comment naming the removed flag, so those checks now read commands with comment lines stripped. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 Build(deps-dev): Update mcp requirement from <3,>=2.1.1 to >=2.2.0,<3 (#705) Updates the requirements on [mcp](https://github.com/modelcontextprotocol/python-sdk) to permit the latest version. - [Release notes](https://github.com/modelcontextprotocol/python-sdk/releases) - [Changelog](https://github.com/modelcontextprotocol/python-sdk/blob/main/RELEASE.md) - [Commits](https://github.com/modelcontextprotocol/python-sdk/compare/v2.1.1...v2.2.0) --- updated-dependencies: - dependency-name: mcp dependency-version: 2.2.0 dependency-type: direct:development ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Pengfei Hu test(mcp): run the SDK injection cross-checks on the mcp the extra allows (#716) (#717) Both SDK cross-checks in test_fastmcp_injection_contract.py imported only mcp.server.fastmcp, the 1.x module. FastMCP became MCPServer in mcp 2.x and the [mcp] extra requires mcp>=2.1.1,<3, so the eight cases comparing the static reader with the real SDK's find_context_parameter skipped on every install satisfying the extra, CI included. They ran only against an out-of-range mcp 1.x. A helper now tries mcp.server.mcpserver first and the 1.x path second. It skips only when mcp is absent, naming the extra, and fails when an installed SDK exposes neither location, so the next rename cannot hide the same way. Two tests pin both branches without needing mcp. Measured: mcp 2.2.0, 8 skipped before and 8 passed after. mcp 1.27.2, 8 passed after on the 1.x fallback. With no mcp (a base install does not pull it in), 45 passed and 8 skipped naming agents-shipgate[mcp]. CI still installs no mcp; recorded on #716. Co-authored-by: Claude Opus 5 Freeze the report contract at 1.0 and publish what it promises (#569) (#642) * Freeze the report contract at 1.0 and publish what it promises (#569) `STABILITY.md` promised that a 1.0 product line would not begin until the report schema reached `1.0` and held without a breaking change. Main emitted `0.43`, and no delivery issue owned satisfying that published condition. Freezing the number is the easy half. `report-schema.v1.0.json` and `report-schema.v0.43.json` are byte-identical apart from `$id`, `title` and the version constant, so a consumer written against `0.43` needs no edit -- and `STABILITY.md`'s own rule reads "major on breaking", which makes "this major broke nothing" a claim somebody has to be able to check rather than a sentence in a migration note. A test holds it. What actually moved is the promise. `1.x` is additive-only; a change that cannot be expressed additively needs `2.0`; a deprecation cycle counts shipped releases rather than time on unreleased `main`; every published schema URL keeps its bytes. `docs/report-1-0-contract.md` carries the stable/provisional inventory of report fields, CLI, exit codes, Action and control surfaces, the migration from the actually shipped `v0.15.0` contract (`report 0.28`, `contract 10`, no control floor at all), and the recorded RC exercise. `tests/test_report_1_0_contract.py` checks that document against the runtime: every field classified exactly once, the presence column equal to the published schema's `required` list, and nothing `STABILITY.md` already promised stable quietly demoted. Pre-freeze reports stop being engine input. `ReadinessReport` is `extra="allow"` with defaults nearly everywhere, so `model_validate` accepts a `0.9` payload and returns an object whose newer blocks hold this build's defaults, with nothing distinguishing a value that was recorded from one that was invented. `scan --diff-from`, `explain-finding`, `findings`, `scenario suggest` and `evidence-packet` now refuse one before validation, by name, with a stable `reason_code` and a regeneration route. Nothing is converted, so no artifact gains current authority through conversion or a restamped digest. Every superseded schema stays published, so archived reports remain validatable, and bundle readers (`org`, `attest`) stay ungated because refusing there would break the diagnostics this freeze must preserve. The two parser directions are not symmetric and are no longer treated as if they were. A reader that only projects its input accepts a later `1.x` minor -- that is what additive means. An evidence comparison refuses it, because the additive rule is a promise to consumers, not a licence to diff against blocks this build cannot interpret. Strict is the default a new caller gets by forgetting. One fail-open defect found on the way, and it is the reason the coupling is now structural: the release decision routed an incomparable `--diff-from` base to `insufficient_evidence` only when its `source_warning` gap matched three substrings of one refusal's prose. Rewording the refusal -- which this change does -- silently downgraded the verdict to `review_required`: the run stopped withholding a verdict it had no evidence for, and nothing that named the decision failed. The refusal now carries its own reason code and the classifier looks it up. Qualification moves with the engine: the production `beta` policy pins `1.0`, and issuance of the `pre_1_0` tier is retired -- the runner produces no artifact carrying it and `--policy-tier pre-1.0` is refused by name rather than dropped from the choice list. Retirement is issuance only. The policy, its thresholds and every reader remain, and it deliberately keeps its historical `0.43` pin: `tier_for_requirements` names a policy by byte-equality, so re-pinning it would demote every existing `pre_1_0` artifact to the unnamed `test` tier and replace an accurate diagnosis with a misleading one. No scoring floor moved. Runtime contract `33 -> 34`; `minimum_control_contract_version` stays `21` because every operational control shape is byte-identical. Packet, verifier, receipt, capability-lock, attestation, preflight and host-grants schemas keep their versions -- renumbering every schema would be work without compatibility. The RC exercise ran against an untagged, unpublished candidate wheel built with the locked `hatchling==1.32.0` backend and installed into an isolated virtualenv: emitted schema equals advertised, a relabelled `0.43` report is refused with its route, a `1.99` report is projected but not compared, and the candidate's own report is accepted. It is automated in `scripts/release_engine_smoke.py` so the final candidate re-runs it rather than anyone repeating it by hand. The wheel is unsigned, unqualified and built from a dirty tree; it establishes the report contract and nothing about qualification, adoption or the final candidate. The design-partner ledger's source-tree column was re-measured rather than carried forward, because a contract bump is exactly what its guard exists to catch. Every cell reproduced unchanged except the contract number. Co-Authored-By: Claude Opus 5 * Address review: honest test scope, structural guards, registry claim Ten findings from an extra-high-effort review of the freeze. The two that mattered were both cases of a claim being stronger than the code. **The compatibility fixtures did not test what they said.** `_isolated_cli` documented itself as running an installed package with "no PYTHONPATH inheritance" and `-I`, while actually setting `PYTHONPATH` to this worktree's `src/` and passing no `-I`. The pin is *correct* -- the `.venv` editable install points at the main checkout, so without it a worktree run tests somebody else's source -- but the docstring, the module header and the PR all read as though a built distribution were covered by CI, when the only run that exercised one was the manual RC build. Renamed `_worktree_cli`, and both docstrings now say what it does and point at `release_engine_smoke.py` for the installed-distribution claim. **`explain-finding` asserted an invariant in a comment.** Removing the `>= 0.12` floor came with "1.x carries `agent_action` by construction". The published schema requires the field, but the Pydantic model has `agent_action: AgentAction | None = None`, so a payload that merely *declares* `1.0` still produces the `"agent_action": null` explanation #58 review P2.2 exists to prevent. It now refuses structurally after validation, the way `findings.py` already refuses a missing `provenance_kind`. The rest: - `classify_report_schema_version` guarded a runtime invariant with a bare `assert`, which `python -O` strips -- leaving `None[1]` to raise `TypeError` out of the one function whose job is to fail cleanly, past callers that wrap only `ValueError`. Now a typed refusal naming `doctor --json`. - `_report_schema_precedes_semantic_diff`, `_SEMANTIC_DIFF_REPORT_SCHEMA_VERSION` and `_schema_version_at_least` were dead after the boundary replaced them. Deleted rather than left to invite a second definition of "comparable base". - The `>= 0.31` check gating `binding_surface_facts` became vacuous once only `1.x` bases are accepted. Presence in the payload is the half that still discriminates. - `test_the_1_0_schema_is_a_promotion_of_the_last_pre_freeze_schema` skipped as soon as a later minor became current -- retiring the only byte comparison guarding `report-schema.v1.0.json` exactly when that file became a frozen published artifact nothing else pinned. It now compares the two frozen documents, which is a permanent fact. - `docs/report-reading-for-agents.md` still taught `report.get("report_schema_version", "0.6")`; `0.6` is a version the same release made unusable as input. - `select_release_requirements` discarded its `wheel_version` parameter with a `del`; it names the wheel in the retirement message instead. - `report_schema_exercise` now states which side each fact comes from: observed values from the installed candidate, `REPORT_CONTRACT_MAJOR` and the refusal recognizer from the source tree as the reviewed claim under test. **`docs/distribution-surfaces.md` gains a `report_schema_pin` claim.** Five registered surfaces restate which `report-schema.v.json` a reader should validate against, and "which schema does this build emit" is an engine answer -- `contract --json` publishes it. Per that document's own rule it belonged in the registry rather than in hand-maintained per-file lists. The proving test checks both directions, and both were verified by perturbation: a pin left behind fails, a pin ahead of the build fails, and a frozen-reference list of older versions correctly does not. Co-Authored-By: Claude Opus 5 * Address review round 2: a refusal that could escape its own classifier Fixing the bare `assert` in round 1 introduced the defect that fix was about. `classify_report_schema_version` now raises for a build whose own schema version cannot be parsed, carrying `report_schema_engine_version_unreadable`. But `report_schema_refusal_code` built the codes it would recognise by walking `ReportSchemaStatus` — a second copy of the producer's vocabulary — and that code is not an *input* status. So the new refusal came out unrecognised, and an incomparable `--diff-from` base carrying it would have routed to `review_required` instead of withholding the verdict: the exact fail-open class round 1 fixed, reintroduced in the one branch that fires when the install itself is broken. The recognizer now matches the bracket marker rather than an enumerated list, so producer and consumer are in sync by construction, and the direct raise carries the marker like every other refusal. `test_no_refusal_this_module_raises_can_escape_the_classifier` states the property over every refusal the module can raise instead of over the statuses somebody remembered to enumerate. Also: `classify_report_schema_version`'s docstring said it decides and says why, with no mention that it can now raise. It says which case raises and why that case is about the build rather than the payload. Co-Authored-By: Claude Opus 5 * Address review round 3: execute the broken-install branch Round 1 replaced a bare `assert` so a build that cannot parse its own schema version fails as a typed, readable refusal rather than a `TypeError` past callers that wrap only `ValueError`. Round 2 made that refusal recognizable to the classifier. Neither round ran the branch: it is unreachable from any input, which is exactly why it needed forcing. `test_a_build_that_cannot_read_its_own_version_refuses_readably` patches the engine's declared version to an unparsable one and asserts the three things the two earlier rounds claimed — a typed `ReportSchemaCompatibilityError`, a message naming `doctor --json`, and a code the classifier reads back. Verified by perturbation: removing the marker from that one refusal fails the test. Co-Authored-By: Claude Opus 5 * Address review round 4: the contract doc named a command that does not exist The migration section listed `agents-shipgate packet` as one of the boundaries that refuses a pre-freeze report. There is no such command — it is `evidence-packet` — so the one document whose job is telling a reader what to do with a refused artifact told them to run something that errors. STABILITY.md's migration note carried the same name. Nothing caught it: the docs tests check links, and the distribution-surface parity harness checks pins. `distribution-surfaces.md` already states the invariant this restores — "a surface that tells a reader to execute something must name something that resolves in the build it names" — so `test_every_command_the_contract_names_exists` checks every `agents-shipgate ` in the contract doc against the CLI's own registered commands and groups. Verified by perturbation. Co-Authored-By: Claude Opus 5 * Keep the report-freeze changelog independent of pending feature PRs * docs(pilot): re-run the route readiness dry run on the reconciled tree The ledger's source-tree column is a claim about this tree, and test_route_readiness_source_tree_row_matches_this_tree refused to let the contract bump be carried forward. So the dry run was re-run rather than the number retyped. Rebuilt the synthetic Route H fixture exactly as documented and observed every source-tree cell on 871786ba: runtime contract 36, host inventory schema 0.4, check block/critical with 4 violations and visible coverage, init --write --ci not_applicable_host_review with no manifest or workflow, init then manifest-free verify both exit 0 with six rows (three added permissions, two removed narrower rules, the added MCP server), and baseline/drift with all four expansion signals. Every cell reproduced unchanged except the contract number, so only that cell and the dated recheck line change. Co-Authored-By: Claude Opus 5 * docs(report-1.0): re-establish the installed-candidate replay on contract 36 The record said 'Engine identity observed: contract_version 34' about a wheel built from 8e81672b. After this branch became contract v36, that line described a build that no longer exists, and retyping the number would have claimed an observation nobody made. So the candidate was rebuilt and the replay re-run. Release Engine Smoke run 34726385422 built the wheel from 8a14470d (sha256 02a7a3fe...) and its report_schema_exercise passed. The downloaded wheel was then installed in an isolated venv and driven with python -I and PYTHONPATH cleared, and all six rows were replayed by hand: emit 1.0 equal to contract --json, findings on a 0.43 relabel exits 3 with report_schema_pre_freeze naming the scan regeneration route, 1.99 and the current report accepted, scan --diff-from a 0.43 base completes with the refusal as a source_warning, every diff enabled=false and the decision not passed, and a 1.99 base refused as report_schema_newer_than_engine with diffs disabled. Every row reproduced, so the table stands. 'Built from an uncommitted working tree' was no longer true; the paragraph now quotes what the run records about itself: qualified false, qualification_claim none, synthetic distribution smoke only. Co-Authored-By: Claude Opus 5 * merge: integrate #704, #710 and #678 from main CHANGELOG is main's Unreleased entries verbatim plus this branch's own entry; llms-full.txt regenerated from its script rather than merged. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 fix(host): tell a widened permission rule from a narrowed one (#657) (#678) Host grants are keyed by rule text, so replacing `Bash(npm *)` with `Bash(npm test:*)` arrived as one removal plus one addition — the same shape as replacing it with `Bash(*)`. Set arithmetic cannot separate those, so drift reported a tightening as an expansion. `core/permission_lattice.py` decides which of two rules is wider for the patterns hosts actually use, and answers `None` for anything else — character classes, interior stars — where a guess would put the same wrong direction in a new place. A narrowing no longer contributes to `expansion_signals`, which `preflight` prints as "Expansion signals" and the drift markdown flags with a ⚠; it stays visible in `changes`. A widening is now named rather than only counted. The same lattice sets severity. Every wildcard allow was `admin`/`critical`, so a carefully written host config rated 7 of 10 grants `high` or above — three of them `critical` for reading files. The same config now rates 1 of 10: the MCP server. This narrows a blocking check. Adding `Read(**)` was a `critical` release blocker; read-only whole-tool grants now raise SHIP-HOST-BOUNDARY-PERMISSION-ALLOW-EXPANDED instead. Execution, network, write, `*` and unknown tools still block. The gate had classified wildcards a second time; both readers now share one lattice. 51 of the new file's 78 cases fail against pre-fix `main`. Co-authored-by: Claude Opus 5 docs(cursor): lead the Cursor instructions with the named rows (#662) (#710) The Cursor rule opened with `check --format agent-boundary-json` and the control envelope, so an agent following it told the user "a human must review" without naming what changed. `AGENTS.md` has led with `shipgate diff` since #651; this brings the generated surface level. Names the row fields the agent should quote, says a covered comparison with no rows is a real answer rather than a missing one, and repeats the rule the control envelope already carries: a row is a description, never a permission to edit, commit, push, merge or report complete. The committed `.cursor/rules/agents-shipgate.mdc` and the copyable snippet in `docs/target-repo-agent-snippets.md` are regenerated from the one renderer. Both are pinned to it by existing tests, which is how this change was caught rather than shipped half-applied. Co-authored-by: Claude Opus 5 Read Claude Code's hook file, and count what is still unread (#689) (#704) Twelve public repositories, 456 first-parent steps, classified twice: once with the engine's own `BOUNDARY_ADAPTERS` and once with a census of host paths written from the hosts' documentation. Compared per path, the two agree on 441 steps. The disagreements were not one thing: - `.claude/hooks/hooks.json` (2 steps) — a real coverage gap, fixed here. - `.github/actions/*/action.yml` (20 steps) — a real gap needing a grant kind that does not exist; #701 with the measurement. - `.claude/hooks/pre-compact` (1 step) — the script a hook command runs, which is read by nothing; #702. - `.vscode/settings.json` — a census bug. The file held `editor.formatOnSave` and `files.eol`. It is host surface only when it carries an `mcp` or agent key, which a census of paths cannot know, so the entry is removed with the reason recorded rather than the reader extended. `.claude/hooks/hooks.json` is byte-for-byte the shape of `.codex/hooks.json`: `_source_kind` already routes any `hooks.json` to the hook reader and `_hooks_grants` already produces the grant, so the registry entry was the whole fix. A `SessionStart` command running `.claude/hooks/session-start` now produces a row. The row said "hook". `_grant_value` recognises a grant by `rule`, `server`, `name`, `value` or `permission`, and a hook grant carries `event`, so every hook row rendered as its own kind — telling a reviewer that a hook changed and not which one. The census is committed at `benchmark/cold-start/census.py` and its independence from the registry is the point: classifying history with the registry makes a path the engine does not read look like a benign change with nothing to report. It compares per path rather than per step — a step where each side names a different file is two findings, not agreement — and separates three counts: unexplained gaps, gaps an issue already owns, and census bugs. It exits non-zero only on the first and third. A test asserts every registry path is named by the census, which immediately found ten it was not naming, including `AGENTS.override.md`, `.cursor/cli.json` and `policies/*.shipgate.yaml`. One limit is now recorded rather than unstated: the script a hook command names is not read, so editing `.claude/hooks/session-start` changes what runs with no row. That is #702, it is owned in the census, and the test that pins today's quiet behaviour says so. 7 hook cases plus 2 census cases; 4 fail against the parent commit. Closes #689. Co-authored-by: Claude Opus 5 fix: report named host configuration directories as incomplete (#715) * Read the base tree the host reader will read, and no more (#686, #688) `shipgate diff` and verify's host comparison materialized the whole base tree to read a handful of files. Two defects followed. **Cost.** Materialization runs one `git cat-file` per blob. On a 26 MB repository with 1,168 files whose host surface is four files, that is 1,168 subprocesses: 25 seconds, against 1 second on the toy scenario. The pack was not the cost — 200 ms — the subprocesses were. **Reach.** Every entry was checked for portability, name collision and external binding, so anything anywhere refused the whole repository. Four of twelve public repositories with committed host configuration were uncomparable on every step of their history: `getsentry/sentry-mcp`, `miantiao-me/bm.md`, `gtkx-org/gtkx` and `awslabs/nx-plugin-for-aws`. One of them for three symlinked PNGs under `website/public/`. `archive_tree` now takes a `scope` predicate. Scoped, it packs the tree rather than the commit — no ancestry, which is also why a shallow clone can be compared against a commit it already holds — and materializes only what the predicate accepts. The predicate is `is_boundary_surface_path`, placed in the registry that already owns path classification, so the archive and the live reader cannot disagree about what the surface is. Scoping narrows what is *materialized*, never what is *verified*: every blob written is still checked against its object ID and the isolated store is still fsck'd. An unscoped archive is unchanged in every respect, including refusing symlinks — verify materializes trees for the release decision, and narrowing the advisory host read must not loosen that. Two further fixes were needed before the four repositories came back: - A scoped archive recreates a symlink instead of refusing the tree. The live reader treats a link as a candidate and opens it with `O_NOFOLLOW`, so the base tree must present the same object or the two sides disagree about what the file is. - `_symlink_may_hide_boundary_glob` returns True for any non-empty path once a pattern starts with `**/`, so every symlink in a repository became a boundary candidate, failed its read and left the inventory incomplete. Only a directory has descendants; a link to a regular file conceals nothing. Unresolvable links still fail closed — `Path.is_dir()` answers False for a dangling link rather than raising, so this uses `os.stat`. No ancestor-escape guard was added. A valid tree cannot put a blob under a symlinked name — the name would be both a blob and a tree, a duplicate entry `fsck --strict` rejects before anything is written — so such a guard could never fire, and a guard that cannot fire is not a guard. The invalid tree is built by hand in the tests and checked to be refused. Measured after: b4 25.3s -> 2.7s; `gtkx-org/gtkx` comparable on every step. Three repositories still refuse, all for a symlink *at* a boundary path (`CLAUDE.md -> AGENTS.md`, `.claude/skills -> ../.agents/skills`). That is own threat review. Two existing tests were rewritten rather than deleted. The shallow-clone case now asserts the comparison it can actually answer, and a second case covers a base the clone genuinely does not have, which still routes to `fetch_base` with `refs_missing`. Static-only line pins moved 2001 -> 2092 and 2245 -> 2336. 13 new cases; 8 fail against pre-fix `main` on behaviour, and the 5 that pass on both are the contract-preservation half. Closes #686. Closes #688. Co-Authored-By: Claude Opus 5 * fix(diff): refuse a shallow checkout only when its base is unusable (#686) `diff` refused every shallow checkout. The scoped archive this branch adds packs the tree and walks no ancestry, so a base the clone already holds is readable at `--depth 1` — the refusal was a repair for a problem the caller did not have, which is the principle this branch already applied to verify's test but not to the command itself. Measured on five public repositories cloned at `--depth 3`: `diff` refused all five, while `archive_tree(scope=...)` read all five in 0.5-2.1 s. `actions/checkout` defaults to `fetch-depth: 1`, and #660's harness clones shallowly, so this was reach lost for no gain. The recovery now fires when the base ref or the merge base is genuinely unavailable. No graft check: `git merge-base` does not guess past a shallow boundary, it reports nothing — verified on two diverged shallow clones whose true base lay past the graft, both returned empty. A check for a returned-but-wrong base could not be made to fire, and a guard that cannot fire is not a guard. `test_shallow_checkout_names_a_recovery_that_restores_the_diff` passed `--base HEAD`, a comparison needing only HEAD's tree; it now names a base the clone genuinely lacks and keeps the full recovery walk. Both new cases fail against this branch's previous head. Co-Authored-By: Claude Opus 5 * fix: prove shallow comparison ancestry and retain bound coverage * fix: retain explicitly named host configuration directories as incomplete (#613) --------- Co-authored-by: Claude Opus 5 Read the base tree the host reader will read, and no more (#686, #688) (#703) * Read the base tree the host reader will read, and no more (#686, #688) `shipgate diff` and verify's host comparison materialized the whole base tree to read a handful of files. Two defects followed. **Cost.** Materialization runs one `git cat-file` per blob. On a 26 MB repository with 1,168 files whose host surface is four files, that is 1,168 subprocesses: 25 seconds, against 1 second on the toy scenario. The pack was not the cost — 200 ms — the subprocesses were. **Reach.** Every entry was checked for portability, name collision and external binding, so anything anywhere refused the whole repository. Four of twelve public repositories with committed host configuration were uncomparable on every step of their history: `getsentry/sentry-mcp`, `miantiao-me/bm.md`, `gtkx-org/gtkx` and `awslabs/nx-plugin-for-aws`. One of them for three symlinked PNGs under `website/public/`. `archive_tree` now takes a `scope` predicate. Scoped, it packs the tree rather than the commit — no ancestry, which is also why a shallow clone can be compared against a commit it already holds — and materializes only what the predicate accepts. The predicate is `is_boundary_surface_path`, placed in the registry that already owns path classification, so the archive and the live reader cannot disagree about what the surface is. Scoping narrows what is *materialized*, never what is *verified*: every blob written is still checked against its object ID and the isolated store is still fsck'd. An unscoped archive is unchanged in every respect, including refusing symlinks — verify materializes trees for the release decision, and narrowing the advisory host read must not loosen that. Two further fixes were needed before the four repositories came back: - A scoped archive recreates a symlink instead of refusing the tree. The live reader treats a link as a candidate and opens it with `O_NOFOLLOW`, so the base tree must present the same object or the two sides disagree about what the file is. - `_symlink_may_hide_boundary_glob` returns True for any non-empty path once a pattern starts with `**/`, so every symlink in a repository became a boundary candidate, failed its read and left the inventory incomplete. Only a directory has descendants; a link to a regular file conceals nothing. Unresolvable links still fail closed — `Path.is_dir()` answers False for a dangling link rather than raising, so this uses `os.stat`. No ancestor-escape guard was added. A valid tree cannot put a blob under a symlinked name — the name would be both a blob and a tree, a duplicate entry `fsck --strict` rejects before anything is written — so such a guard could never fire, and a guard that cannot fire is not a guard. The invalid tree is built by hand in the tests and checked to be refused. Measured after: b4 25.3s -> 2.7s; `gtkx-org/gtkx` comparable on every step. Three repositories still refuse, all for a symlink *at* a boundary path (`CLAUDE.md -> AGENTS.md`, `.claude/skills -> ../.agents/skills`). That is own threat review. Two existing tests were rewritten rather than deleted. The shallow-clone case now asserts the comparison it can actually answer, and a second case covers a base the clone genuinely does not have, which still routes to `fetch_base` with `refs_missing`. Static-only line pins moved 2001 -> 2092 and 2245 -> 2336. 13 new cases; 8 fail against pre-fix `main` on behaviour, and the 5 that pass on both are the contract-preservation half. Closes #686. Closes #688. Co-Authored-By: Claude Opus 5 * fix(diff): refuse a shallow checkout only when its base is unusable (#686) `diff` refused every shallow checkout. The scoped archive this branch adds packs the tree and walks no ancestry, so a base the clone already holds is readable at `--depth 1` — the refusal was a repair for a problem the caller did not have, which is the principle this branch already applied to verify's test but not to the command itself. Measured on five public repositories cloned at `--depth 3`: `diff` refused all five, while `archive_tree(scope=...)` read all five in 0.5-2.1 s. `actions/checkout` defaults to `fetch-depth: 1`, and #660's harness clones shallowly, so this was reach lost for no gain. The recovery now fires when the base ref or the merge base is genuinely unavailable. No graft check: `git merge-base` does not guess past a shallow boundary, it reports nothing — verified on two diverged shallow clones whose true base lay past the graft, both returned empty. A check for a returned-but-wrong base could not be made to fire, and a guard that cannot fire is not a guard. `test_shallow_checkout_names_a_recovery_that_restores_the_diff` passed `--base HEAD`, a comparison needing only HEAD's tree; it now names a base the clone genuinely lacks and keeps the full recovery walk. Both new cases fail against this branch's previous head. Co-Authored-By: Claude Opus 5 * fix: prove shallow comparison ancestry and retain bound coverage --------- Co-authored-by: Claude Opus 5 An empty frontmatter value means absent, not invalid (#712) * fix(instructions): an empty frontmatter value means absent, not invalid `globs:` with nothing after it is how Cursor writes a rule that is not glob-scoped. YAML reads that as None, and `_valid_metadata` rejected None for every field, so the canonical Cursor rule file classified as `frontmatter_invalid_structure`. That reason is a *blocking* inventory issue, and one blocking issue makes the base inventory incomplete, and an incomplete base inventory makes the entire comparison `incomparable`. Measured on `Doist/todoist-mcp`: eleven readable host files, two Cursor rules in the format Cursor itself generates, and `shipgate diff` answered `incomparable` with zero rows on all forty steps of its history. It now answers `comparable`. An explicit null is the same statement as an absent key. The checks that should still fire do: a wrong type is still a wrong type, an unknown key is still unknown, and a skill with an empty `name` or `description` still fails as `skill_identity_missing` rather than sliding through — pinned, because forgiving null in a shared type check would be a hole if identity were only enforced there. Eight cases; five fail against pre-fix `main`, and the three that pass on both are the contract-preservation half. Co-Authored-By: Claude Opus 5 * fix: treat optional null instruction fields as absent --------- Co-authored-by: Claude Opus 5 Build(deps-dev): Bump securesystemslib from 1.4.0 to 1.5.1 (#709) Bumps [securesystemslib](https://github.com/secure-systems-lab/securesystemslib) from 1.4.0 to 1.5.1. - [Release notes](https://github.com/secure-systems-lab/securesystemslib/releases) - [Changelog](https://github.com/secure-systems-lab/securesystemslib/blob/v1.5.1/CHANGELOG.md) - [Commits](https://github.com/secure-systems-lab/securesystemslib/compare/v1.4.0...v1.5.1) --- updated-dependencies: - dependency-name: securesystemslib dependency-version: 1.5.1 dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Build(deps-dev): Bump pip-api from 0.0.34 to 0.0.35 (#708) Bumps [pip-api](https://github.com/di/pip-api) from 0.0.34 to 0.0.35. - [Release notes](https://github.com/di/pip-api/releases) - [Changelog](https://github.com/di/pip-api/blob/master/CHANGELOG) - [Commits](https://github.com/di/pip-api/compare/0.0.34...0.0.35) --- updated-dependencies: - dependency-name: pip-api dependency-version: 0.0.35 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> fix: show host capability changes on the first PR review (#697) fix: compare effective workflow grants and inherited secret recipients (#685) (#695) Read a TypeScript tool's description in both SDK shapes (#680) (#682) * fix(mcp): read a TypeScript tool's description in both SDK shapes (#680) `ts_sdk_register_tool` resolved a registration's name and nothing else, so every tool declared on the reference MCP SDK was reported undocumented and earned a SHIP-DOC-MISSING-DESCRIPTION finding — one per tool, on the most common server shape there is. `_TS_DESCRIPTION_RE` looks like it covered this and does not: it serves the class-property idiom, not a call. Both documented shapes are now read: the options object of `registerTool("n", { description: "…" }, fn)`, quoted key included, and the positional second argument of `tool("n", "…", fn)`. The description is taken from the options object's own `description` and never a nested one. That restriction is the substance of the change: the options object carries `inputSchema`, and a JSON Schema describes every parameter, so taking the first `description:` found would publish an argument's documentation as the tool's — a confident wrong answer in place of an absent one. A tool that describes only its parameters now reads nothing. The zero-install detector carries the same extractor; the shared conformance corpus holds the two copies together. 6 of the new file's 13 cases fail before the fix; the 7 that pass are refusals the old reader satisfied by reading nothing at all. Co-Authored-By: Claude Opus 5 * fix(mcp): preserve TypeScript description member boundaries and overrides * fix(mcp): invalidate descriptions after unknown member keys --------- Co-authored-by: Claude Opus 5 Give shallow diff checkouts a runnable recovery instead of a traceback (#683) (#692) * fix(diff): route shallow checkouts to runnable recovery (#683) * fix(diff): keep shallow recovery copyable in colored terminals docs: scope v1 readiness to the advisory host workflow (#687) Read Go MCP descriptions from options and translated struct fields (#658) (#679) Read Go MCP descriptions from direct options and struct fields, including complete calls to declared translation-helper parameters. Respect option order, preserve expression and scope boundaries, and keep the package and zero-install readers aligned. Regression tests, a current upstream source sample, and final-head CI passed. Related to #658. Source annotation projection and the broader server qualification table remain open. build(release): two lines, two promises, two cadences (#648) (#676) Measure advisory preview and qualified-release cadence independently. Count only publisher-shaped preview tags, preserve existing top-level release JSON fields, and report the advisory measurement alongside them. Cadence boundary, operator exit, release-pipeline, and final-head CI checks passed. Related to #648. The operational publication and scheduled enforcement acceptance criteria remain open; this change establishes measurement and does not prove that two releases have shipped. Build(deps): Bump build from 1.5.0 to 1.6.0 (#529) Update build to 1.6.0 and synchronize the development declaration and release sealing lock. Use the compatible securesystemslib 1.5.1 closure. Dependency-lock validation, wheel and source-distribution build smoke, release-pipeline tests, and final-head CI passed. Build(deps): Bump uv from 0.12.5 to 0.12.9 (#527) Update the publication toolchain to uv 0.12.9 with a consistent hash-locked closure. Replace the yanked, incompatible securesystemslib 1.5.0 pin with 1.5.1 while retaining unrelated pins. Dependency-lock validation, isolated hash-locked publication installation and Sigstore startup, release-pipeline tests, and final-head CI passed. feat(cli): six commands in --help, and reader strings in the reader's language (#652) (#674) Expose diff, check, verify, audit, init, and doctor in root help, and provide complete command discovery through --help-all using Typer's supported context API. Replace internal binding-graph prose with reader-facing explanations and preserve machine enum values. CLI, fixture, workspace-guard and performance checks, Ruff, and final-head CI passed. Closes #652. feat(diff): promote reviewed authority comparison to main (#651) (#677) * feat(check): compare against the detected base by default (#649) `shipgate check` compared the working tree with `HEAD`. On a branch whose changes were already committed it therefore answered `allow` with an empty change set — truthfully reporting "nothing is uncommitted" to someone asking "what does this branch change". Reaching the real answer took `--base --head HEAD`, a pair nothing in `--help` suggested and which the CLI rejected unless both were supplied. A flagless run now compares the detected default branch's merge base against the working tree: one comparison spanning committed branch work and uncommitted edits, which is the shape `working_tree_context` already provides and the verifier already uses for exactly this reason. `--base` and `--head` became independent, and `subject.base` names the ref actually compared — a verdict whose base is unnamed cannot be reviewed. Two things the issue's sketch got wrong, handled deliberately: `detect_default_base` refuses a local `main` on purpose, because a remote is the authority it might be stale against. Rather than overriding that, the fallback is narrowed to the case the rationale does not cover — a repository with no remote at all, where the local branch is the only base in existence. Opt-in, so `verify` keeps its behaviour exactly. That detector also returns nothing when HEAD *is* the default branch, which would have made an honest stop out of the commonest situation there is. Being on the default branch and having no base at all are now separated: the first keeps the complete working-tree answer, only the second stops and names `--base`. `verify_command_for` required both refs or neither — a mirror of the CLI rule this change removes. It now emits whichever ref was resolved, because `verify` accepts either alone, and a command that disagrees with the check that emitted it is the recurring second-implementation defect. 12 cases in `tests/test_check_default_comparison.py` covering every rung. Replayed against pre-fix `main`: 8 fail, and the 4 that pass are the default-branch, empty-ref and option-smuggling cases that must not change. `tests/test_codex_boundary_check.py::test_codex_check_rejects_one_sided_git_refs` pinned the removed rule; it now pins the new one, plus the empty-ref refusal it was also covering. Two line-pinned entries in the static-only allowlist moved with the code they point at. Closes #649. Co-Authored-By: Claude Opus 5 * feat(diff): show what a change does to the agent's authority (#651) The first-run measurement that opened this milestone reached its first actionable output in sixteen commands. The one moment that worked was `audit --host --drift`: four widenings named in twelve lines, no noise. It was also the least reachable thing in the product, because it needed a baseline recorded in advance, on the base ref, and reachable from the branch — two checkouts and a committed file before the first answer. `shipgate diff` supplies the other side from Git instead. Materialise the base tree, read it with the same host readers, hand both inventories to the same comparator. No manifest, no committed baseline, no second checkout: Agent capability diff main (17a958ba) -> working tree ⚠ critical added claude-code .claude/settings.json Bash(*) matches any command of this kind, without a prompt ⚠ critical widened github .github/workflows/ci.yml read → write, contents: write, pull_request_target runs with repository secrets on fork pull requests; … Surface discipline, answered before writing it: 1. **Metric:** time-to-first-value. Sixteen commands to two is this milestone's exit criterion, and this is the command it turns on. 2. **Could an existing surface carry it?** The decision logic is *entirely* reused — this adds no comparator, no severity model, no second verdict. `audit --host --base ` was the smaller alternative and was rejected on one ground: `--help` shows the product's prominent commands, and a flag on a subcommand cannot be one. Discoverability is the metric here. 3. **Non-goals:** no second verdict, no adapter, no network, no LLM. This command publishes no verdict at all. It is a projection, and the tests hold it to that. `risk` is the engine's severity, `expansion_signals` is the engine's word on widening, and neither is recomputed: a wildcard add the engine did not signal is reported as an addition, not promoted to an expansion, and a change on both sides stays `changed` until the engine calls it widening — guessing narrowing needs the pattern lattice in #657. `expands` is a separate field from `direction` because presence is a fact both sides show and widening is a judgement only the engine may make. Registered in `docs/distribution-surfaces.md` under rule 2 — a row with no claims and a note saying why: it restates none of the engine's answers, so there is nothing to drift. Contract `primary_commands` registration and `--help` prominence are deliberately left to #652, which owns that question and can bump the contract once instead of twice. Twelve cases in `tests/test_capability_diff.py`, including the ones that would fail if a field were ever invented rather than read. Closes #651. Co-Authored-By: Claude Opus 5 * Keep the diff changelog entry independent of other pending PRs --------- Co-authored-by: Claude Opus 5 fix(control): make every emitted next action lead somewhere (#650) (#672) On a repository with recognized host configuration and no manifest, following the tool's own advice went round three commands forever: verify --preview -> "Next: init --write" init --write -> refuses; host-only needs no manifest, "Next: audit --host" audit --host -> "Next: verify --preview" <- back to the start The host-audit footer was unconditional. It is now the baseline comparison where there is no manifest, `verify` where there is, and `--drift` once a baseline exists — and it names the base ref deliberately, because recording a baseline from the changed checkout acknowledges the very change under review instead of comparing it. Writing the property test found a second defect the cycle had hidden. That footer emitted its command *without* `--workspace`, so following it ran the next step against the caller's current directory rather than the repository just audited. Walking the chain did not loop; it wandered into a different repository and terminated there. Every emitted command now carries the workspace it was produced for, which is what the invocation policy already required of every other control surface. A preview that evaluated no gate no longer calls itself failed. "failed" was the catch-all whenever no release decision existed, so a preview with nothing to gate printed "Agents Shipgate verify: failed" directly above "Exit code: 0" while its own machine fields said `execution: "not_run"` and `applicability: "not_evaluated"`. Only the human line disagreed. `tests/test_next_action_chains_terminate.py` walks every entry command on a host-only and a manifest repository and fails on a repeat, a lost workspace, or a step this CLI cannot run. The guard took three passes to stop being vacuous, and each pass is recorded in its comments. Anchoring the `Next:` reader to end-of-line missed a footer that continues after the backtick. Keying a step on raw argv made `audit --host` and `audit --host --json` different nodes, so the loop looked finite. Only the workspace assertion finally reproduced the defect: replayed against pre-fix `main`, both `audit --host` chains now fail, along with the two verdict-word cases. Closes #650. Co-authored-by: Claude Opus 5 feat(check): compare against the detected base by default (#649) (#671) `shipgate check` compared the working tree with `HEAD`. On a branch whose changes were already committed it therefore answered `allow` with an empty change set — truthfully reporting "nothing is uncommitted" to someone asking "what does this branch change". Reaching the real answer took `--base --head HEAD`, a pair nothing in `--help` suggested and which the CLI rejected unless both were supplied. A flagless run now compares the detected default branch's merge base against the working tree: one comparison spanning committed branch work and uncommitted edits, which is the shape `working_tree_context` already provides and the verifier already uses for exactly this reason. `--base` and `--head` became independent, and `subject.base` names the ref actually compared — a verdict whose base is unnamed cannot be reviewed. Two things the issue's sketch got wrong, handled deliberately: `detect_default_base` refuses a local `main` on purpose, because a remote is the authority it might be stale against. Rather than overriding that, the fallback is narrowed to the case the rationale does not cover — a repository with no remote at all, where the local branch is the only base in existence. Opt-in, so `verify` keeps its behaviour exactly. That detector also returns nothing when HEAD *is* the default branch, which would have made an honest stop out of the commonest situation there is. Being on the default branch and having no base at all are now separated: the first keeps the complete working-tree answer, only the second stops and names `--base`. `verify_command_for` required both refs or neither — a mirror of the CLI rule this change removes. It now emits whichever ref was resolved, because `verify` accepts either alone, and a command that disagrees with the check that emitted it is the recurring second-implementation defect. 12 cases in `tests/test_check_default_comparison.py` covering every rung. Replayed against pre-fix `main`: 8 fail, and the 4 that pass are the default-branch, empty-ref and option-smuggling cases that must not change. `tests/test_codex_boundary_check.py::test_codex_check_rejects_one_sided_git_refs` pinned the removed rule; it now pins the new one, plus the empty-ref refusal it was also covering. Two line-pinned entries in the static-only allowlist moved with the code they point at. Closes #649. Co-authored-by: Claude Opus 5 fix(host): keep machine-written tool caches out of the host inventory (#598) (#670) `shipgate check` run while the pytest suite was active returned `human_review_required`, every permission false, with "Directory inventory could not complete at tests/__pycache__". A quiescent rerun passed. The refusal was correct. `IdentityBoundReadSession` revalidates every directory it scanned, so a `.pyc` appearing — or the directory being removed — between the scan and its revalidation is exactly the incoherent read the session exists to refuse. The defect is upstream of that: a bytecode cache was an identity-bound input at all. No host reads configuration from one, so walking it buys no coverage while exposing every local run to unrelated churn, and one such directory failing discards the whole repository inventory. `__pycache__`, `.pytest_cache`, `.mypy_cache`, `.ruff_cache`, `.tox` and `.nox` join the existing skip set beside `.git`, `node_modules`, `site-packages` and `.venv`. This is the boundary the issue asked to examine — whether generated cache inventory is needed for recognized host coverage — not a blanket skip and not an incomplete read turned into success. Nothing else moves: a recognized host directory that changes mid-read is still refused and yields no grants, an unreadable one is still fail-closed, and a rerun after the writer goes quiet performs a real read rather than serving a cached denial or a cached completion. 10 cases in `tests/test_host_inventory_stability.py`. Replayed against pre-fix `main`: the 7 cache-boundary cases fail and the 3 integrity guards pass in both states. Deliberately not done: a new `input_unstable` failure reason. It was drafted, then reverted — the existing model already names a typed cause, phase and source with no exception-text inference, and the new value was reachable only from a narrow parent-identity path I could not demonstrate end to end. An enum value that no test can reach is surface, not a diagnostic. Closes #598. fix(scan): one availability contract for required tool sources (#585) (#669) * fix(scan): one availability contract for required tool sources (#585) `docs/diagnostics.md` and `AGENTS.md` both publish it: a required `tool_sources[].path` that does not resolve is `InputParseError(3)` from `scan`. That was true of the shared loaders and `mcp_server_source`, and false of `openai_agents_sdk`, which returned a source warning and let a required, absent entrypoint finish as an advisory exit-0 scan with `insufficient_evidence`. An integration reading execution status to tell bad input from a completed scan got a different answer depending on which reader the source happened to use. Scan now applies one availability precondition before any adapter runs, reading the same resolver `doctor` renders as `SHIP-DIAG-MISSING-SOURCE-FILE`, so the two agree by construction instead of by coincidence. The error names each offending source id, its declared path, and whether it is missing or escapes the manifest directory. Deliberately unchanged: `optional: true` sources keep their warning and `coverage_recovery` evidence. Nothing is invented to replace a missing input — it is named. `verify` applies the precondition per tree. A base commit whose manifest declares a path absent from that tree now reports `base_status: "scan_failed"` with the reason in `base_notes` and no capability delta, instead of claiming the base scanned successfully with zero tools and reporting a delta it had not read. Measured on the same fixture: the merge verdict is `insufficient_evidence` either way, and the note says the head gate is unchanged. Two existing tests moved with the decision rather than around it. `test_required_sdk_source_keeps_existing_advisory_execution_result` was and its comment named #585 as the owner of the contract itself; it keeps that invariant and now asserts exit 3 on both paths. Four dispatch-order fixtures in `test_adapter_registry.py` used placeholder paths that never existed; they are marked `optional`, which is what they always were. 15 cases in `tests/test_required_source_availability.py`, including one that fails if any adapter runs before the check — the point is a precondition, not another reader that happens to raise. Replayed against pre-fix `main`: 10 fail, and the 5 that pass are the optional-source and doctor cases that must not change. Closes #585. Co-Authored-By: Claude Opus 5 * fix(scan): say "not found" for an absent source, matching the shared vocabulary CI caught it: `tests/test_absent_input_messages.py` pins one vocabulary for an *absent* input across manifest, policy-pack, baseline and tool-source messages, so an absence is never described in the words used for a malformed input. The new refusal said "does not exist", which is the workspace class's wording, not the tool-source class's. The missing case now reads "'src' at 'agent.py' was not found"; the containment case is unchanged. Guard, the new tests, the CHANGELOG and llms-full.txt move together. fix(diff): read every counted hunk row, so header-shaped content is not lost (#611) (#668) The shared unified-diff parser decided what a line was from how it was spelled. A removed `---` renders as `----` and a removed `-- note` renders as `--- note`, so both matched the file-header prefixes: the first dropped a counted row, and the second replaced the file's identity with the invalid-path sentinel — a file's own content rewriting its path. Hunk state and the header's declared row counts now decide. While a hunk still owes rows, any line opening with a unified-diff marker is that hunk's data and is dispatched on its first character alone; `\ No newline at end of file` is a marker and consumes no budget. A plain Markdown horizontal-rule removal therefore reaches a complete structural comparison instead of an unnecessary `human_review_required`. Nothing is loosened. A truncated or over-declared hunk ends at the first line that cannot be hunk data, keeps its short row list, and is still refused by the declared-count comparison in `core.instruction_structure`; the next file's header is still read as a header. Path identity, rename, new/deleted and malformed-header contracts are unchanged. 16 regression cases in `tests/test_boundary_diff_hunks.py`, built from real `git diff` output where the defect is about what Git emits and from handwritten text where the point is input Git would never produce. Replayed against pre-fix `main`: the 11 defect cases fail, the 5 contract-preservation cases pass in both states. Closes #611. docs(roadmap): add the roadmap to v1.0 as plan of record (#643) (#667) * docs(roadmap): add the adoption roadmap to v1.0 as plan of record (#643) Insert an "Adoption roadmap to v1.0" section ahead of the release-qualification record: v1.0 is defined by external adoption (install, first diff in two commands, zero-noise on benign PRs, repeat use), with two release lines (advisory pre-releases on a 14-day cadence; qualified gate on evidence) and three milestones (#644 M1 installable and navigable, #645 M2 diff first value, Refs #643, #644, #645, #646. Co-Authored-By: Claude Fable 5.1 * docs(roadmap): source the adoption targets from the registry, not from prose `test_no_published_adoption_number_is_larger_than_the_registry` failed on the new section: "10 named external users" matches the adopters claim grammar (#475 rule 1) and ADOPTERS.md can source 0 external entries. The guard is right — a target written in claim shape reads as a claim. State both adoption thresholds as registry states instead: ADOPTERS.md reaches >= 10 external entries with >= 5 of them active weekly, and the gate track continues while ADOPTERS.md shows >= 5 weekly-active entries. The numbers now name where they will be read from, and the file produces no adoption-claim match at all rather than relying on an exemption. Co-Authored-By: Claude Opus 5 * docs(roadmap): v1.0 is achievable alone — adoption numbers move after the release Gating v1.0 on adopter counts is circular: people adopt what has been released, so a release that waits for adopters never ships. The earlier draft made ">= 10 external entries" and second-change observations exit criteria; both are now removed from the v1.0 bar. The rule the section now states: v1.0 contains only what this repository can achieve on its own, measurable against public repositories and committed fixtures. Milestones are renamed v1.0 M1 / M2 / M3 and their exits carry only self-verifiable evidence (installable build, two-command first value, row precision and benign quiet on the host-config corpus, cold-start reach, reader precision, host-family coverage, a completed human route). Conversations, the pilot, the registry, the GitHub App and delta consumers move to an explicit "After v1.0" track that never blocks a release; conversations can still start immediately. The gate track stays eligible for v1.0 because it is work this repository controls; whether it ships there is a product decision in #572, not an adoption threshold. Refs #643, #644, #645, #646. Attribute each finding to the bound this change moved (#515) (#641) * Attribute each finding to the bound this change moved (#515) `finding_deltas` answers a question about identity: is this fingerprint in the base report? A weakness the pull request never touched, a change that bounds the capability without perfecting it, and a change that widens the capability all get the same answer — one row in `unchanged_findings`. `tool_surface_diff.finding_attributions[]` answers the other question, for the findings the shipped comparison profiles can speak about. It reads no source and runs no second diff: every direction is copied verbatim out of a row `compare_operations` or `compare_guard_dependencies` already published, and `evidence[].effect` is the one place the two vocabularies are reconciled. Three refusals decide most rows. Only evidence naming a finding's own fingerprint classifies; same-capability evidence can withdraw a negative claim but never establish one; and absence is never agreement — unattributed findings are counted in the diff notes, a run with no base emits no rows at all, and two active findings sharing one fingerprint are unresolvable. Dependency coverage stays `incomplete` and `finding_exclusion_eligible` stays `false` on every row: `standing_weakness` says no *modeled* bound changed, never that nothing changed. No `--scope` option, no receipt question, no decision consumption, no check id, no schema version bump, and no finding, severity, baseline, exit code or release decision changes. Co-Authored-By: Claude Opus 5 * Address review: publish the uncompared count where it cannot be truncated Four findings from the review pass on #641. The count of findings no profile could compare was published only as a trailing diff note, and `report.md` renders three notes. On any enabled diff three are already spoken for, so the one statement that stops an empty attribution reading as a clean one never reached a reader. It is now `tool_surface_diff.unattributed_findings`, printed under the section it qualifies, with one shared spelling of the sentence. A finding carrying neither a fingerprint nor an id was skipped before that counter, so it appeared in neither the rows nor the count — both fields are optional and a plugin check can supply neither. It is counted now. The `improved_not_resolved` reason said "the finding still stands" for rows whose own `identity` is not a base match, contradicting the identity buckets on the same page. It now says what the identity supports. `_guard_pointer` substituted `tool_path:tool_line` when the guard location was unknown, publishing the tool declaration under a `guard_predicate` axis. An unknown guard location now publishes no pointer. Co-Authored-By: Claude Opus 5 * Address review round 2: state the withdrawal rule the code actually applies Two findings from the second review pass on #641. The contract said any capability-linked movement, "including an axis the profile could not compare", retracts both negative claims. The code retracts `improved_not_resolved` only on a demonstrated widening. The code is right and the sentence was not a guard: the two claims claim different amounts. `standing_weakness` says nothing anywhere on the capability got worse, so an axis nobody could compare defeats it. `improved_not_resolved` says one named bound demonstrably narrowed, which an uncomparable axis elsewhere does not refute. Both halves are now pinned by a test. The Markdown section always led with "What this change did to the bound each finding depends on", then printed only counts whenever no row reached a direction — the normal shape for a repository whose only profile publishes no fingerprint. The lead now says what the section actually contains, and the stray blank line before the counts is gone. Co-Authored-By: Claude Opus 5 * Address review round 3: name both causes, and where the projection stops Two findings from the third review pass on #641. `unattributed_sentence` now covers both causes it counts. After the round-1 fix it also counts a finding that published neither a fingerprint nor an id, and saying those had "no comparison profile evidence" sends a reader to look at profile coverage for a finding no coverage could have reached. The contract did not say where the projection is published. It is `report.json` and `report.md`, like the two profile comparisons it reads. `packet.json` and the PR comment carry the diff notes and both truncate them, so neither restates a class; the packet's "Finding identities" line stays an identity statement and the contract now points to report.json for its attribution. Co-Authored-By: Claude Opus 5 * Address review round 4: correct three comments that describe an older join Three findings from the fourth review pass on #641, all the same class. `AttributionLink` and the contract described the same-capability join as a shared tool id. It is canonical identity: a shared tool id, or a capability id the finding cites in `capability_refs`. Both paths are live on the paired fixture, and the rule decides whether a `standing_weakness` is withdrawn, so a test now pins the capability-id path on a finding carrying no tool id at all. The `improved_not_resolved` comment kept the unconditional "the finding still stands" that round 2 removed from the emitted reason, and the `unresolved` comment did not name the unchanged-bounds-under-a-changed-identity case. `FindingIdentity` documented its `unresolved` value as arising when there is no base, which is the one condition under which no rows are emitted at all. fix: bind base-scan cache to effective engine identity (#639) fix: bind implicit Codex plugin default selection (#635) (#637) Preserve Codex plugin component lookup identity before resolution (#636) * fix(verify): preserve Codex plugin component lookup identity * fix(verify): retain wrong-kind plugin component refusals * fix(verify): bind direct skill-file kind before optional recovery fix: bind enumerated input directory membership (#630) (#634) fix: reconfirm recorded verification input origins (#627) (#631) Connect declared OpenAPI operations to approval predicates (#628) * Connect declared OpenAPI operations to approval predicates * Refuse noncanonical method keys in operation attribution Fill four beta sourcing gaps from pinned source changes (#626) Require the existing CI checks before merging to main (#625) Add a separate beta sourcing inventory with explicit gaps (#624) Align release recovery with enforced tag protection (#622) Record merged delivery, hosted smoke and effective release protections (#621) * docs: reconcile merged delivery and hosted distribution evidence (#572) * docs: record effective release immutability and remaining protections (#573) Record maintenance responsibilities and release recovery obligations (#619) Route supported host-only repositories into their first review (#614) * fix(discovery): route host-only workspaces to host review (#568) * docs: preserve mixed-workspace routing in the first-run guide * fix(discovery): stop host applicability from breaking settled answers Independent review of #614. The host census reached two ordinary repository shapes and turned a correct, settled answer into a stop on each of them. A symlink to a regular file defeated the product-wide negative. Every symlink landed in `host_discovery_incomplete_paths`, because a `**/` glob makes `_symlink_may_hide_boundary_glob` true for any name, so a plain `docs/README.md -> ../README.md` suppressed `SHIP-DIAG-NON-AGENT-LIBRARY`, flipped `setup_not_applicable` to a human stop, and withheld the trigger stop conditions — about a link with no descendants to conceal. A link is now concealment only when it resolves to a directory; the type is read, the link is never followed into, and anything untypable stays concealment. A synthesized child under an ancestor link that resolves to a non-directory is dropped rather than published as `unresolved`. An unreadable directory made `detect`, `init` and `bootstrap` exit 4, with no way back: the census raised on any `HostInventoryReadError`, ahead of every other classification step, and its published recovery ("rerun detect") could never succeed. `audit --host` owns the same walk and does not do this — it records the failure as an `inventory_failures` row and carries on — so an applicability hint was stricter than the audit it routes to. The census now names the path that stopped it in `host_discovery_incomplete_paths` and publishes no candidates, which withholds the negative and the trigger stop while leaving the framework, source and scope answers standing. Also: `detect`'s human output no longer prints a host heading over a `doctor` route on an adopted workspace, prints the excluded-sources block the other negative routes print, and gives an adopted workspace a `Next:` line at all; `not_applicable_host_review` is added to every published `manifest_status` enum; truncated subject lists say how many they left out; `bootstrap` calls the one `needs_host_route_payload` instead of restating it; `init`'s hand-off payload keeps `tool_surface_origin` and `control_pack`; the ledger reason is `census_not_seen_through` rather than `source_rejected`; and the parity harness gains host-only shapes so its two new comparison rows can fail. Co-Authored-By: Claude Opus 5 * Close the three deferred items from review loop 2 `--local-review` hand-off: behavior kept, reason stated. The line is not how explicit a flag is, it is whether the flag uses discovery at all. `--minimal` short-circuits before `detect_workspace` is called, so there is no classification for a host-only route to act on; every other mode (`--ci`, `--claude-code`, `--agent-instructions`, `--local-review`) renders its manifest from that classification, and on a host-only repository would write a manifest declaring nothing — the dead end #568 removes. STABILITY.md and the recipes now say so instead of listing `--minimal` as an unexplained exception. Incomplete-census rows are `route_blocked`, not `not_claimed`. The run acted on them: withholding the product-wide negative is the act, and `total - gated` is documented as what was "recorded and deliberately not acted on". Unconditional, following `walk_capped`, which is `route_blocked` while `detect` publishes a runnable higher-cap retry — so the token means this stage withheld its verdict, not that routing stopped. The layering objection raised in review does not apply; no predicate is needed. `host_config_adapters_for_path` still carries an instruction-prefix denylist beside a registry lookup, which cannot be made registry-driven without a new registry field. Made the drift loud instead: every registry entry is pinned as host-config or not, so a new adapter entry fails the test rather than silently widening the candidate set. Each expansion ends in `.json` so a negative cannot pass through the suffix gate, and the two prefix exclusions are asserted directly. Verified against a hypothetical `.cursor/hooks/*.json` adapter, which the guard rejects by name. Bind candidate-generated CI to the intended release distribution (#616) * fix(release): bind candidate CI to its source and package version (#570) * test: classify candidate provenance build hook * docs: point candidate provenance to real backend tests * docs: reconcile release roadmap delivery and remaining gates * fix(release): scope candidate-provenance failures to what they decide (#570) Review follow-ups on the candidate source record. None changes what the record binds or what the provenance gate accepts. * The record decides one thing — the Action ref `init --ci` writes into an adopter's repository — but it was read while rendering a module-level constant, so a corrupt record in a stamped wheel raised during the import of `agents_shipgate.cli.discovery`. Every command died with a traceback, `doctor` included, which is the command the runbook sends an operator to when an install looks wrong. Serving `WORKFLOW_TEMPLATE` lazily keeps the refusal on the emission path and nowhere else; a subprocess regression test imports the CLI with a corrupt record and asserts the workflow is still refused. An operator override no longer has to be set before import to be honoured. * `os.open()` without `O_BINARY` is a text-mode descriptor on Windows. The Action's exact-wheel route read a zip through one, so every CRLF pair was rewritten and the first 0x1A byte ended the capture: a correct wheel could never install on a Windows runner. Added at both call sites; the stdlib ORs it for the same reason, and it is 0 elsewhere. * `release_engine_smoke.py prepare` commits fixture history onto the current HEAD. Standing at the candidate commit is exactly what a release operator's own checkout looks like, so that check cannot tell a runner from their tree; it now requires `--disposable-checkout`. It also wrote a commit identity into `.git/config`, which would have outlived the run and re-authored whatever they committed next — passed per-commit instead. * The stamp env var was only checked for absence inside the sealer workflow. A preview or CI build that acquired it would emit adopter pins no published channel carries, so the guard now sweeps every workflow in the repository. * Registered the Action's local-wheel install route in the surface registry — it names no channel and claims no `executable_pin` — and recorded both operator-facing behaviours in the runbook. Full suite: 9,462 passed, 5 skipped. Ruff clean. Co-Authored-By: Claude Opus 5 * test(release): keep the distribution smoke on the hash-locked backend (#570) `python -m hatchling build` does not enforce `[build-system] requires` the way the sealer's `python -m build` does, so the only thing holding the smoke to the release backend is the hash-locked install ahead of it. An index-resolved hatchling stamps a different `Generator:` into `.dist-info/WHEEL`, and the "exact source candidate" would then not be the bytes the sealer compares. Compare instruction structure consistently across review and edit workflows (#612) * Compare supported instruction structure across review workflows * Correct the current preflight discriminator in discovery documentation Record source comparison delivery and deferred attribution prerequisite (#608) Compare complete Boolean source functions and literal Agent membership (#606) * Compare closed Boolean source functions and literal Agent membership (#557) * Bound helper evaluation with a validated Boolean table Make sample goldens reproducible from the source tree (#605) * Make sample golden regeneration reproducible * Keep the CRLF golden test valid on CRLF checkouts Record v1 delivery and remaining release obligations (#604) Bound FastMCP injection claims to whole-signature evidence (#603) * Bound FastMCP injection to whole-signature evidence * Keep invalid typing arity visible in Context signatures Replace historical MVP defaults with accepted decision index (#602) Define compact current-control versus full receipt validation (#600) Compare imported SDK guard predicates with bound input evidence (#599) * Track imported SDK guard predicate evidence (#557) * Refuse looping guard dependency paths without crashing Use a canonical private path for pilot baselines (#550) (#595) Preserve product identity recovery from discovery evidence (#594) * Preserve product identity recovery from discovery evidence (#543) * Resolve identity recovery beside the loaded manifest (#543 review) Identify failed authorization context prerequisites (#548) (#591) Explain host inventory failures from captured read facts (#547) (#589) Explain source coverage recovery from typed SDK evidence (#561) (#587) fix(discovery): match same-build integration surface enumeration (#553) (#583) fix(miner): retain typed scoped base input obligations (#564) (#582) docs(release): reconcile qualification instructions with approved policy (#579) fix(qualification): score the same bytes validated by input digests (#578) fix(control): revalidate full live currency before returning authority (#576) docs: define evidence-backed v1.0 release gates and roadmap (#574) Benchmark: record fixed-history regression and sequence remaining core work (#565) * bench: record fixed W37 history and unresolved release-exit regression (#312) * docs: sequence remaining core work from the measured workflow failures * test: preserve the historical driver identity across future improvements Qualification: report actual coverage misses without changing safety thresholds (#562) * fix(qualification): expose actual coverage misses and metric applicability (#520) * fix(qualification): validate malformed legacy metric containers (#520) fix(report): distinguish provisional effects from static evidence (#357) (#560) fix(diff): retain evidence behind matched finding identities (#515) (#558) feat(review): show current questions in PRs and preserve publication fallback (#337) (#556) feat(review): evaluate bound external human decisions (#537) (#554) feat(review): bind a checkable human review request without granting authority (#536) (#551) Deprecate the path-only agent instruction weakening check (#516) (#549) Route init generation defects to product review instead of repository repair (#328) (#546) Reject product identities declared only in fixtures or templates (#533) (#544) FastMCP signature projection resolves the injected Context and discloses unsupported schema (#539) (#541) * Resolve the injected Context by its binding, not its spelling (#539) The FastMCP signature projection decided whether a parameter was the framework's request context by comparing the last token of its annotation with `Context`. The SDK does not: its injection reader resolves the type hint and checks class identity against its own `Context`, and the tool builder skips only the parameter that check identified. Matching the spelling instead was wrong in both directions, and both were reproduced through the production loader on `20968551e`: * `class Context(BaseModel)` — the caller's own model — was erased, and the catalog published `update() -> str`: a tool presented as taking no arguments, with a required `context.account_id` behind it, on `enumerated` evidence and no warning. * `from mcp.server.fastmcp import Context as RequestContext` was not matched at all, so the value the framework supplies became a required `string` the caller is asked for. The reader now answers with the module's own import and class table — the same machinery, the same prefix rule and the same refusal to answer on a doubly-bound name that already resolve a `FastMCP(...)` construction. It adds one thing: the base list, because the SDK injects any *subclass* of its context, so a class written in the module is a caller input only once the bases beside it say so. The context module family is derived from `PYTHON_SERVER_CONSTRUCTORS`, because the request context ships beside the server class and a hand-written list would go stale on exactly the rename that hid 229 tools in #484. `SignatureParameter.injection` is a tri-state, not a boolean: "not established" is a different answer from "established as a caller input" and publishing them as one is how a parameter disappears. Where the module does not settle it — a name two statements bind, a relative import outside the walk — the parameter stays in the inventory and the tool carries `unresolved_context_identity` in `extraction["surface_gaps"]`, which holds its surface at `partial` without touching the route, the tool name or the `medium` ceiling. Nothing is imported and nothing is executed: a forward reference is parsed as the source it is, and every name resolves in the scope the annotation was written in. Both readers, and the shared corpus that keeps them honest: the seven new Python cases live in `tests/mcp_idiom_corpus.py` and the parity comparison now includes `injection`, so a port that returned the dataclass default would agree on every name and disagree about every signature. Co-Authored-By: Claude Opus 5 * Publish the type an annotation denotes, or none at all (#539) `json_schema_type` reads the *rendered* annotation and falls back to `string` for everything it does not recognise, so `int | None`, `Annotated[int, Field(ge=1)]`, `typing.List[str]` and a Pydantic model all shipped as concrete `{"type": "string"}` property schemas — a guess presented as read evidence, on a route reporting `enumerated` with no warning. The same fallback built `output_schema` from the return annotation. The reader now answers from the annotation's **tree**, against the same binding table the injection identity uses, so a spelling it does not understand answers "no type" rather than quietly becoming a scalar. Optionals are the type they wrap — our property carries one type, and `T | None` is the framework's own spelling for "may be omitted", so naming `T` is the kind the source names while `string` for an integer is a different one. Containers are represented when their elements are, and `Annotated[T, ...]` is `T`, because that is the spelling FastMCP's own documentation uses for every constrained parameter. `str`, `list` and the `typing` aliases are canonical only while the module has not rebound them. Where the type was not read, nothing is published: the property is `{}` — JSON Schema for "any value" — `ToolParameter.type` is `None`, and the tool carries `untyped_parameter` or `unrepresentable_annotation` in `extraction["surface_gaps"]`, the vocabulary `google_adk` established for the same two facts and which a test now pins equal. An *absent* return annotation stays the honest omission it was: `output_schema` is `{}` and that is not a gap. `parameters`, `input_schema`, `function_signature` and the action projection's input fields therefore describe one set of evidence, and a `-> SomeModel` tool no longer draws `SHIP-SCHEMA-FREEFORM-OUTPUT` for a string it does not return. The determinism-boundary page is regenerated from the coverage cell, which now says what the route does with a signature it cannot fully read. Co-Authored-By: Claude Opus 5 * Measure the annotation reader against two live servers (#539) A self-review pass ran the finished reader over `redis/mcp-redis` and `chroma-core/chroma-mcp` and read every annotation it refused. Two of the refusals were the reader's own defects rather than the source's. **A builtin outside the published-type table read as an unresolved injection identity.** The identity check asked whether an unbound name was in `PYTHON_ANNOTATION_JSON_TYPES`, but that table answers a different question — what type is *published* — so `bytes`, `object` and `bytearray` fell through it and `Union[str, bytes, int, float, dict]` came back as a question about who supplies the parameter. Nothing named by a builtin is the framework's request context, so identity now reads its own set. The same change closes a second hole in the other direction: `List` was canonical in that table, so an unbound `List[str]` — a `NameError` in any module Python will run — was published as an array. A `typing` alias means something only when an import into `typing` binds it. **A container's kind does not depend on what it holds.** Requiring every element to be representable reported a type this reader *did* read as one it did not: `Dict[str, Any]` is the most common return spelling in `mcp-redis`, and `{"type": "object"}` is true of it. This projection publishes no element schema for any annotation — a bare `list` included — so the element requirement bought nothing. The exception is a mapping's key, which is the part that decides whether the value is a JSON object at all, so `dict[int, str]` is still refused. Measured after: `mcp-redis` 53 tools named, 21 → 14 held at `partial`, and every annotation still unreadable there is a union of genuinely different JSON kinds or an opaque library alias; `chroma-mcp` reads 13 of 13 with no gap. Four adversarial corpus cases for the branches this pass exercised — the context re-exported from its defining module, a rebound builtin and a rebound `typing` alias, and a union whose arms disagree — and a 16-case perturbation sweep over every decision both slices added: all 16 caught. Co-Authored-By: Claude Opus 5 * Assert the published action names the same inputs as the schema (#539) The last of the issue's acceptance criteria, and the one that says why the first reproduction mattered downstream: `input_fields` and `required_input_fields` on the published action are derived from `tool.parameters`, so a parameter erased from the inventory is erased from what a reviewer reads and from the diff that would have flagged its arrival. Asserted through `build_action` rather than inferred from the loader. Co-Authored-By: Claude Opus 5 * Prove the base cycle guard, and correct a comment (#539) `# pragma: no cover - a class cannot be its own base` was wrong twice over. A class cannot, but two can be each other's: `class A(B)` beside `class B(A)` raises at import time and **parses**, and this reader reads source rather than running it. Deleting the guard now raises `RecursionError` out of the corpus case that covers it, where before it raised nothing because nothing asked. Also corrects the comment beside the bound-name branch of `_python_annotation_symbol`: a bound spelling is never the builtin; what it can still be is a `typing` alias, which is the only way `List` and `Optional` arrive. Co-Authored-By: Claude Opus 5 * Preserve nullable and unresolved Context evidence in FastMCP signatures (#539) --------- Co-authored-by: Claude Opus 5 Name the ADK MCP endpoint and credential-reference change (#538) (#540) * Name the ADK MCP endpoint and credential-reference change (#538) An ADK agent that mounts a remote MCP server declares its authority in the constructor: *this agent will call whatever `https://…/mcp` advertises, under ``, restricted to ``*. The reader discarded all of it. Two workspaces differing only in the endpoint literal and the `os.environ[...]` key — same lines, same filter — produced byte-identical reports: empty `capability_facts`, empty `tool_surface_facts`, and a finding carrying the toolset kind, the source line and one agent name. Three layers, all reusing machinery that already exists. **Read.** `_read_mcp_connection` parses `connection_params=` with `ast` and keeps the literal endpoint, the `os.environ` / `os.getenv` *names* the credentials are read from, the transport and the literal `tool_filter`. Every axis carries a status, so "the reader did not look", "the argument is not there" and "it is there and is not readable" stay three different answers, and a per-binding limitation code names each doubt. **Carry.** Each binding rides base-to-head as four independent per-axis `ToolSurfacePolicyFact` rows — the carriage `core/toolkit_scope.py` already uses, so no report schema bump, no new check id, no new committed inventory or command. `inventory_path` is deliberately not one of the four, so supplying the reviewed inventory the scan asks for cannot clear a connection delta. **Project.** Into the existing `capability_change` block, which never gates. `tool` stays empty: no remote leaf is synthesized. A changed endpoint or credential reference takes the block's documented opaque-direction bucket with a rationale that says, in the reviewer's own output, that the direction is not established and that a host or variable name proves no privilege level. The tool filter is the one axis where a direction is claimed. Identity is `::` and excludes the source line, so moving, reflowing or commenting the call produces no delta, while two agents on one endpoint — and one toolset shared by two agents — keep distinct attribution. Credentials written literally into source are redacted at the reader and never hashed; that limit is published rather than implied. Nothing is imported, constructed, connected to or looked up in the environment. `release_decision.decision` remains the only gate and every qualification threshold is untouched. Co-Authored-By: Claude Opus 5 * Name the remote-binding member shape in the agent contract (#538) A consumer that switches on `subject_kind` and `tool` needs to know what an empty `tool` on a `scope` member means, and what it deliberately does not claim. Also drops a private alias `is_sensitive_key` had just made redundant. Co-Authored-By: Claude Opus 5 * Address review: three framework semantics and two silent losses (#538) Five findings on 912419f4. Three were the same mistake — reasoning about a framework's semantics from the shape of the data instead of from the framework. Verified against adk-python 2.8.0's own source, not from memory. **`tool_filter=[]` is not a filter of nothing, it is no filter.** `BaseToolset._is_tool_selected` returns `True` for any falsy filter, so an empty list exposes every advertised tool. Comparing it as the empty *set* reported the widest state as the narrowest and emitted a high-confidence `narrowed` member saying "search no longer reachable" about a change that made search reachable along with everything else. Direction now asks whether a side *bounds* anything; two unbounded sides produce no member at all, because the source text moved and the authority did not. **A stdio connection does not carry `env` inline.** `StdioConnectionParams` nests a `StdioServerParameters` under `server_params`, so reading only the outer call answered "this binding has no credential" about one that plainly did — and a changed reference then produced no delta. The nested call is now resolved, inline or through a module-level name bound exactly once, and only when it is a recognized constructor: `server_params=build_params()` reports `unresolved` rather than the same false absence one level down. **A URL query key is classified after decoding, and against a credential-name rule.** `api%5Fkey` is `api_key` to every server that reads it, and `access_token` was in no vocabulary the redaction pass held; both reached `report.json` intact. New `privacy.is_credential_key` is the exact vocabulary plus a suffix rule, and deliberately excludes a bare `key` so `sort_key` stays an ordinary parameter — claiming a credential where there is none is its own false statement. The same predicate now classifies header and `env` keys. Two more were silent losses rather than wrong answers. **A capability member's id is hashed from its subject**, and a subject without the configured source id merged two same-named agents from different sources into one member, deleting one source's endpoint change in `_dedup_members`. The subject now carries the whole identity, spelled `agent [source_id]`. **A hardcoded credential added beside an existing environment reference was recorded only as a limitation**, which the carried summary and hash never saw, so a credential being *added* produced no delta. The credential axis now lists every entry, with `` in place of a value it will not publish. Seven perturbations, one per fix, all caught by the new guards. Reach one useful review from the human entry path (#498) (#534) * Reach one useful review from the human entry path (#498) The README was 933 lines with 28 second-level sections, and the quickstart opened with three commands before the reader had seen a single result. Both were also wrong in ways nobody was checking, because `README.md` and `docs/quickstart.md` were recorded in `docs/distribution-surfaces.md` as "repository documentation" and registered as a surface nowhere: - the quickstart's first command was `shipgate check --format agent-boundary-json`, and the release the same page tells a reader to install rejects it outright — `v0.15.0` accepts only `codex-boundary-json`; - its placeholder step sent a coding agent to the README for `agent.declared_purpose`, a declaration only a person may make; - the README's flagship "what your PR sees" block quoted a comment with an `### Agents Shipgate result: block` heading and an `Impact | Change | Subject | Why` table, called it verbatim, and no code path rendered any of it. `block` is not a value any verdict field takes. The README is now a landing page: one before/after capability change, one demo, one install, the accuracy numbers that are still zero, and an audience-routing table. `docs/quickstart.md` is one review end to end on the committed `ai_generated_refund_pr` sample — what changed, why the top result matters, what the run did not establish, that exit zero is not merge permission, and who owns the next action — before either adoption route, and it opens with a channel table saying which build provides which commands. Displaced material moved into the docs that own it, with a mapping table from every retired anchor. Three guards keep it there: - `README.md` + `docs/quickstart.md` are a registered `human_entry_path` surface and `AGENTS.md` + `docs/agent-recipes.md` + `docs/agents/` + `docs/target-repo-agent-snippets.md` a registered `agent_instructions` surface, so #497's placeholder-ownership and pin rules cover them; - `test_the_human_entry_path_states_what_the_published_build_provides` reads the `--format` values the published tag accepts out of the tag, the way #506 reads that tag's `CONTRACT_VERSION`, and fails if the entry path teaches one it does not; - `test_the_entry_path_quotes_lines_the_pr_comment_actually_renders` runs the fixture and compares every quoted comment line against the artifact. The vocabulary reader learned to see a Markdown reference table, the only shape the entry path states a verdict set in; a dropped `unknown` row now fails. Four unresolvable `uvx agents-shipgate@0.18.0` pins in `docs/incidents/` and `samples/README.md` named a version the index does not carry. Co-Authored-By: Claude Opus 5 * Address review: two guard defects and an incomplete anchor map Three findings from a review pass over the branch, all in this change's own work. `_first_column_sets` required a trailing newline on every table row, so a verdict table that ends the file lost its last row: the complete five-row merge-verdict table returned four members and failed a document that was correct. The newline is optional now. `_published_check_formats` understood only `format_ == "literal"`. The next release will likely spell acceptance the way this tree does — `format_ not in CHECK_FORMATS` — and the old reader would have returned the option's default alone. A *partial* set is worse than an empty one: it passes the emptiness assertion, and `test_the_published_format_reader_sees_the_difference_it_exists_for` then cannot tell that the quickstart's "accepts only" caveat has gone stale. A name is now resolved when its module-level binding is a set of string literals, and any other shape raises rather than reading as "no values". The README's "Moved out of this README" table — this change's own answer to keeping public anchors reachable — omitted five of the 21 retired anchors and claimed the framework notes had moved when half of them stayed. Every retired anchor is now listed, checked by diffing the old README's heading slugs against the table. Co-Authored-By: Claude Opus 5 * Second review pass: a stale contract floor and a misordered strict recipe `docs/agents/README.md` carried "require contract v15 or newer" for the Codex plugin. That number arrived in #271, is stated nowhere else, and contradicts the kit it describes: `adoption-kits/codex-skill/SKILL.md` renders `{{ minimum_control_contract_version }}`, which is 21. A reader checking against 15 accepts a build the plugin's own workflows do not run on. Replaced with the check rather than a literal — run `contract --json` and compare against the floor the kit states, which the kit renders from its own build — so there is no second copy of the number to go stale. `docs/integrations.md`'s strict-mode block told a reader to "save a baseline first" and then showed a bare strict scan as step one, which fails on the backlog. The three commands are now numbered and say what each is for; the first is expected to exit 20, and that is the point of running it. Co-Authored-By: Claude Opus 5 * Address review: four P2 findings on the entry path All four reproduce; each fix is backed by a run against the build in question. **Ownership was routed through a field the CLI does not emit.** `placeholders[]` entries carry `path`, `current` and `line` only — `collect_placeholders()` emits no `owner`, and `placeholder_owner()` is internal. Replaying `init --write --json` on a FastMCP scratch repo returns exactly that, and publishes the routing on `control.next_action` instead: `actor: "human"`, `permissions.edit: false`, and a `why` naming the human-owned fields and their line numbers. That is what `AGENTS.md`, the quickstart, the recipes and troubleshooting now tell a reader to read; `llms-full.txt` regenerated. **The walkthrough promised the published release and quoted the source tree.** Running the fixture from an extracted `v0.15.0` shows the gap: its comment says `Capability delta: +2, 3 modified, -0` with a flat change list and no `support.search_kb` row, its report says `Evidence coverage: static (human review recommended)` with no counts, and it prints the static-verdict boundary nowhere. Steps 3, 4, 5 and 7 now name `v0.15.0` and show what it prints; the scoping sentence claims only what is true of both channels — same commands, same verdict, same exit status. `test_the_entry_path_quotes_lines_the_pr_comment_actually_renders` now renders **both** engines: a quoted line must come from one of them, and a line only one renders must sit in a step that names the release. The reviewer was right that one engine could not catch this. Extracting the tag skips cleanly if its dependencies do not resolve here. Its own `_enclosing_section` had to learn to mask fenced blocks — the PR-comment excerpt quotes the comment's `## Agents Shipgate` heading, so the scan placed the block inside itself. **A scaffold was read as a route.** `tool_surface_origin: "scaffold"` says discovery could not *read* a surface, not that there is none: the FastMCP reproduction publishes real `@server.tool` functions, scaffolds because they sit under an import package, and `audit --host` on it returns three generic instruction trust roots and never mentions the tools. The quickstart and the runbook now keep those repositories on Route A and ask for an export or inventory. `test_runbook_routes_host_boundary_partners_away_from_a_manifest` pinned the disproven sentence as a literal; it keeps its property, pins the corrected instruction, and now also asserts the old one is gone. **The pointer does not validate itself.** After an empty commit in an isolated fixture, `agents-shipgate agent control` exits 4 with `workspace_changed` while reading `current-control.json` still returns the old `review_publishable`. The compressed wording attributed that refusal to the file; `integrations.md`, `docs/agents/README.md` and quickstart step 7 now put the command first and say what reading the file alone cannot tell you. Co-Authored-By: Claude Opus 5 * Address follow-up review: the elision fail-open and an unrunnable command Both findings reproduce, and both are in the fix for the previous round. **`control.next_action.why` is bounded, so absence from it grants nothing.** `_placeholder_review_why()` fits the sentence to the envelope's prose budget: eight unresolved placeholders produce three named paths and `and 3 more in placeholders[]`, and the elided ones — `agent.prohibited_actions[1..2]`, `permissions.scopes[0]` — are human-owned too. "Take the fields `why` names to a person; the rest of `placeholders[]` is repository reading" therefore handed real declarations to the coding agent, which is the fail-open the previous round removed, reintroduced by the sentence that removed it. `actor` routes the *turn*, not individual fields, and that is now what every surface says. Verified in both directions: any unresolved human-owned value gives `actor: "human"` with `permissions.edit: false` — surface everything and stop; only when they are all supplied does a re-run give `actor: "coding_agent"` with the field to replace. `test_surfaces_say_the_routing_prose_elides` holds it, scoped to the surfaces that route on the bounded prose rather than every file naming both tokens — the setup prompts and the runbook enumerate the human-owned fields by name and are not deriving anything from `why`. **Step 7's command did not run where step 2 left the reader.** `fixture run` writes into a temporary copy and reports under `/reports`, so `--workspace .` looked for `./agents-shipgate-reports` in the caller's directory and exited 3 with `Current control is unavailable (missing)` — or validated an unrelated run. The step now prints the `cd` and `--reports-dir reports`, sourced from the run's own `Fixture copy left at` and `Reports:` lines, and names the envelope's real spelling: `agent control` returns top-level `control_state`, while the pointer file is nested `control.state`. `test_the_walkthrough_control_command_runs_in_the_walkthrough` reads the flags out of the page and runs them where step 2 leaves the reader, taking the `cd` literally. Replayed against the previous text: exit 3. Also fixes the second copy of the scaffold-implies-Route-H inference, in the runbook's partner agent prompt — one file over from the copy corrected last round. Read the Python `@mcp.tool` decorator, in both detectors (#484) (#535) * Read the Python `@mcp.tool` decorator, in both detectors (#484) decorator, the largest single registration shape by repositories and by call sites. `redis/mcp-redis`, `chroma-core/chroma-mcp` and `neo4j-contrib/mcp-neo4j` were all in the state MongoDB and Grafana were in before #431 — `detect` returned `is_agent_project: false` for a first-party server publishing dozens of tools. A sixth idiom, `py_fastmcp_decorator`, and a different extraction mechanism rather than a sixth pattern. The name is usually not written down (`@mcp.tool()` on `def dbsize()` registers `dbsize`, which is all 53 of `redis/mcp-redis`'s registrations); the input schema is the annotated signature; and whether the decorator registers anything at all is a binding fact, since `@app.tool` is a registration when `app` is a server and somebody else's decorator otherwise. So it is read with the standard library's parser, it follows the decorated name back to a server construction — across modules, because `redis/mcp-redis` constructs in `src/common/server.py` and decorates in `src/tools/*.py` — and it refuses out loud when it cannot. Two findings changed the design after it was measured. The official Python SDK's v2 renamed `FastMCP` to `MCPServer` and left `mcp.server.fastmcp` as a module that only raises; 41 of `awslabs/mcp`'s servers had already moved, and knowing one name read 105 tools where there are 334. And every one of `neo4j-contrib/mcp-neo4j`'s 40 tools is `name=namespace_prefix + "…"`, so the lexical idioms' rule — no route without a resolved name — would have reported that repository as "not an agent project" because its names are dynamic. A Python site carries its own provenance, so the route is offered and the evidence says so; a committed export cannot displace it either, because containment is the test and an export contains an empty set vacuously. The ceiling stays `medium`, a runtime-built name stays an omission the exclusion ledger accounts for, and a decorator on an object this reader cannot follow — `@self.mcp.tool` on an injected server, `app = create_server()` on a factory's result, both measured in `awslabs/mcp` — is recorded rather than guessed at. `tools/shipgate-detect.py` moved in the same commit (0.5.0 → 0.6.0), because `tests/mcp_idiom_corpus.py` is what keeps it from becoming a different one: the whole Python probe list and the cross-module trees live there once and both readers are driven through all of it, site by site and span by span. Measured, on live checkouts, identical from both detectors: mcp-redis 53 tools / 0 unenumerated, chroma-mcp 13 / 0, mcp-neo4j 0 / 40, awslabs/mcp 334 / 107. A 23-case perturbation sweep caught 22; the one it cannot is the index prefilter, which is a cost bound rather than a decision, and both readers say so at the line. The sweep's first pass found nine untested decisions, including a corpus case that proved "no match" where it meant to prove "two matches" and a unicode regression whose non-ASCII sat on a different line from the offset it was testing. `IDIOM_REGISTRY_VERSION` is 2 and `TRIGGER-MCP-TOOL-REGISTRATION-SOURCE` routes `@mcp.tool`. #441's scaffold tests moved to the factory-bound shape, since their original reproduction is a workspace `detect` now reads. Closes #484 Co-Authored-By: Claude Opus 5 * Address review: dead Poetry gate, class-scope lookup, duplicated route entry Twelve findings from a self-review of the branch. Four changed behaviour: - **The Poetry half of the Python language gate was dead.** `_toml_requirements` returned a table's *constraint* where the distribution is the key, so `mcp = "^1.6.0"` produced `^1.6.0` and never `mcp` — three of the five dependency tables named in `_PYPROJECT_DEPENDENCY_PATHS` could never match, and a Poetry-managed Python MCP server got no route at all. The comment above the branch claimed the opposite. Both halves of a table are admitted now, and the reason it shipped is closed too: every fixture in the file wrote PEP 621, so a parametrised test now drives all five tables plus `requirements*.txt`. - **Class scopes were searched from a nested function.** Python does not close over class scope, so `@mcp.tool()` on a function nested inside a method resolves to the module-level server — this reader walked the enclosing class first and reported `server_binding_not_proven` for a real registration. It now searches a class body only when that is the scope the name is written in, which keeps `@mcp.tool()` written directly in a class body working. Both directions are corpus cases. - **The route repeated its proving module** once per importing file, so `candidate_files` came back with `src/common/server.py` three times for three tool modules. De-duplicated, and the fixture now has two tool modules, because one importer is the shape in which the defect is invisible. - **The test-file *prefix* rule was applied to every language.** It exists because pytest collects `test_*.py`; widening it silently dropped `test_helpers.ts` and `test_util.go` from the surface with no omission to show for it, since an excluded path is never read. Scoped to Python, with both other languages pinned. The rest are cost and clarity: Python files are read once rather than twice (the index pass hands its texts to the scan pass under an explicit `MAX_CACHED_SOURCE_BYTES`, and streams into the index so the whole tree is never held); the line table is built on the first site rather than per file; `_python_site` no longer takes a `receiver` it never used; `_python_module_paths` returns a one-element tuple instead of an empty-string sentinel; PEP 695 `type_params` are visited. `docs/mcp-registration-idioms.md` republished #431's 61/114/110 vendor counts; the current heads measure 61/115/114. Confirmed by running the same clones against `main`'s reader, which gives the same numbers — the servers gained tools, the reader did not change its mind — and the document now says so. Two of the 29 sweep perturbations stay uncaught and both say why at the line: the index prefilter is a cost bound, and a `type_params` subtree cannot bind a name in any module Python will compile (`ast.parse` accepts a walrus in a TypeVar bound; `compile` refuses it). Measured again after the fixes, identical from both detectors: mcp-redis 53/0, chroma-mcp 13/0, mcp-neo4j 0/40, awslabs/mcp 334/107. Co-Authored-By: Claude Opus 5 * Second review pass: an empty module is a cached answer, not a missing one `text = python_texts.pop(path, None) or _source_text(path)` re-reads an empty module, because `""` is the right answer and a falsy one. Asks `is None` instead, in both detectors. The port keeps the single-read index too, rather than only the CLI: leaving the two discoveries structurally different is how a mirror starts drifting, and `MAX_CACHED_SOURCE_BYTES` is now pinned equal across both by the shared vocabulary test. Co-Authored-By: Claude Opus 5 * Say at the line why the cached-read guard cannot be tested by outcome Reading an empty module through `or` returns the same text, so the fix costs a read and never an answer. A sweep reports it untested; the comment now says that is the correct result rather than a gap, alongside the two other lines in this reader that cannot move an answer. Co-Authored-By: Claude Opus 5 * Address review: three ways the registered identity is not the one on the page Five inline findings on #535. All five reproduce, and four are the same class: this reader asserted a name or a schema that something between the `def` and the server can change. - **A decorator below the registration.** Decorators apply bottom-up, so `@mcp.tool()` over `@replace` registers `replace(harmless)` — whose `__name__` and `inspect.signature` need not be the decorated function's. A `functools.wraps` wrapper does preserve both, and all 72 such sites in `awslabs/mcp` are that shape or a plain `return func`, but that is a fact about the decorator's *body*, in another module for most of them, and this reader does not read bodies across modules. The site keeps its provenance and loses its name and signature. Measured cost: awslabs 334 → 262 named, the 72 in the ledger; the six other measured servers have no such site. Proving the wrapper is the deferred increment, and 72 is the evidence for whether it is worth it. - **An unpacked mapping.** `_python_keyword` ignores a keyword whose `arg` is `None`, so `@mcp.tool(**options)` read as "no name given" and fell back to the function's — publishing `harmless` for a tool the mapping registers as `delete_all`. Any `**` in the decorator call now withholds the name. - **An export that only looks like containment.** The empty-names guard protected an all-dynamic surface but not a mixed one: an export naming every tool the reader *could* read displaced the route, so the registration nobody could name was never scanned and reached no exclusion ledger — a measured miss turning back into a silent one. Now one rule for both readings: an export displaces this route only when the route has nothing it left out, unreadable registrations included. - **Ambiguity measured against the wrong universe.** The index held only modules that export a server, so a conflicting `src/common/server.py` binding `mcp` to something else vanished before the uniqueness check and an `archive/src/common/server.py` supplied binding evidence for an import that demonstrably did not name it. Every scanned module is in the index now, exporting nothing where it constructs nothing. - **Conventional parameter names are ordinary MCP inputs.** The server injects its `Context` on the *annotation*; `SKIPPED_TOOL_PARAMETERS`, shared with this package's other Python adapters, holds `config`, `context` and `runtime`, so `configure(config, context, runtime)` published an empty schema. Detection is annotation-only now — `Context`, `Context | None` and `Context[ServerSession, None]`, but not `list[Context]`. Every fix is in both readers and in the shared corpus, including the positive case each boundary must not swallow: a decorator *above* the registration is applied after it and still resolves. Sweep is 34/37; the three uncaught are the inert lines that say so at the line. Measured after, both detectors agreeing on all seven repositories: mcp-redis 53/0, chroma-mcp 13/0, mcp-neo4j 0/40, awslabs/mcp 262/179, mongodb 61/3, grafana 115/0, github 114/3. Scope the unresolvable-root rejection to its own project (#398) (#532) * Scope the unresolvable-root rejection to its own project (#398) A declared application root whose identity cannot be established statically makes nothing selectable — every name still standing is by construction not that root (#324). The rule was applied with repository scope: on adk-samples one file rejected roughly 55 candidates across 25 projects and named a file in a project the reader was not adopting. The rejection is now scoped to the project the root sits in, using the grouping `agent_project_candidates` already publishes — moved behind one `_ProjectAttribution` object so "which project" has a single answer for both readers. A name several projects declare is rejected when any of them is blocked; that is the fail-closed direction. Scoping alone would not have unblocked the reproduction: one culprit was `eval/test_eval_arize.py` inside the project being adopted, the other a scaffolding template under `resources/templates/`. Neither is the application a project ships. The ranker already discounted both for selection and not for rejection, so one predicate now answers both — a template is demoted exactly as a fixture is rather than standing in for the product. The template rule is the directory pair, not a bare `templates/`, which real packages use. Preserved: a project whose own product root is unreadable still refuses with CHANGE_ME, and an unreadable root outside the two non-product shapes still blocks. `tools/shipgate-detect.py` carries the same change (`script_version` 0.6.0) — `agent_name_candidates` is pinned byte for byte. Restructuring both sides surfaced a pre-existing parity break: with no marker above the evidence, the script named the workspace's `requirements.txt` as the project marker where the CLI reported `null`. Fixed and covered. Co-Authored-By: Claude Opus 5 * Address review: name the shared declaration, case-fold, hoist Eight findings from a high-effort review of the previous commit. Fixed: - A value rejected because another project that also declares it is blocked published a rationale naming that project while every other field pointed at the clean site it ranked from — #398's own complaint, one project narrower. The candidate's own project is now preferred when it is blocked, and when it is not, the sentence says the name is also declared in the project being quoted. - `resources/templates` was matched case-sensitively, so `Resources/ Templates` — the same directory on a case-insensitive checkout — kept the #398 symptom. - `attribution.of(fact.path)` was called once per name-evidence site although it is invariant across a file, paying a stat before its cache. - `test_a_bare_templates_directory_still_blocks_selection` asserted only that nothing was selectable, which any other reason satisfies. - The precedence between the two `_block` passes was untested; the new guard puts the resolution-time file first in walk order so it pins precedence rather than ordering. Documented rather than changed: where a workspace's only agent evidence is non-product code, an unreadable root there no longer forces CHANGE_ME. The guard was accidental — a readable root in the same file already had its name written — so this makes the two cases agree. The general rule is `--allow-unresolved-scope`, whose contract already adopts one of several unrelated agents. No change needed: the `zero_install_detector` row in docs/distribution-surfaces.md still states exactly what is true — same claim, same proof, same narrowness column. Co-Authored-By: Claude Opus 5 * Second pass: skip attribution for a file that declares no agent Hoisting `attribution.of(fact.path)` out of the inner evidence loop (the previous commit's efficiency fix) moved it from once per name to once per *parsed file* — including the overwhelming majority that declare no agent at all, where nothing below reads the answer. On a 1000-file repository with five agent files that is 995 stats and project walks traded for the handful the fix set out to save. The loop now skips a file with no name evidence outright, which is both the cheaper and the clearer shape. Also records on both copies of NON_PRODUCT_DIR_SEQUENCES that its entries must be lowercase: the path is case-folded before the comparison, so a capitalised entry would silently match nothing. Co-Authored-By: Claude Opus 5 * Decide eligibility against the established agent projects (P1, P2) Two regressions this PR introduced, both reproduced on the branch and absent on db63e42b. **P1 — a name under a non-agent marker escaped a blocked scope.** Scoping resolved "which project" with `find_project_root`, the nearest project *marker*. That is the right boundary for grouping evidence and the wrong one for deciding whether a name is disqualified: a utilities package carrying its own `pyproject.toml` but no agent evidence is not a manifest scope. With the workspace's only application root unreadable and `agent_scope: single`, a bare `Agent(name="CrmHelper")` under such a package was attributed to `util/` — not the blocked `.` — so nothing rejected it, and plain `init --write` (no --allow-unresolved-scope, no test or template code) wrote `agent.name: CrmHelper`. Base kept CHANGE_ME. `_ProjectAttribution` now carries the projects `_agent_project_candidates` actually established, and ranking resolves a name to the nearest of those, falling back to the workspace. `scope_of()` raises rather than answering before the grouping has run, since answering early would silently attribute everything to the workspace. **P2 — the two implementations used different evidence sets.** The script counted every `Agent(name=…)` literal as project evidence; the CLI counts only framework-attributed ones. Feeding attribution into ranking turned that into an `agent_name_candidates` parity failure, and the underlying mismatch was already a *verdict* break: a phantom `agent_project_candidates` entry, and `agent_scope: "ambiguous"` where the CLI says `"single"`. The script now qualifies literals the same way, snapshotting each Python file's framework attribution right after the Python pass — exactly what `_score_python_signals` records on the CLI side. Four regression tests, including one that drives `init --write` end to end and two parity cases covering both the strong- and weak-marker shapes. All four fail on 6a9c019b and pass here. Publish an opt-in adopters registry, and the rule that every adoption number comes from it (#475) (#531) * Publish an opt-in adopters registry and the rule that every number comes from it (#475) Adoption evidence has been this project's #1 named risk in every strategic review, and the reason is structural: Shipgate collects nothing, so downloads measure CI caches and stars measure sentiment. Neither is adoption. This adds the counting mechanism the no-telemetry stance requires. `ADOPTERS.md` is opt-in and self-served: one row per adopter, added by the adopter through an issue form or a pull request, naming what they gate, the tier (local evaluation / advisory CI / blocking CI), the date, and a link to the public act where they asked. Private repositories are listed at organization granularity. Entries are removed on request, without a reason. The deliverable is the claims policy, not the rows. Eight rules: no entry, no claim; every number carries its as-of date; external and dogfooding are never summed; an entry is a dated statement rather than a measurement; a tier is self-reported and never upgraded; a badge is not an entry; the design-partner ledger and the historical corpus stay separate; removal is unconditional. Today's numbers are 0 external entries and 1 maintainer dogfooding entry. That row reads `advisory CI`, not `blocking CI`, because `main` requires no status check and a red run there stops no merge — the smallest available demonstration that the tier is what is true rather than what is flattering. `tests/test_adopters_registry.py` enforces it: row shape, closed tier vocabulary, ISO dates that are not in the future, entry links that point into this repository, this repository never listed as an external adopter, counts that equal the rows, an as-of date not older than the newest entry, one tier-neutral badge, an issue form that cannot drift from the registry's vocabulary or stop requiring consent, and any adopter number in the repository's prose that these rows cannot source. A 42-case perturbation sweep confirmed each guard fails on the weakening it exists to catch, that a correctly added adopter still passes, and that benchmark prose ("5 sentences of user instructions", "27 national company registries") does not trip the claim scanner. Co-Authored-By: Claude Opus 5 * Address review: keep the registry's guards usable by the first adopter Four findings from a review pass over the guards, all reproduced before fixing and all pinned by new perturbation cases (sweep now 45 cases). - The README external-count assertion was a bare substring requiring the plural "entries", so the first real adopter — count 1 — had to write "1 external adopter entries" or take a red build over grammar. Both counts now accept entry/entries, as the dogfooding assertion already did. - The self-entry agreement guard demanded the literal opener "That row says ``" for every dogfooding row, which cannot be written naturally once there are two. It now pins the claim ("row says ``") rather than one sentence shape; a second maintainer row with no explanation still fails. - Dates were compared against the runner's local `today`, so a contributor east of UTC dating an entry with their own today would fail CI in a UTC runner. One day of skew is allowed; "2027-03-01" still fails. - `_untraceable_claims` re-read and re-parsed ADOPTERS.md once per scanned file (294 today). The counts are read once and passed in. Co-Authored-By: Claude Opus 5 * Address second review round: verify the tier claim, and stop restating counts Three findings from a content pass over the prose, which the first round did not cover. - **The self-entry's tier claim was checked against the wrong mechanism.** Classic branch protection reports `main` as unprotected, but an active ruleset governs it. The ruleset requires a pull request and linear history and carries no `required_status_checks` rule, so `advisory CI` stands — but the file now says how the claim was made, with the command, so the next reader can disagree with it rather than trust it. - **ROADMAP.md restated the external count in prose.** Rules 1 and 2 say a number is published from the registry with its date; a second copy in a file nothing checks is how the first one goes stale. It now points at the file instead of repeating its number. - The changelog claimed a 42-case sweep; it is 45 since the last round's fixes. Two added README lines were also left unwrapped. Co-Authored-By: Claude Opus 5 * Address PR review: qualify claims by sentence, reject duplicate rows, accept any host Three P2 findings from review of eb71c0bd, each reproduced before and after. - **A neighbouring sentence could authorize an external claim.** The dogfooding exemption used a 120-character window, so "Agents Shipgate is used by 1 external team. Maintainer dogfooding is counted separately." passed the whole suite: the external `1` borrowed the maintainer total, which is the sum rule 3 exists to forbid. A qualifier now only qualifies its own sentence, and a claim that calls itself external never takes the exemption at all — prose that discusses both counts together is exactly where this had to hold. Both halves are pinned as regressions. - **Duplicate rows inflated the counts.** The counts were compared to list lengths, so the same adoption and the same consent record could be counted twice with every guard green. Rows now carry an identity — adopter plus repository link, compared across both tables — and a repeat fails. A `private` row is identified by its adopter alone, which is all the registry knows about it, so one organization holds at most one; the file says so, and says a second repository is a second row. - **Only GitHub repositories were accepted.** A correctly filled row linking `https://gitlab.com/acme/agent` failed, leaving a GitLab, Codeberg or self-hosted adopter to write `private` about a public repository in order to be counted — a registry mislabelling its own rows to fit its validator. The repository link is now validated independently of the consent record: any host for the repository, still this repository for the public act that created the row. Both rules have their own accept/reject unit cases. Sweep is 50 cases; the three reported reproductions are in it. Refuse an oversized detector candidate instead of reading it (#530) * Refuse an oversized detector candidate instead of reading it `tools/shipgate-detect.py` is fetched over `curl | python3` and run against repositories nobody has inspected, and it read every glob-matched MCP/OpenAPI/Conductor candidate — and every n8n and Conductor workflow it scored a framework off — with an unbounded `read_text()`. A multi-hundred-megabyte `*mcp*.json` was pulled into memory whole. The input adapters have always refused such a file before parsing it (`read_static_input_bytes`, `MAX_INPUT_FILE_BYTES`, 10 MB), so this was also a parity break: the file was *excluded* by `agents-shipgate detect --json` and *suggested* by the zero-install script, and an agent following the script would write a `tool_sources` entry the next `scan` rejects — the cold-start break the parse probe exists to prevent. Both sides now ask the same content-independent question from `stat` alone, and the script spells the refusal exactly as the CLI does, so `excluded_sources` agrees on the reason and not only on the split. Excluding the same file under a different explanation only moves the divergence: "too large" and "wrong shape" call for different work from the agent that reads it. The gate is asked *before* the script's `.yaml`/`.yml` early return. That early return exists because stdlib has no YAML parser and the script must never wrongly drop a spec it cannot judge — but size needs no parser, so an oversized YAML spec was the one case where "keep it" was wrong. CLI discovery had the mirror-image hole, and fixing only the script would have been a parity regression of its own: - `artifacts._looks_like_n8n_workflow` and the Conductor probe in `signals._collect_glob_hits` read their glob hits whole to score a framework whose adapter (`load_structured_file`) would then refuse the same file. Scoring a framework off a file `scan` cannot read names one nobody can verify. - `artifacts._probe_failure_reason` re-read an MCP candidate whole to sniff for `mcpServers` shape — including the file the adapter had *just* rejected for being too large, which would also have reported it as host configuration where the script reports it as oversized. Tests monkeypatch the bound down to 512 bytes rather than writing 10 MB fixtures; the rule under test is "compare the file's size against the configured bound". A third test pins `MAX_INPUT_FILE_BYTES`, `MAX_STRUCTURED_FILE_BYTES` and the script's copy to one value — that equality is what makes the split parity hold at all. `tests/test_adapter_static_only.py` allowlists the `git` subprocess call sites in `cli/discovery/artifacts.py` by line number; the inserts above them shifted both, so those entries are corrected here. Co-Authored-By: Claude Opus 5 * Address review: narrow the docstring claim, cover the manifest consumer Two issues found reviewing the size-bound change. **The module docstring overclaimed.** It said "Every structured candidate is size-gated at MAX_STRUCTURED_FILE_BYTES before it is read", but `_mcp_framework_evidence` reads `package.json` / `go.mod` unbounded and `_package_tokens` reads `pyproject.toml` / `requirements.txt` unbounded, both in the same file. A universal claim is the kind a later maintainer trusts instead of re-deriving, so it would have licensed the next unbounded read. It now names exactly what is gated — the suggestion candidates and the n8n/Conductor scoring reads — records that the MCP registration-site walk applies its own smaller `MAX_SOURCE_FILE_BYTES`, and states that the remaining reads are deliberately uncovered and symmetric with the CLI, so nobody bounds one side alone and reopens the divergence this change closed. **The n8n guard had an untested consumer.** `_looks_like_n8n_workflow` feeds framework scoring (pinned in `test_zero_install_detector`) *and* `discover_n8n_artifacts` -> `template.py`, which is what `init --write` writes as the manifest's `n8n.workflows` list. That second path carries the actual stated consequence — an entry `scan` refuses to read — and nothing exercised it, so a refactor moving the check into the scoring loop alone would have kept every test green while `init` regressed. Verified the guard is load-bearing there by reverting it. Writing that test surfaced two ways it could have passed for the wrong reason: the shared `_write_workflow` fixture is already ~1.8 KB, over the first bound tried (the explicit `st_size` assertions caught it), and a bare `"workflows/bulk.json" not in manifest_text` is too weak because both files also match the Conductor glob and are named in the rendered excluded-sources hint comment. The assertion pins the `n8n.workflows` entry form instead. The new import is aliased `artifacts as discovery_artifacts`: five tests in that file bind `artifacts` as a local, so an unaliased module-level import would be shadowed inside each of them. Separately verified, no code change needed: a boundary probe at the bound minus one, exactly at it, and one byte over shows both surfaces agree on the suggested/excluded split and on the byte-identical reason at all three points, and every remaining unbounded read in the script has an equally unbounded CLI counterpart, so no parity asymmetry is left. Measure first value and the next eligible change, and publish the zero (#521) (#523) * Measure first value and the next eligible change, and publish the zero (#521) The design-partner pilot counted three partners through one PR each and exited on a first-run feedback note. That cannot answer the next product question: after the first verdict, did a reviewer make a better decision, and did the team run it again on the next eligible change? `docs/design-partner-verifier-pilot.md` now carries the measurement protocol: two routes under test (host boundary with no manifest, agent tool surface with one), six denominators that failures stay inside, first value defined as four things a reviewer who did not write the change can name plus a recorded decision, a four-week window for the second eligible change, three separately-granted consents, a review-required ledger that reports the missing continuation honestly until #337 lands, and a pre-registered continue/narrow/stop rule. Ten minutes to first value is stated as an experiment target throughout. Dry-running the runbook's own commands produced the first result. The newest published build carries contract 10, and the runbook required "runtime contract 14" in the paragraph that installed it with `pipx install` — a precondition no published build has ever satisfied. On a host-boundary change that widens three permission rules to `Bash(*)` / `Read(**)` / `WebFetch(*)` and adds a remote MCP server, that installable build's `check` returns `warn` / `none` and names none of the widened rules, while this tree returns `block` / `critical` with all four named. The two entry failures are mutually exclusive by build: the published one writes a CI pin that resolves and an evaluator that cannot see the change; this tree has the evaluator and writes a pin to a tag that does not exist (#506). The documented `init` → `verify` order dead-ends on exactly the repositories the host-boundary route is for, because they have no tool surface to declare. `docs/design-partner-pilot-results.md` publishes six external denominators at zero, a dated enrollment shortfall naming that cause rather than a recruiting gap, the reproduced blockers routed to #506, #497, #520 and narrow: invite the baseline/drift pair that does run on the installable build, withhold `check` and `verify` until a published evaluator matches this tree's. `tests/test_design_partner_pilot.py` pins what is easy to soften by accident — a dropped denominator, dogfooding folded into the external count, the target restated as an achievement, a tracker field quietly removed — and fails the day a new tag ships, so the build-dated findings are re-measured instead of carried forward. A 19-case perturbation sweep confirms each guard bites. Co-Authored-By: Claude Opus 5 * Review: close three holes in the pilot guards, and one over-claim Self-review of the change, with a second perturbation sweep over the guards themselves. **The build-dated guard went quiet at exactly the release it exists to catch.** It asserted the newest published version appeared *anywhere* in the route-readiness section — and that section also names the in-tree build. So tagging v0.16.0 would have satisfied the check by coincidence, leaving the "newest published build carries contract 10" comparison standing untested and false. The ledger now states `**Published build measured: `0.15.0`.**` on its own line and the guard compares that value with the newest tag that exists. **The contract-floor guard was scoped to one range of numbers.** `contract \s+1[0-9]` catches the "contract 14" that was there; a future "requires runtime contract 21" would have passed. Widened to any digits. **`_section` terminated on a shell comment inside a fenced block.** Two `# ` lines in column one already live under § Evidence to preserve, so any later assertion on that section would have been silently truncated — vacuous rather than failing. The helper now tracks fences. **One claim was inferred rather than measured.** "No published build has ever carried contract 14" rested on assuming older builds are older still. What was actually measured: 0.15.0 carries 10, this tree carries 29, so 14 was first reached on the unpublished 0.16 line — and 0.8.0, the next-newest published build, has no `contract` command at all. Perturbation sweep now 23 cases: the original 19 still caught, plus a simulated v0.16.0 release (caught), a contract floor outside the old window (caught), a dropped measured-build line (caught), and a fenced shell comment in an asserted section (correctly tolerated). Co-Authored-By: Claude Opus 5 * Review round 2: bind the thresholds to the denominators, answer the question asked Three follow-ups from a second read against #521's acceptance and non-goals. **The decision rule could have been met by selecting successful users.** A non-goal is a numeric retention target reachable that way, and "≥ 2 reach first value" said nothing about the denominator it sits over. The rule now reads its counts against the reported denominators — two of three and two of thirty are different results and the decision must name which — and states outright that it is a threshold this project holds itself to, not one asked of a partner. **The standing decision answered a different question than the one asked.** the persona to invite next, which is not that. It now leads with the honest answer: none — nothing has been observed once, let alone twice. **"Keep the tracker private" named no place.** `.agents-private/` is already gitignored here for exactly this, so it says so. Co-Authored-By: Claude Opus 5 * Review round 3: the reproduction block did not reproduce The results ledger promises the next person can disagree with its numbers, and prints the fixture so they can. The `.mcp.json` "change" snippet showed the new server at the top level rather than inside `mcpServers`, so anyone following it literally would have built a repository with no second MCP server and then failed to reproduce the finding that depends on one. Corrected, and then checked the way the section asks to be checked: a fresh repository built by transcribing the block verbatim reproduces every published number — `warn` / `none` / 0 violations on 0.15.0, `block` / `critical` / 4 on this tree, and 4 expansion signals from the published build's baseline/drift pair. Co-Authored-By: Claude Opus 5 * Address review: reach the baseline from the change, and make the ladder pick one rung Both P2s from the Codex review of cd95f09f, reproduced before fixing. **The Route H setup could not produce a first result on a real PR.** (P2, docs/design-partner-verifier-pilot.md:284) A partner arrives with a branch cut before the baseline commit, so checking it out removes `.agents-shipgate/host-grants.json` and `--drift` exits 2. Reproduced on the ledger's own two-file fixture. The trap is sharper than a missing step: the CLI's recovery line says "Record one first: audit --host --save-baseline", and running that from the changed checkout acknowledges the very expansion under review, after which the drift reports nothing and the run looks clean. The command block and the partner prompt now carry both remedies — merge or rebase the reviewed baseline into the change branch, or keep the snapshot outside the repository and pass `--baseline-file` to the drift command — plus the prohibition and its reason. Both were verified on the fixture: each returns all four expansion signals. **The pre-registered rule licensed any answer.** (P2, docs/design-partner-verifier-pilot.md:545) The review's cohort — two repositories reach first value, one runs again, a third fails on a product defect — satisfied continue, narrow *and* stop simultaneously, and no precedence was given. Three independently-worded conditions are not a pre-registered rule. It is now an ordered ladder, first match wins, total by construction: 1 Continue (`first_value` ≥ 2 unaided and ≥ 1 repeat, same route), 2 Stop (`first_value` = 0 *and* `first_valid_result` ≥ 1 — a result nobody could act on), 3 Narrow (anything else). Rung 2 deliberately excludes the cohort that never reached a valid result: stopping because our own entry defects blocked entry would read a fact about this repository as a fact about the market. Six worked evaluations cover the review's counterexample, mixed success, single- vs multi-route, unaided vs translated, and all-failed-at-entry — which is the cohort the 2026-09-04 checkpoint is in, and the ledger now shows the ladder applied to reach it. Guards for both, and the sweep grew to 37 cases. One of the new guards was itself too weak on the first pass: it accepted a snapshot written with `--baseline-file` and never read back, because shell continuations put the flag and its command on different source lines. Co-Authored-By: Claude Opus 5 * Put the baseline caveat before the block it guards The reachability warning sat after the per-change commands, so a partner reading top-to-bottom would run the drift command, hit exit 2, and only then find the explanation — in a step whose whole subject is a partner's first ten minutes. Moved ahead of the block, with the out-of-tree drift invocation alongside the block it modifies. Co-Authored-By: Claude Opus 5 * Record the channel, and what the partner was told about it The standing decision now turns on which channel a partner is invited onto, and the tracker had no row for it — "published or preview build" was a guess at the distinction rather than the distinction #497's channel table draws. A row that does not say which channel it ran, and whether the partner was told the build carries no qualification, cannot be read against the decision later. Register the ten distribution surfaces and test that they agree (#497) (#525) * Register the ten distribution surfaces and test that they agree (#497) One engine is published through `action.yml`, `plugins/`, `skills/`, `adoption-kits/`, `harness/`, `examples/`, `prompts/`, `policies/`, `tools/` and the MCP server, and nothing checked that they said the same thing. #485 is what that costs: after #431 taught the CLI to read an MCP server's tool surface out of TypeScript or Go source, the zero-install detector went on answering `is_agent_project: false` for the vendor servers the CLI now accepts, and CI stayed green. `docs/distribution-surfaces.md` lists every surface, what it claims, and which test proves it. `tests/test_distribution_surface_parity.py` is that test, and `CONTRIBUTING.md` and `CLAUDE.md` point at both. A surface that answers nothing the engine answers still gets a row saying so. A new top-level directory fails the suite until somebody classifies it. Four surfaces were disagreeing with the engine, and are repaired here: - The bundled setup prompt told a coding agent to derive `agent.declared_purpose[]` from the README. `init` returns `control.next_action.actor: "human"` and `permissions.edit: false` for that field. The Codex kit's recipes page and the design-partner runbook carried the same instruction as a blanket "replace every CHANGE_ME". - The Claude Code kit rendered `@v0.16.0` and `agents-shipgate==0.16.0` into an adopter's CI — a tag and a release that do not exist. The drift was invisible because `test_claude_code_skill_source_matches_renderer` skipped that one file; removing the exemption is the actual repair. - The plugin's `plugin.json` and marketplace entry both demanded "runtime contract 15" beside `pipx install agents-shipgate`, which yields contract 10. The number was a second copy of what the bundled skill already states, so it is removed rather than re-synced. - The design-partner runbook named `v0.15.0`, demanded a contract that build has never implemented, floored pip at `>=0.13`, and gave a read order starting at `control.state`, which `shipgate.agent_handoff/v1` does not emit. `tests/fixtures/distribution_parity/` and a parity row that fails today and passes when the port lands, with `xfail(strict=True)` so the exemption itself fails once it is unnecessary. #506's two unpublished-pin gaps are recorded the same way, against a ledger of exactly the files that diverge. Co-Authored-By: Claude Opus 5 * Make the registry checkable in both directions (#497 review) Seven findings from the review pass, all real: - The `harness` registry row named `test_surface_states_only_engine_merge_verdicts`, which had been renamed out of existence during the first self-review. The registry's central promise — which test proves the claim — was already false in the change that introduces it. Claims are now claim-keyed to their proving tests, `test_every_claim_names_a_test_that_exists` resolves every name against this module, and `test_every_scanned_claim_actually_has_rows` catches the weaker failure: a claim naming a real test whose scan matches no file on that surface. That caught `design_partner_runbook`'s `executable_pin`, which carried only a `>=` install floor no pin pattern looked at — so floors are now checked too, for reachability rather than equality. - The tag corroboration read `CONTRACT_VERSION` with an unanchored regex. `CONTRACT_VERSION` is a suffix of `MINIMUM_CONTROL_CONTRACT_VERSION`, so it survived only because v0.15.0 has no minimum and line 155 precedes 156. - `ParityGap.surface` held one id for a gap spanning four surfaces, so the xfail reason pointed readers at the wrong place; `_GAP_ROW` parsed the doc's surface column and threw it away. - The doc/code claim comparison ran one way, so the registry could advertise a claim nothing tested. - `github_action`'s roots included `scripts/github_action_outputs.py` in code and not in the doc, and nothing compared roots at all. - The placeholder-ownership guard only ran on files containing the literal `CHANGE_ME`, so a prompt naming a human-owned field without that literal was unchecked. - `_surface_files` walked the working tree while the classifier read `git ls-files`, so an untracked scratch file could fail the suite. Two guards were replaced rather than fixed. The verdict scan matched nothing anywhere in the repository; `test_surface_enumerations_match_the_engine_vocabulary` replaces it, keyed to braced set literals (a comma-joined run is often a correct partial statement) with the two overlapping vocabularies disambiguated, plus `test_vocabulary_guard_is_not_vacuous`. The contract-floor table was hand-written and covered two of the four shipped copies; it is now scanned, spelling-agnostic so the plain-JSON "runtime contract 15" that shipped in plugin.json would be caught. Every guard was replayed against the pre-fix text from `main` and catches the defect it was written for. Co-Authored-By: Claude Opus 5 * Close three fail-open shapes found reviewing the review (#497) - `test_every_scanned_claim_actually_has_rows` matched a claim only when its proving tests were exactly one name, so adding a second test to a claim would have dropped it out of the coverage check silently. - `_release` parsed an install floor with a bare `int()`, so a pre-release in `release_status.latest_release` would have crashed inside the pin scanner rather than saying what was wrong. - `_tracked_files` shelled out to `git ls-files` on every call; collection asks dozens of times. Both new registry guards were mutation-tested: a claim naming a nonexistent test, and a claim whose scan matches no file on its surface, each fail. Co-Authored-By: Claude Opus 5 * Pin the ownership guard against going vacuous (#497) Widening the placeholder scan from CHANGE_ME-bearing files to every Markdown file on the surface fixed one hole and opened the possibility of another: the guard now runs over 48 files, and would read as green if a rewording left it with no human-owned field to look at. Seven mentions exist today; the floor is asserted, as it already is for the vocabulary guard. Co-Authored-By: Claude Opus 5 * Address the three review findings (#497) [P1] `release-verify.yml` checked out the candidate with the default shallow, tagless fetch and then ran the whole suite, which now requires the `v0.15.0` tag when `CI` is set. `release.yml` and `release-rehearsal.yml` both call it, so the release path would have gone red on a green PR — the fetch flags landed in `ci.yml` only. That checkout now asks for full history and tags, and `tests/test_action_metadata.py::test_jobs_running_the_whole_suite_check_out_history_and_tags` asserts the contract for every job that runs the suite, whichever workflow adds one next. It classifies a run by whether the pytest invocation has a positional path, so `--ignore=tests/x` does not read as a scoped run; both current suite jobs fail the assertion when their fetch flags are removed. [P2] The pin scanner never read the Action's own `shipgate_version:` input, although `action.yml` turns it into `pip install agents-shipgate==`. A workflow could therefore name a valid Action ref beside a package version that was never released, and the older public-surface guard checks that key against an explicit file list which cannot cover a workflow added later. Verified by adding the reviewer's tracked example: the registered parity test now fails on it and reports `shipgate_version input 9.9.9`. [P2] The vocabulary guard projected each documented set onto the expected values before comparing, so an unsupported member vanished — `needs_a_wizard` added to the setup prompt's otherwise-complete release-decision set still compared equal. Sets are now judged by all of their members, and a literal mixing the two vocabularies fails rather than being skipped by both, which was the same fail-open from the other side. The projection helper is deleted rather than kept. All three reproductions were replayed against the fix and now fail. Co-Authored-By: Claude Opus 5 * Point the registry at the real sources of truth after the merge (#497) The claims table still named `.well-known` for `executable_pin` and `schemas.contract.CONTRACT_VERSION` for `contract_floor`. Both are now `agents_shipgate.published_release`, which is where #506 put them and what the code imports — a registry that names the wrong source of truth is the same defect as a surface that restates one. The release-channel note also still described the rendered-prompt pin gap as open. It closed; the honest-statement rule that replaced it is what the note should point at. Read MCP registration sites in the zero-install detector (#485) (#526) * Read MCP registration sites in the zero-install detector (#485) `tools/shipgate-detect.py` is the documented first command run against a repository that has *not* adopted Shipgate — which is every repository #431 was about. #431 taught the installed CLI to read a tool's name out of a TypeScript or Go registration site; the script did not gain it, so the two disagreed on the one question the script exists to answer: `mongodb-js/mongodb-mcp-server`, `github/github-mcp-server` and `grafana/mcp-grafana` were agent projects to the CLI and "Stop, not an agent project" to the script. The masking lexer, the five idioms, the path predicate, the dependency gate, the export-precedence rule and the scoring are now in the script too, stdlib-only. On live checkouts of all three vendor servers both detectors now return `is_agent_project: true` with identical suggested sources, identical excluded sources and byte-identical evidence (61, 114 and 114 tools). Porting a load-bearing matcher means a second implementation of it, which is this repository's recurring bug class. What makes it affordable is that the two are not allowed to become *different* implementations: every case either reader has ever been asked about now lives once in `tests/mcp_idiom_corpus.py` — every idiom's positive sample, the whole adversarial sweep, the path predicate's cases and both escape grammars — and both readers are driven through all of it, compared site by site including each site's byte span. `samples/mcp_source_only_server` puts the route inside the existing `samples/` parity sweep, nine constructed workspaces pin the branches around it, and `test_framework_vocabulary_names_every_cli_omission` passes with an empty `known_omissions`. One defect surfaced while porting and is fixed in both readers: with no MCP export in the workspace at all, `_covering_export` returned every resolved name as "uncovered", and the caller renders a shortfall as "An MCP tool export is also present and does not name N of these registrations". A server whose surface exists only as source is the population this input was built for, so that claim about a file that does not exist was published into the adoption evidence for every one of them. Co-Authored-By: Claude Opus 5 * Address review: bound the export read, and one more route branch (#485) Three things from the review pass over #526. **The ported export reader had no size bound.** `load_mcp_tools`, which `_mcp_export_tool_names` mirrors, refuses any input over 10 MB before parsing it, so an oversized export names nothing on the CLI side while the port read the whole file and withheld the source route from it. `MAX_STRUCTURED_FILE_BYTES` is the same number and the script already applies it elsewhere. The remaining half of that divergence is `_probe_suggested`, which reads every glob-matched candidate unbounded and predates this change; it is filed separately rather than widened into this PR. **A route branch no fixture reached.** A single-file server registers at the workspace root, so the route is `"."` — its own branch in the unresolved-count rollup, and the one route path that reaches `agent_project_candidates` as a bare workspace. **The docs claimed less than the tests enforce.** `docs/zero-install.md` said evidence strings and scores are simplified; for `mcp_server_source` they are pinned, because the declared dependency is what carries the label to `medium` and the unenumerated-count line is a claim rather than prose. Co-Authored-By: Claude Opus 5 * Second review round: close four vacuous guards (#485) A perturbation sweep over the ported reader found four changes no test could see. Three were gaps in the shared corpus or its fixtures and are now closed; the fourth turned out to be a comment claiming more than the code does. - **The export size bound had no test.** Driven by lowering the bound rather than by writing a 10 MB fixture — the file being over it is the whole condition — and asserted against the outcome the CLI reaches by a different route: the loader refuses the file at the probe, so the source route stands. - **Three escape cases refused by a guard rather than by falling through.** `\u{zz}`, `\u{110000}` and Go's `\U00110000`: drop the check and `int(digits, 16)` raises out of `scan_source`, losing the whole file's surface instead of one name. The corpus had `\u{}`, which the emptiness check catches on its own. - **The wildcard fixture now carries the `tools: []` the CLI's own wildcard test uses.** And the honest note: `_mcp_export_tool_names`'s wildcard branch cannot be distinguished from its fall-through by any input, because every wildcard shape the probe accepts also has no usable `tools` array and the one that does is refused before this function is reached. That is written down where the branch is, rather than left as an untested-looking line. Co-Authored-By: Claude Opus 5 * Third review round: pin the readers on a CRLF checkout (#485) Git for Windows translates line endings on checkout by default, so the detector a maintainer curls onto their own repository is often reading `\r\n` source — and this repository has lost time to CRLF in a corpus reader before. The whole corpus now runs through both readers a second time with CRLF endings, asserting the CLI's own answer does not move and that the two readers still agree site for site. It passed on the first run, and one perturbation showed why that was luck rather than coverage: `_literal_is_whole_value` reaches the character *after* the literal only when the statement ends at a line break, and every case in the corpus ended in `;`, `,` or `}`. Deleting `\r` from its skip set changed no answer at all. So the corpus gains the shape that reaches it — a semicolon-less `static toolName = "…"` — and the perturbation now fails there. Co-Authored-By: Claude Opus 5 * Fourth review round: three route branches, and the reason a route vanished (#485) A perturbation over the discovery layer found three more changes no fixture could see. All three are claims a reader acts on, and none of them moves a verdict. - **The unresolved count is scoped to the route directory.** A registration outside the subtree the manifest will point at is one `scan` never reaches, so counting it is the mirror of the over-claim the count exists to prevent. Every fixture had its registrations inside the route. - **The evidence names the language the tools were *read* in, not the languages the workspace declared.** `grafana/mcp-grafana` is the shape: a `ui/` package declares an MCP dependency and every tool is in Go. Reading the gate instead would report a TypeScript surface this reader never found. - **The withheld route's reason is pinned, not just its path.** It is the only thing a reader gets when a route disappears, and a route that vanishes without one is indistinguishable from one nobody implemented. Also adds the both-languages fixture, which is what makes the second point testable in both directions. Co-Authored-By: Claude Opus 5 * Fifth review round: two masker fail-opens the shared corpus could not see (#485) A sweep over the masker itself. Both cases invent or lose a tool, and both are now corpus cases, so they pin the CLI's reader as much as the port's. - **A Go raw string is not escape-processed.** ``MustTool(`raw\137name`, …)`` registers that literal, which is not a tool-name shape and is recorded as an omission. Decode it the way an interpreted string is decoded and it becomes `raw_name` — a name the server does not serve, entered into the catalog as though it did. The corpus had a raw string, but with no backslash in it, so decoding was a no-op. - **A `/` inside a regex character class does not end the regex.** Stop tracking the class and the pattern ends early, leaving `']/;` as code — an apostrophe that opens a string and swallows the line, which is how one regex costs the registrations after it. One perturbation stayed invisible and should: dropping the `)` early exit in `_call_sites` routes a one-argument call through the "not a literal, so it needs a second argument" check, which drops it for a different stated reason and the same observable answer. No input distinguishes them. Co-Authored-By: Claude Opus 5 * Sixth review round: the word boundaries, and operationType's own guard (#485) The last sweep over the reader. Four more perturbations changed an answer with no test failing; all four are corpus cases now, so they pin the CLI's reader too. - **Three word-boundary lookbehinds.** `mcpgrafana.MustTool(` matches because a `.` is not a word character; a helper of the repository's own whose name merely *ends* in `MustTool`, `NewTool` or `Tool` is a different construct, and reading one invents a tool nobody serves. The existing `ToolDependencies{` sample does not reach the struct boundary — that name has text between `Tool` and the brace, so it never matched at all. - **`operationType` is safe only because its pattern requires `static`.** That is written in the module and was checked by nothing: `description` had the identical defect and a regression, and its sibling had neither. The regression fixture now carries a `this.operationType = "delete"` beside the `this.description` it already had, and asserts the field stays unset — a scratch assignment must not infer a `delete` risk tag. Two perturbations remain invisible and both provably: the wildcard branch in `_mcp_export_tool_names` (noted at the line) and the `)` early exit in `_call_sites`, whose removal routes a one-argument call through the second-argument check for a different stated reason and the same answer. Co-Authored-By: Claude Opus 5 * Document the whole walk on the source-only sample (#485) A route is only worth suggesting if the step after it can act on it, so the sample now records what `init` and `scan` actually do, not just what `detect` reports: the manifest gets `tool_sources: [{type: mcp_server_source, path: src}]`, and `report.json`'s `tool_catalog` carries both tools at `medium` confidence with the file and line each was registered at, with the runtime-named registration in `surface_exclusions`. Writing it down caught an over-claim in my own first draft. The terminal says `Surface: 0 tools` and stops at `insufficient_evidence` — the catalog has two tools and the reviewed *surface* has none, because an MCP server has no agent object to bind them to (`0/2 catalog tools reachable`). Telling an adopter to expect two tools in the terminal would have been exactly the kind of over-claim this input exists to prevent. `samples/mcp_only_server`, whose surface is a committed export, stops in the same place with the same `tool_sources[].binding` next step, so this is the shape of an MCP-server workspace rather than anything the new route introduces — and the README says so. Co-Authored-By: Claude Opus 5 * Fix four lexer defects the review found, in both readers (#485) All four reproduce in the reader #431 shipped as well as in the port, so this corrects the original rather than the copy. Two invent a tool name and two lose a whole file's surface. **A `${…}` holds code.** A brace inside a string, comment, regex or nested template is not a structural brace. ``const msg = `brace: ${"{"}`;`` left the substitution open, and from there the rest of the file was consumed as one unterminated template: every registration after that line silently gone, and a workspace declaring an MCP dependency reported as "not an agent project" over a brace in a string. Rewritten as an iterative scan with one stack entry per open template — a nested template is reached through a substitution, and recursion on attacker-shaped input is a crash rather than a wrong answer. **A line break does not always end the value.** `static toolName = "safe"` with `+ "_delete"` on the next line published `safe` at `medium` confidence for a tool the server registers as `safe_delete`. The check now looks past the break, and consults the continuation set only for characters the caller's own terminators do not claim — Go ends a struct field with a `,` on the next line. **The regex heuristic reads the mask, not the raw text.** A comment between `if` and its condition hid the keyword, so the slash was read as division and the pattern scanned as code — a tool invented out of a regex body, which is the one outcome masking exists to make impossible. `typeof /*c*/ /…/` had the same shape. **A backslash before CRLF is one line continuation.** Stepping over two characters left the `\n`, which ended the string: the identical file resolved its registration on a Unix checkout and lost it on a Git-for-Windows one. Nine expected-result cases in `tests/mcp_idiom_corpus.py`, so both readers are pinned to the corrected behaviour rather than to each other's agreement — and the CRLF sweep finally has a continuation case that exercises it. Reverting each fix fails its own case; four of the nine did not discriminate on the first draft, because a *closing* brace inside a comment or regex ends the substitution early and the reader lands on the same backtick anyway. They carry an opening brace now. Unchanged on the three vendor servers: 61, 114 and 114 tools, same routes, same evidence, from both detectors. Co-Authored-By: Claude Opus 5 * Add the grammar sweep the review's findings called for (#485) The four defects the maintainer found all live in the lexer, and six rounds of code perturbation had missed every one — for the reason those rounds could not reach. Mutating a branch only re-asks the questions the corpus already holds, so it measures the corpus rather than extending it. The exercise that finds a lexer defect is the other one: enumerate the *grammar* and ask what construct neither reader has ever been shown. Eighteen such constructs, run through both readers — a regex holding a backtick or a template opener, a comment holding a backtick or an apostrophe, a template holding a comment opener, an escaped backtick, a `$` that opens nothing, a regex beginning a statement after a block, division after an index, a private static field, and on the Go side a rune holding a backslash or an escaped quote, a raw string holding a quote or a comment opener, and an interpreted string holding a backtick. All eighteen passed on the first run, so this adds no fix. They are in the corpus because the next change to the masker has to keep them passing, and because the list is the record of what has actually been asked — which is the only thing that distinguishes a construct this reader handles from one nobody has tried. Pin what init writes to a release that exists, and say what it reports (#506) (#524) * Pin what init writes to a release that exists, and say what it reports (#506) `init --write --ci` wrote `uses: ThreeMoonsLab/agents-shipgate@v0.16.0` into an adopter's repository, and no such tag had ever been cut. GitHub fails that job at action-resolution time, before any step executes, so a first-time adopter's very first Shipgate run was a red check carrying an error about our repository rather than theirs. The bundled onboarding prompt had the same defect one layer up: `uvx agents-shipgate@0.16.0`, which the index answers with a 404. Every render since 2026-07-09 named a nonexistent ref -- 56 days. Two conventions coexisted and only one was correct. `llms.txt`, `.well-known`, the docs and the Action examples tracked the latest published tag; the one artifact that gets *executed* by a stranger's CI tracked `__version__`, which for the whole interval between releases is a version nothing can fetch. `LATEST_PUBLISHED_VERSION` moves out of the test file into `src/agents_shipgate/published_release.py`, and every surface -- documentation and emitted artifact alike -- now derives from that one constant. Pinning the published release keeps the pin resolvable but does not make it sufficient, and conflating those is what the previous fix got wrong in the other direction: the bundled prompts demand runtime contract 21, and v0.15.0 emits contract 10. So the prompts state that gap where they state the pin, rendered from `LATEST_PUBLISHED_CONTRACT_VERSION` -- read back out of the tag itself by the suite rather than asserted about it. When the newest published build predates the floor, the honest output is to say so; it is never to pin a build that cannot be fetched. `tests/test_init_ci.py` asserted the defect, which is why it shipped. It now requires a published tag, and `tests/test_adopter_pins_resolve.py` sweeps every pin shape across everything `init` emits, refuses to answer from an empty tag list, and carries a negative control that re-introduces the exact defect. Cutting v0.16.0 is not what fixes this: the rule is "names a tag that exists", so it holds on the first commit after the tag too. Co-Authored-By: Claude Opus 5 * Address review: sweep every target, and stop the guards being spelling-deep (#506) Five findings from the review pass on this branch. The pin sweep claimed to cover "everything init writes into an adopter's repository" and covered two of eight targets. `AGENTS.md`, `CLAUDE.md`, the slash command, the Cursor rule, the PR template and the local contract carry no pin *today*, which is exactly why a sweep that skipped them would look correct indefinitely and fail only on the day one of them grew an install snippet. It now drives off `SPECS` -- the registry `--agent-instructions` itself selects from -- through `render_targets`, the function `init` renders with, so a target added there is swept the day it is registered. A second negative control injects a pin into `AGENTS.md` and proves that coverage is real rather than nominal; without the widening it goes uncaught. The runner-pin pattern required a literal `uvx ` prefix, narrower than the repo's own `UVX_PIN_PATTERN`, so `uv tool run agents-shipgate@0.16.0` would have carried an unfetchable pin past it while the non-vacuity check stayed green on the other prompt's `uvx` occurrence. It is prefix-free now, for the reason that pattern already documents: a digit must follow `@`, so the `@v` Action form cannot collide. The "must not promise the pinned release is new enough" guard matched two exact spellings, and neither was the one `contract_floor_prose` renders -- a hand-edit to ``(`0.15.0` or newer; ...)`` restored the precise false claim the guard exists to forbid. It is a pattern now, and it names the numbers in its failure message. `published_release.py` pointed a maintainer at a runbook section that does not exist; the instruction is step 8 of `## Cutting the release`. And the two git checks disagreed about a missing git binary -- one skipped, the sibling raised `FileNotFoundError`, reporting an environment limitation as a code failure. One helper states the precondition once. What still does not skip: a working git that answers with no tags. Co-Authored-By: Claude Opus 5 * Second review round: name the file an adopter sees, and resolve before joining (#506) Three from re-reading the widened sweep. `render_targets` resolves the workspace before joining and reports absolute paths against *its* view, so labelling the emitted files needed the resolved workspace. `tmp_path` is already resolved, which is why the suite passed on the unresolved form -- and why it would have raised `ValueError` for any caller whose temporary directory goes through a symlink, which is every `tempfile.TemporaryDirectory()` on macOS. Found by running the helper outside pytest, which is the only place the difference shows. The labels are workspace-relative now. A violation message naming `/private/var/folders/.../adopter/.claude/skills/...` tells the reader nothing about which shipped file to fix. `AGENTS_SHIPGATE_WORKFLOW_REF` is a documented escape hatch, and exported in a shell it turned every sweep in the module red for a reason unrelated to the code under test. An autouse fixture clears it, so the sweep measures what `init` writes by default; the one test that wants the override still sets it, after the fixture runs. Verified with the variable exported. And the allowlist comment now names its guard test exactly rather than by a prefix of its name. Co-Authored-By: Claude Opus 5 * Bind the drought count to when it was measured, and state the final sweep (#506) The changelog claimed "56 days" in the present tense for a window that keeps growing until this lands, and described the sweep as it stood before the review widened it to every registered `--agent-instructions` target. Land the five rulings and force the insufficient_evidence line (#508) (#519) * Land the five rulings, both calibration re-runs, and the three cases that force the evidence-gap line (#508) **PR #514 merged an earlier state of its branch**, so main has the packet integrity work and the codex read-boundary audit but not the owner's rulings, not the round record, and not the round-2 re-run. This carries all of it onto main and adds round 3. Nothing under `src/` changes; #513's files are left as main has them. 1. **Judge the diff, and the gate's deterministic deliverable is the capability delta, not risk.** A `review_required` must hand the person a *named* capability or it does not get to block. `LABELING.md`: "What you are deciding" rewritten around that; `review_required` redefined as "a capability you can name"; the incomplete-guard bullet removed and its case moved to `passed`. Engine: #515, #518. 2. **Agent instruction files are out of scope** — semantic, and a static gate cannot judge them. Labeled by what *else* the diff does, else `passed`. Engine: #516. 3. **`insufficient_evidence` only when it names what would resolve it.** Engine: #517. 4. **Identical copies, mechanically.** `identical_files` in the manifest by sha256; `run_rater` rewrites a citation of any copy to the canonical path and records what was cited. 5. **The codex read boundary as context management.** A shell-bearing rater is refused on a host carrying `strata-inventory.csv`; `--working-material` proceeds for calibration and the label records which it was. **κ = 0.4444 → 1.0000.** Both splits resolved, to the side the rulings predict, and the rater that moved cited the rule it moved on. Twenty labels across two rounds chose `insufficient_evidence` zero times, for a rule governing 15 of 60 slots; refinement 1 was equally untouched. Three constructed cases on one fleet-ops base: `cal-6` (nothing nameable survives), `cal-7` (`cal-6` plus a gate made non-blocking), `cal-8` (`cal-6` plus a literal tool that bills an account). Both families agree on all three: `insufficient_evidence`, `blocked`, `blocked`. They were run for what the raters *write*. On `cal-6` both named the same two artifacts — the capability profile and the OpenAPI spec it references — with claude naming the two keys inside the first, and noticing unprompted that the manifest and reviewed inventory still describe the deleted tools. On `cal-7` and `cal-8` neither rater let the opaque remainder swallow the visible finding. A fixture defect was found mid-round by a rater and fixed: the three cases first shipped `cal-5`'s manifest and inventory, which describe a different agent. Decisions unchanged after the fix; the resolving sentence got more specific. Refs #508, #515, #516, #517, #518. Co-Authored-By: Claude Opus 5 * Address PR #519 review: staging erased link changes, equal bytes are not identity, and the recorded model was the announced one All three reproduce; each fix has a test that fails on the previous commit. - **[P1] Staging decided what git recorded, and it was dropping dangling links.** `copy_tree_excluding` stages both `base/` and `head/` into the throwaway repository, so whatever it did to a link *was* the state git saw. Reproduced: a link only `base/` carried, deleted by the change, gave identical tree hashes for both pins, a 0-byte `diff.patch`, and no `broken_symlinks` — a packet asserting nothing changed, for a change whose whole content was that deletion. git stores a link as its target bytes without following it, so staging now recreates links as links (`preserve_symlinks=True`) and the packet boundary — where a rater's world actually begins — decides what survives. The broken-link list is computed once, from the tree that is about to become `repo/`, for both kinds of case. Covered for base-only deletion and for a changed target. - **[P1] Equal bytes are not semantic identity.** Canonicalisation rewrote `IndependentHumanLabelV1.evidence_references`, which is what everything downstream reads, so a citation of `repo/runtime/hooks/pre.sh` would be rewritten to `repo/a.txt` whenever those happened to hold the same bytes — destroying the thing the citation establishes, since a path says which hook, provider or package is loaded. The rater's citations are authoritative again; the folded form is `evidence_comparison_key`, derived and beside them, for the one job it is good for. Grouping now keys on content **and** the executable bit, so `pre.sh` never joins a group with `a.txt`. - **[P2] The recorded model was the configured one, not the serving one.** Every Claude label recorded `claude-opus-5[1m]` from the `init` event — decoration the API never returns — while the assistant message events in the same transcripts consistently name `claude-opus-5`. `claude_final` now takes the model from the message events, refuses a transcript naming more than one serving model, and keeps the announced value as `announced_model` with a diagnostic when they differ. The thirteen archived working label records were re-derived from their own transcripts and carry `model_corrected_from`; the round record is corrected and records this as finding 6. Co-Authored-By: Claude Opus 5 * Widen the answer-key check past the one file it knew about Answering "why a host without the checkout?" showed the constraint was both overstated and under-enforced. **Under-enforced:** the guard knew only `strata-inventory.csv`. The checkout carries ten answer-stating files — the inventory in both formats, six adjudicated `benchmark/miner/results/*.labels.csv` (`pr_url,label` for real PRs, several of them later pinned as corpus candidates), and the calibration records, which now state every `cal-*` decision. `answer_keys_on_host()` looks for all of them. **Overstated:** "a host without the checkout" cannot work at all, because the harness imports from `src/`. What must be absent is the answer-stating files, and a trimmed deployment — `src/agents_shipgate/`, the two `rater/` scripts, the packets — already is that. Recorded in `cut-c-preconditions.md`, along with why `src/` and `docs/checks.md` are deliberately not on the list: condition 2 forbids a rater seeing verifier *output*, and source is not that. The test suite models a clean host by moving `REPO_ROOT` to an empty directory rather than stubbing the lookup, so the tests still exercise the real `answer_keys_on_host`; the refusal tests put files back under it. Co-Authored-By: Claude Opus 5 * Make the answer-free host a script instead of a thing to remember The owner chose local rounds on OAuth over a one-time GitHub Action with API keys: the Action would have been stronger isolation — a fresh runner, and `--home-mode isolated`, which OAuth cannot use — but it costs a key and metered spend for sessions that run free under the existing subscriptions. Two consequences, both written down rather than left implicit. **The rounds run in `shared` mode**, so blindness is *checked* rather than *structurally impossible*. `cut-c-preconditions.md` now states exactly what that mode does give (packet root and every ancestor checked for instruction files, no auto-memory for the packet path, `--setting-sources ""`, codex loading none of `config.toml`) and what `isolated` would have added (an empty `HOME` with `--bare`, a Codex home built from nothing) — in the terms the Amendment 1 disclosure block has to use. **`deploy.py` builds the host.** It copies `src/agents_shipgate/` and the `rater/` scripts into a layout whose root *is* the deployment, which is what makes `answer_keys_on_host()` search the host the rater actually runs on rather than somewhere harmless — the layout is the check, and a test pins that relationship. It then runs the check against what it just built and refuses if anything turns up. Packets are built elsewhere on purpose: choosing which to build needs the inventory, and that is the answer. Verified end to end: a corpus-mode session — no `--working-material` — ran from a deployment and recorded `host_isolation: no answer key on host`. Co-Authored-By: Claude Opus 5 * Run the first corpus round: 96 labels, and two blockers that stop the tag (#508) 96 admissible blind labels over 48 cases, both families, the corrected guide byte-identical in every packet, from a `deploy.py` host carrying no answer-stating file — every label records `host_isolation: no answer key on host`, and none used `--working-material`. Sharded by case so `claim_family` could compare. Labels and transcripts stay on the owner's machine. It did not reach the bar, in two independent ways. **Only 48 of the 60 slots can be a packet.** Twelve are shipped samples under `samples/`, which are a single tree: no `base/`, no `head/`, so not a change, and the packet is defined as head plus the diff that produced it. They are cold-start cases — the gate runs `scan` on a state — and making them ratable needs a packet form with no `diff.patch` and a rubric that asks about a repository rather than a change. Both change what the corpus measures. 48 < 56, and the per-decision floors cannot be met either. **κ = 0.6111 against a floor of 0.80**, and adjudication cannot repair it: κ is a property of the two blind primaries. The disagreement is one ambiguity, and this branch created it. Six of fourteen splits are purely `review_required` ↔ `insufficient_evidence`; collapsing that line gives κ 0.7322. Codex reached for `insufficient_evidence` 13 times to claude's 5, and the rationales agree on the facts and divide on one question: a tool the diff **registers by name**, whose endpoint and credential are citable, but whose advertised operations live outside the packet — a capability you can name, or a surface that cannot be established? Both raters are following the guide, because the guide says both. Ruling 1 rewrote `review_required` as "adds a capability you can name" and the `insufficient_evidence` list was not revisited — it still reads "an integration is **mounted by name** and its capabilities live somewhere the repository does not include". Those overlap on the commonest shape in real history. It is the same structural defect that produced ruling 1, introduced when the rulings landed rather than found by them. No threshold is moved and nothing is adjudicated toward the inventory's `target_decision`, which was chosen with the engine's verdict in view and is a sourcing guess, not the answer. Co-Authored-By: Claude Opus 5 * Guide: a binding is a capability you can name; insufficient_evidence is a signal, not a step (#520) The owner's sixth ruling, from first principles: `insufficient_evidence` should be rare and avoided where possible. The first corpus round measured why (κ 0.6111): two rater families split consistently on a tool registered by name whose advertised operations live outside the packet, because ruling 1 had rewritten `review_required` as "a capability you can name" while the `insufficient_evidence` list still said "mounted by name … lives elsewhere". The guide now says what the four agreed-IE cases already showed: "I cannot enumerate the operations" and "I cannot establish the authority" are different claims, and only the first is true of a remote or runtime binding. **The authority is the binding.** A new section, *Naming a binding*, says how to cite it and how to judge it — unbounded and unguarded → `blocked`, bounded and attributable → `review_required`, narrowed or untouched → `passed` — and the three shapes that used to be filed under `insufficient_evidence` move there. `insufficient_evidence` is no longer a step in the decision procedure. It is what a rater reaches for when the packet is incomplete, and packets are checked complete before a session starts; reaching for it is a signal that the guide has a gap, and the rationale must say what was unnameable so the gap can be fixed. The illustration that used to end in it now ends in `review_required`, and a second one — one remote endpoint, one credential — says why. `tests/test_labeling_guide_is_rater_safe.py` passes. Corpus and requirements changes (28 → 21 cells, the twelve cold-start slots retired) follow the re-label, not precede it. Co-Authored-By: Claude Fable 5.1 * Corpus round 2 on the corrected guide: kappa 0.8048 on 47 pairs, one case without a label (#508, #520) Same 48 cases, families, models, host layout and sharding as round 1; the one change is `LABELING.md` carrying the sixth ruling. Packets rebuilt so the guide inside each is that text. **κ = 0.8048 on 47 complete pairs, against a floor of 0.80.** Round 1 on the same 47 was 0.6036. `insufficient_evidence` went from 18 labels to 1; the four cases both raters had filed there in round 1 are now `review_required` or `blocked` from both, naming the binding. Disagreements fell from fourteen to six, and they are no longer one axis: four sit on slots the inventory sourced as `insufficient_evidence`, where the raters now split on whether the binding is worth a human or nothing changed — a judgement call for adjudication, not a guide contradiction. One case has no `security_governance` label after three independent attempts at three fresh packet paths, refused each time for omitting `evidence_references`. Not random (three of three on this case), not the ruling (each attempt said `passed`, on an instruction-files-only change where the rater seems to conclude there is nothing to cite). The parser was not loosened and `TASK.md` was not forked for one case; the 48th pair is open and could move κ to either side of 0.80. Recorded as it stands. Still open: the inventory and requirements changes the ruling implies (28 → 21 cells, the twelve cold-start slots retired, the IE-sourced slots re-targeted), which follow this record as their own PR. Co-Authored-By: Claude Fable 5.1 * Record the six adjudications, and the third blocker they expose (#508, #520) The owner adjudicated all six disagreements as third identity, disclosing a personal walk on every one. Each upheld one of the two primaries, so each frozen record carries that rater's citations — siding with a rater is adopting their evidence, and nothing is invented to satisfy `HumanAdjudicationV1.evidence_references`. **Not one adjudication landed on `insufficient_evidence`**, including the four on slots the inventory had sourced as exactly that. Forty-seven cases now carry a final decision (41 agreement, 6 adjudicated) and none is `insufficient_evidence`; the sixth ruling holds end to end. **The strata no longer balance, and that is the third blocker.** Removing the label redistributed the cases: 14 of 21 cells reach two, seven do not. `review_required` runs to 21 while `blocked` falls to 12, and three profiles hold a single `blocked` case each. Totals are fine (47 against 42 needed); the shape is not. This is a sourcing result that could not appear before the labels existed — every slot was sourced against a target decision, and the blind raters put the cases somewhere else. The 28 → 21 requirements change is deliberately **not** committed with this: writing two-per-cell into `pre_release_safety_requirements()` today would encode a target the corpus is known not to meet. The owner settles the shape first — source into the seven short cells, lower the per-cell count where the material does not exist, or narrow the profile set. Adjudication records are with the labels on the owner's machine, uncommitted. Co-Authored-By: Claude Opus 5 * Retire `insufficient_evidence` as a target, and shorten four cells to the material that exists (#520, #508) The first blind labelling round put both of these under load and answered them with measurements rather than argument. `insufficient_evidence` is what the gate says when its own extraction failed. That is a statement about shipgate, not about the change, so it cannot be the answer to "what should a correct gate do here?" Both named policies stop demanding cases that expect it: `beta` 28 strata / 100 cases -> 21 / 80, `pre_1_0` 28 / 56 -> 21 / 38. The value stays in the enum and the verifier still emits it; it is scored as a miss against whatever the case expected. Four `blocked` cells now ask for one case, not two. Of the seven profiles only four produce a real-world `blocked` case at all, and each produces exactly one -- Cut B recorded the cause and the W36 sweep confirmed it, since a change that should have been stopped usually was. The second case is reachable only by building another construction, and a cell filled with constructions measures our imagination rather than the world. `strata-inventory.md` carries the per-cell evidence. No rate moved. Every exact-match floor is still production's rate over the population it governs, rounded up; `minimum_blocked_exact` falls to 10 because there are 10 `blocked` cases at the same 100% demand. The origin floor stays 40% of the corpus (32 / 16), which is why holding it at a fixed count was refused -- that would have raised production's demand to half the corpus as a side effect of deleting a decision. The stdlib restatement in `scripts/_release_support.py` splits the four decisions a case may *carry* from the three it may be *targeted* at, so a zero-count cell is absent rather than present-and-empty; the sealer compares strata by equality and a present-and-empty cell would reject every conforming artifact. Co-Authored-By: Claude Opus 5 * Address PR #519 review: a guide seam that split two raters, and three guards that were narrower than their own reasons (#508, #520) **The rubric could be read two ways, and it was.** `passed` carved out "read-only additions whose reach is fully visible and plainly within the agent's stated purpose"; the `review_required` headline and procedure step 2 said "adds, widens, or unguards a capability you can name" with no exception. A new read-only tool is both. It bit a real case: on one adjudicated pair the two raters split on that seam, and the `framework_tooling` rationale reads almost verbatim from the carve-out. The owner's adjudication settles the direction, and not the obvious way -- it upheld `review_required` on a rationale that turns on *reach*, not on "a new tool": the org allowlist only enforces when an org argument is present, that call takes none, so the bound constraining every other read does not reach it. So the exception is real and its third condition was missing. The guide now states the **bounded-read exception** in three establishable conditions -- read-only, fully visible reach, and no reach the agent did not already have -- in `passed`, referenced from both other sites so all three agree, with a constructed pair showing where it stops. No existing final moves: every `passed` resting on the carve-out was re-read against condition 3, and the one that adds a new read-only tool reaches a fixed path template under the scope that already bounded the agent. **The command audit asked how a path was spelled, not where it lands.** `packet.parent` is a prefix of every absolute path *inside* the packet, so `cat /repo/agent.py` -- a session reading only its assigned input -- produced no label at all, and the deployment layout makes `REPO_ROOT` an ancestor of the packets too. Each token is now resolved against the session's working directory (which is the packet) and judged by containment. Resolving also fixes what the substring test got wrong in the other direction: `repo/../repo/x` never left, and two spellings of a symlinked directory now compare equal. **The answer-key check knew four filenames.** It missed the corpus round records -- one publishes case id -> both primary labels -> final decision -- and it missed them *alone*: with only that file present the check passed. Two classes it never knew about at all: each construction's `CASE.md` names its target cell and slot id, and the miner's `*-mined.csv`/`.jsonl` sweeps carry `head_decision`/`verify_decision` per `pr_url`, which is verifier output for the very PRs the inventory pinned. Matched by directory now, because the next round record is written by someone who will not think to come back here. 8 files found before, 52 now. Each class is asserted to refuse on its own. **`--working-material` was not recorded.** Only `host_isolation` was, and it cannot carry the flag: for claude it is always "no shell", and on a clean deployment "no answer key on host" either way -- so two runs differing only in this wrote identical records and a freeze could not tell which labels the caller had excluded from evidence. Prioritize first value, completed review and repeat use in the roadmap (#522) Clear what Cut C's preconditions can clear, and stop the change under judgement subtracting from its own packet (#508) (#514) * fix(benchmark): make a rater packet something the change under judgement cannot subtract from (#508) Cut C's four preconditions, cleared as far as this machine can clear them, and one class of defect found while clearing the fourth. **Precondition 4 — the packet contents.** Decided: head tree plus diff, no base tree. The argument for that was already written down, and it quietly assumed `diff.patch` is a complete textual description of base -> head and that `repo/` is the commit's tree. Neither was true. Both of git's readers obey `.gitattributes` *from the tree under judgement* -- that is, from the change being labeled: - `export-ignore` made `git archive` drop paths, so `repo/` was missing files and `MANIFEST.json` hashed only what survived; - `-diff` or `binary` made `git diff` reduce an ordinary text change to `Binary files ... differ`. The second is the one this corpus cannot survive: most of what the rubric calls `blocked` is a *removal* -- an allowlist that no longer bounds, an approval step deleted -- and a removal is visible only in the diff. A repository could hide exactly that, in a plain text file, and the packet would still build and still verify against its own manifest. Shipping the base tree would not have fixed it; `export-ignore` subtracts from a base tree too. So the tree is now read with `ls-tree` + `cat-file`, which consult no attribute, and the diff with `--text --no-textconv`. What survives that -- genuinely binary content, and submodules, whose content is in neither tree -- refuses the build and names the paths. A fourth consequence of the same cause also goes away: `git diff` read attributes from the clone's *worktree*, so the same two pins produced different `diff.patch` bytes depending on which commit the clone happened to have checked out. `git archive` is gone, and with it tarfile's `data` filter, so the `../` refusal it used to provide is restored explicitly. **Precondition 3 — the OpenAI harness.** The recorded cause was an empty vendor directory. That is a symptom: the cached `@openai/codex@0.85.0` `darwin/arm64` binary extracts cleanly and passes `codesign -v`, but `spctl` reports `CSSMERR_TP_CERT_REVOKED`, so macOS kills it at exec with no output. Still owner-gated -- it needs a current install and a credential -- but `run_rater.py` now probes the CLI before handing over a packet (`--check-cli`), so a missing binary, an empty vendor directory and a binary macOS refuses to load say which remedy each needs instead of failing obscurely mid-run. The version it reports is recorded with the label. **Precondition 2 — the Claude harness.** Verified against `claude 2.1.126` short of one live session: every flag exists with the spelling used, the restrictions take effect (`tools` is exactly Glob/Grep/Read, no MCP servers, `dontAsk`, no user or project settings), `_encoded_project_dir` computes the directory name the CLI actually uses, and the 401 arrives as `subtype: "success"` with `is_error: true` -- a shape a `subtype`-only check would pass, and which `claude_final` refuses because it gates on both. **Precondition 1** is the owner's sign-off and is not touched. Nor is the `review_required` / `insufficient_evidence` rule: #508 orders the calibration round to test the draft, and a correction written without the round is the guess 56 labels would then be produced against. **Also found:** Amendment 1 condition 1 -- the two roles on different model families -- was enforced by nothing. `SafetyCorpusCaseV1` asks only that the two `reviewer_id` values differ, and two sessions of one family differ anyway, in the session uuid. The runner now refuses the second role when the sibling label records the same family, before that session starts. Nothing under `src/` changes. Co-Authored-By: Claude Opus 5 * Address review: pin the diff against git config, and record what the family check could see Five findings from the review pass, all real. - **`diff.patch` still varied with the operator's `~/.gitconfig`.** Attributes were pinned; configuration was not. Reproduced: `diff.context=7` turns a 13-line patch into a 21-line one, `diff.noprefix` rewrites every header, `core.abbrev` widens every `index` line. That put `MANIFEST.json` -- and every `diff.patch:` a rater cites for an adjudicator to re-read -- at the mercy of whose machine built the packet: the same defect class the change set out to fix, left open on the neighbouring input. Every diff-shaping key is now pinned with `-c`, plus `--full-index`; `diff.renames` is pinned *off*, because rename detection summarises a move instead of showing both sides and what a rater must see is content that left. The new test fails on the previous commit. (`diff.orderFile=` is not the neutral value -- git then tries to open a file called `""` -- which the existing suite caught immediately. `/dev/null` is.) - **The family-independence check could be bypassed by two `--out` directories.** It compares against the sibling role's label record, which it can only find under the same `--out`; splitting them left condition 1 unenforced and silent. Silence was the wrong answer: each label now records `family_independence`, reading `"unchecked"` when there was nothing to compare with. The first role of a case is legitimately unchecked; a case whose *both* records say so is one nobody ever compared, which a freeze step can see. `calibration.md` says to use one ``. - **`cli_version` recorded the probe's build, not the session's.** The Claude `init` event carries `claude_code_version` for the run that actually produced the label, and `claude_final` was already parsing that event for the model. It now takes both, and the probe's answer stays the fallback for the family whose stream names neither. - **`suppressed_diff_markers` had no known trigger and no test.** With `--text` nothing is known to reach it (checked: a low `core.bigFileThreshold` still renders NUL-laden content textually). It stays, because `--text` is one edit away from being dropped and the failure it guards is silent everywhere else -- but it now says it is a backstop, and its behaviour is pinned at its own boundary rather than left unexercised. - Test hygiene: a conditional expression used as a statement, and two dead assignments. Co-Authored-By: Claude Opus 5 * Address review round 2: a sibling that names no family is a refusal, not a pass Two findings, both in what the last commit added. - **`check_family_independence` claimed a check it could not perform.** It compared `sibling.get("family")` against the current family, so a sibling record carrying no `family` key never matched and the run recorded `"checked against framework_tooling (None)"`. Executed against the branch to confirm. That is worse than the silence `family_independence` was added to replace: a freeze step looking for two `"unchecked"` records would see a positive claim instead. A sibling that names no recognised family is now a refusal, before the session starts. - **`probe_cli`'s docstring described the superseded behaviour**, still saying the version it prints is what the label records. It is the fallback; a family whose stream names its own build supplies it instead. In a module whose docstrings are the protocol's specification, that is a wrong specification sitting beside correct code. Co-Authored-By: Claude Opus 5 * Address review round 3: read the diff through a git dir that carries no attributes One finding, and it was the claim the last two commits made. `--text` and `--no-textconv` answer the attributes that *hide* content. They do not answer `diff=`, whose funcname pattern chooses the text printed after every `@@` -- and git's built-in drivers need no configuration, so the tree under judgement selects one on its own. Reproduced: `*.md diff=markdown` turns `@@ -7,4 +7,4 @@ line b` into `@@ -7,4 +7,4 @@ # Title`, and which one you get depends on whether that `.gitattributes` is in the clone's worktree -- the same checkout-dependence this branch set out to remove. Nothing is hidden and no cited line moves, so it is cosmetic for a rater; it is not cosmetic for an artifact whose identity is its hash, and "a function of the two pins alone" was written in three places while it was not yet true. The diff is now read through a throwaway **bare** git directory: no worktree to read a `.gitattributes` from, `objects/info/alternates` pointing at the clone's real object store so it reads the same commits without copying them, and one line in `info/attributes` -- `* !diff` -- which outranks every `.gitattributes` in every tree. Nothing is written into the operator's clone. The flags stay, restating for the two attributes that hide content outright. The new test fails on the previous commit, and its companion checks the neutral git dir reads the commits it was built from rather than passing by reading nothing. Co-Authored-By: Claude Opus 5 * Address review round 4: pin core.quotePath, and test the config channel that is still live `_DIFF_CONFIG` pinned six keys and missed `core.quotePath`, which decides whether a non-ASCII path reaches every header as itself or as `"caf\303\251.py"`. `core.quotepath=false` is a common global setting, so a case touching an i18n fixture still hashed differently on two machines. Writing its test found that the two config tests were pointed at the wrong channel. They set the *clone's local* config, and since round 3 the diff is read through a bare git dir that never sees it -- so the quotePath test passed against the very code it was written to fail. Both now set `GIT_CONFIG_GLOBAL`, which is what an operator's `~/.gitconfig` actually is, and which does reach the neutral git dir. Checked against two builds: on the branch's first commit (no `_DIFF_CONFIG`) both fail; on the previous commit (`_DIFF_CONFIG` without `core.quotePath`) the quotePath one fails and the other passes. Each names the defect it is for. Co-Authored-By: Claude Opus 5 * Address review round 5: pin the last two diff knobs, and make the test carry the exhaustiveness claim `_DIFF_CONFIG`'s comment said it pinned "every key git reads from configuration that changes the bytes of a diff". It did not. - `diff.interHunkContext` merges nearby hunks. Reproduced: at `20` on a 40-line fixture with two separated edits, two hunks become one and the patch goes from 22 lines to 34 -- so every `diff.patch:` a rater cites after that point names a different line on an operator who sets it, and the adjudicator re-reading the citation reads something else. - `diff.suppressBlankEmpty` strips the leading space from blank context lines. Both are pinned to their defaults now. More usefully, the universal claim has moved out of the comment and into the test that can actually hold it: `_gitconfig` builds a hostile `~/.gitconfig` and the packet built under it must be byte-identical to the one built without, so a knob that reaches the diff and is not pinned fails there rather than being asserted away in prose. The comment now says the list is incomplete until that test says otherwise. The widened fixture fails against the previous commit. Co-Authored-By: Claude Opus 5 * Address review round 6: a config key could hide a submodule change from the packet entirely `changed_submodules` ran `git diff --raw` without the config pins, and `diff.ignoreSubmodules=all` empties that listing. Reproduced end to end: for a change that bumps `vendor`, `--raw` prints `:160000 160000 15d7664 4138482 M vendor` by default and nothing at all under that key. In the builder that means the refusal finds nothing to refuse, the patch carries no `Subproject commit` line either, and `materialize_tree` skips gitlinks by design -- so the rater is handed a packet in which one of the change's edits is not present at all, and it verifies against its own manifest. People set that key to quiet noisy submodule diffs, so it reaches a real machine. Unlike the five rounds before it, this one is not byte drift. It is the same failure the branch opened with -- content the rater cannot see and cannot tell is missing -- reached through the operator's config instead of the tree's attributes. `_DIFF_CONFIG` now carries `diff.ignoreSubmodules=none` and goes to the `--raw` call too, and `diff.srcPrefix` / `diff.dstPrefix` (git >= 2.45, which outrank the `diff.noprefix` already pinned) are pinned with it. The hostile `~/.gitconfig` in the tests sets all three; both new guards fail against the previous commit. Co-Authored-By: Claude Opus 5 * Clear precondition 3 against a codex that runs, and make the documented command work `npm install -g @openai/codex@latest` + `codex login` worked, so the harness could finally be checked against the CLI instead of against memory. Three things came out of that, and one of them was a live blocker. - **The documented command did not run.** `run_rater.py` imports `agents_shipgate.schemas` at module scope with no path bootstrap, so `python benchmark/safety-qualification/rater/run_rater.py …` — the spelling in `calibration.md` and `cut-c-preconditions.md` — failed. On this machine it failed *confusingly*: an uninstall left `site-packages/agents_shipgate/` as an empty husk, which imports as a namespace package, so the error was `No module named 'agents_shipgate.schemas'` rather than `'agents_shipgate'`. The script now puts this checkout's `src/` first on `sys.path`, which also shadows any stale install. Verified with the ambient conda interpreter that produced the failure. - **Shared mode would have refused on any ordinary machine.** It rejected a profile whose `config.toml` mounts MCP servers; the real profile mounts two, and shared mode is the *only* mode an OAuth login can use — isolated mode authenticates strictly through `ANTHROPIC_API_KEY` / `OPENAI_API_KEY`. So the round could not have run in either mode. `codex exec --ignore-user-config` ("do not load `$CODEX_HOME/config.toml`; auth still uses `CODEX_HOME`") is exactly that split, so the config refusals go and the flag replaces them. What is left is a *non-empty* `AGENTS.md` / `AGENTS.override.md`, which the flag is documented not to cover; an empty one instructs nobody. - **`--strict-config` and `--ephemeral` are now passed.** The second is codex's `--no-session-persistence`. The first is the one that changes an outcome: without it a config key codex does not recognise is ignored in silence, so `sandbox_mode`, `tools.web_search = false` and `history.persistence` could each have been absent from a session that reported nothing wrong. `_ISOLATED_CODEX_CONFIG` passes it, and a control run with one bogus key is refused, so the pass is not vacuous. Precondition 3 is cleared: every flag exists on 0.153.0, the config is accepted, and a live session completed with the event shape `openai_final` parses. Precondition 2 still needs `claude login` — re-confirmed on 2.1.259, where the restriction checks still hold but the token is still expired. Co-Authored-By: Claude Opus 5 * Address PR review: web search, an atomic family claim, NUL-but-decodable, and the env channel All four review findings reproduce, and each fix has a test that fails on the previous commit. - **[P1] Shared-mode codex had web search.** `--ignore-user-config` removes the profile's `config.toml` and the branch supplied no replacement, so `_ISOLATED_CODEX_CONFIG`'s `web_search = false` did not apply and codex's documented default for an unset value is `"cached"`. `-c 'web_search="disabled"'` is now passed on the command line in **both** modes. The spelling is pinned by `--strict-config`, which accepts `web_search="disabled"` and rejects `web_search=false` (measured). - **[P1] The family gate was time-of-check around a minutes-long session.** Two roles started together each read no sibling *label*, each recorded `unchecked`, and both wrote same-family labels. Replaced with `claim_family`: an `O_EXCL` claim under `/claims/` written **before** the session, and the sibling read **after** — so whichever way two runs interleave, at least one sees the other, and if they share a family that one refuses. Re-running a role is still allowed while the family is unchanged. The regression test is barrier-backed, and fails on the previous commit with two same-family labels written. - **[P1] UTF-8 decodability is not a text test.** `b"before\x00tail"` -> `b"after\x00tail"` is binary by git's NUL heuristic (`--numstat` reports `-`), yet `--text` emits the NULs and `decode("utf-8")` succeeds, so the "genuinely binary refuses" boundary was open and a rater's Read and Grep may truncate the patch without saying so. A section is now non-text when it fails to decode **or** contains a NUL, and the refusal names which. - **[P2] `GIT_DIFF_OPTS` reopened the machine-dependent manifest.** `-c` outranks every config *file*, but that is not a config file: measured, the same two pins under `GIT_DIFF_OPTS=-u20` gave a 48-line patch where the default gives 22. Git subprocesses now run with that and the repo-redirection variables dropped. `GIT_CONFIG_*` is deliberately kept — it is config, `-c` does beat it (measured), and the determinism test sets `GIT_CONFIG_GLOBAL` precisely to prove that. Also records the owner's role assignment for the round, which Amendment 1 asks to be kept with it: `security_governance` → `claude` (2.1.259), `framework_tooling` → `openai` (`codex-cli` 0.153.0). Precondition 2 is cleared — `claude auth login` (not `claude login`, which is a prompt, and not `claude auth status`, which reports `loggedIn: true` on an expired token). docs: correct two stale qualification claims, and bind them to what they claim (#513) Both were found while explaining the release-evidence gate, and both are the same class: a current-state fact restated in prose with nothing checking it. - `benchmark/safety-qualification/README.md` said the policy pins report schema `0.42`. `pre_release_safety_requirements()` pins `0.43`. The pin itself was never wrong — a test holds it equal to what the engine emits — only the sentence describing it. - The same file cited "the exact six-variable contract" for `docs/distribution.md`, which documents four. The other two moved into `.github/release-trust-roots.json` when the signer identity and OIDC issuer were deliberately taken out of variables, so an actor who can set variables cannot also replace the identity that vouches for the evidence. The count is now dropped rather than corrected: an unguarded number in prose beside a table that owns the fact is what drifted in the first place. - `README.md` contradicted itself — one section said reports carry `0.43`, Limitations still said `0.42` and named `v0.41` as the frozen reference. `docs/agent-contract-current.md` and `docs/architecture.md` were already bound to `CURRENT_REPORT_SCHEMA_VERSION`; these two sentences were not, which is why a bump left them behind. `test_the_prose_that_states_the_current_report_schema_states_the_current_one` now covers both, matching the sentence that makes the claim rather than searching the file for a number — the files also carry legitimate frozen-schema references that a bare number search would trip over. The runbook's claim is checked against the *policy* value, since a separate test already holds the policy equal to the engine. Verified by reverting each fix: both reversions and a deletion of the runbook sentence fail the new guard. Not touched: `prompts/fix-top-finding.md` carries "(contract v26 / report v0.42)" in four byte-identical copies pinned by `tests/test_agent_instructions_renderers.py`. Changing it needs the bundled- prompt bump procedure, and it may be a deliberate provenance stamp rather than drift — that is a decision, not a typo fix. Close out Cut B: pin the walked candidates and source the last cell (#456) (#507) * feat(benchmark): close out Cut B — pin the walked candidates and the last cell (#456) Sourcing for the `pre_1_0` corpus is finished: all 60 slots are `pinned`, and nothing in the inventory needs mining before the Cut C calibration round. Three things stood in the way, and one sweep (`benchmark/miner/results/2026-W36-closeout`, 45 rows over three repositories) closes all of them. - **The last gap.** `langchain_crewai × insufficient_evidence` is claimed by `bytedance/deer-flow#4868`, following Session B's lead to mine an *application* that calls `MultiServerMCPClient` at agent construction rather than the adapter library. Its tools are assembled at run time from an out-of-tree extensions config, and the change makes which credential a tool call carries depend on the run-time user and a `$ENV_VAR` map. - **Three walk candidates, pinned.** `github-mcp-server#3020` and `#3076` and `grafana/mcp-grafana#1080` were carried with abbreviated or absent SHAs. `#3076` is why the convention matters: its walk note's head `5ea9a0e8…` is `refs/pull/3076/head`, which after a squash merge is not reachable from the default branch at all. - **The guard the pinning section names was not in the tree.** It was added in Cut A and removed by Cut A's own review commit, which added the sentence citing it. Restored and generalized: every external pin is re-read from the sweep that resolved it, a pin no sweep corroborates is refused, and a subject a sweep did resolve may not sit `unpinned`. Two smaller things fall out. Filling the last gap makes the "a gap that names a PR plans the origin that PR can supply" tripwire vacuous, so it is re-pointed at the reserve — which states the same origin-beside-state pair, and is where the next candidate comes from. And the Cut B finding that the SDK repositories' example trees are "no longer picked up by cold-start `init`" is corrected: reproducing it at the recorded pin gives `refused_unresolved_scope`, the monorepo behavior working as designed. No engine change: nothing under `src/` moves, per the Cut B rule that a fix made in response to a candidate is what turns it into tuning material. Co-Authored-By: Claude Opus 5 * fix(benchmark): close two holes the close-out's own guards left open (#456) Review of the close-out found four things; two are guards that failed open on the shape they exist for, and two are claims stated more confidently than the evidence supports. **A sweep file that contradicts itself was read as agreement.** `_swept_pins` keyed recordings by sweep file, so one file holding a subject twice with different SHAs kept whichever row came last: a corrupt duplicate ordered *before* the good row disappeared. Reproduced against the pre-fix keying — the disagreement check and the pin comparison both passed. Keyed by the pins instead, with the sweeps that recorded them as the value, so any two recordings that disagree fail whether they sit in one file or two. **A reserve row whose state can supply no origin was skipped rather than failed.** `_reserve_claims` dropped every row whose `State` was outside `STATE_ORIGINS`, which includes `open` — so `adk-samples#1745`, the candidate that taught this project an open PR is not history, could have been reserved as `real_history` and passed. A row that states an origin now has to state a state that can supply one. **Two claims corrected.** deer-flow#4868 merged five days before the latest-40 window's oldest PR, not six weeks: that repository merges forty PRs in six days, where crewAI-examples' forty reach back eighteen months, so `--limit 40` is not a time window and a busy repository's silence under it means very little. And `init`'s refusal reports the agent-defining projects it found while warning the list is incomplete, so the narrative no longer states the count as an enumeration. Guard sweep re-run: eight perturbations, all fail closed. Co-Authored-By: Claude Opus 5 * docs(benchmark): say what the close-out counted, and what the pin guard refuses (#456) Second review pass, three prose corrections and one addition. - "The three walked MCP servers all read `tools_scanned=0`" counted candidates as servers: three pinned candidates, two repositories. The walks covered four candidates across three repositories, so the sentence contradicted the inventory it summarizes. - The close-out run note now says *why* it stays off the `*-mined` glob in its own terms rather than by analogy: deer-flow's trigger-skip rate is 14 of 40, because it is an agent application chosen for one cell, not the unselected sample the noise bound measures. - The pinning section now states what the restored guard does to a candidate that cannot be mined at all — a private design-partner repository fails it rather than passing on a hand-written pin, and accepting one anyway is an owner's decision rather than a silent exception. Co-Authored-By: Claude Opus 5 * docs(benchmark): tighten the close-out narrative (#456) Co-Authored-By: Claude Opus 5 * test(benchmark): make every results run load-bearing, and commit the perturbations (#456) Addresses the five review findings on this PR. **[P2] Every results CSV is pin authority, so every results run is now integrity-checked.** `test_miner_corpus.py` validated well-formedness, CSV/JSONL parity and LF endings over `*-mined.jsonl` only, while `_swept_pins` reads *every* committed CSV as the authority for a corpus candidate's pins — so a half-regenerated `-cutb`, `-closeout`, `-reeval` or `constructed` pair could leave stale pins certifying the inventory with nothing in CI to notice. The three integrity guards move to an all-runs enumeration; only the KPI aggregation stays on `*-mined`, because the trigger-skip noise bound is only meaningful over unselected sweeps. Two guards keep the split honest: the enumerations may not converge, and a results CSV with no JSONL sibling — which would escape the enumeration entirely while still supplying pins — fails. **[P2] The two holes fixed on this branch now have regression tests.** Every committed input is valid, so the suite passed either way: reverting `_swept_pins` to per-sweep last-write-wins, or `_reserve_claims` to skipping a state outside `STATE_ORIGINS`, went undetected. Both parsers now take their input, the row-versus-sweep rule moved into `_pin_complaints` so a fixture runs through the same code the inventory does, and the perturbations are committed: a contradictory duplicate inside one sweep **in both row orders** (only one of them ever failed), an uncorroborated pin, a resolved subject left `unpinned`, an `open` PR reserved with a real origin, and a misspelled reserve origin. Each was confirmed to fail against the reverted implementation and pass against the fixed one. **[P3] Three consistency fixes.** The `#3076` row's note still instructed the reader to resolve both SHAs and led with the discarded `5ea9a0e8` PR-head abbreviation; it now states the resolution. The close-out's run-table row still counted three candidates as three servers. And the inventory carried two current statuses: the opening paragraph, the origin-floor reading and the order-of-work line all still described sourcing as future work. Make a preview wheel report one version instead of two (#491) (#505) The first published preview — `0.16.0+preview.20260902.gcc59410`, cut from `cc594109` — shipped the defect this fixes, and it was found by installing it rather than by reading it. The version lives in two places. `pyproject.toml` decides the wheel's METADATA; `src/agents_shipgate/__init__.py` carries `__version__` as a literal, which is what `--version` and `doctor` report. `release-preview.yml` stamped only the first, so the published wheel's METADATA said `0.16.0+preview…` — C1's index refusal held — while the installed CLI said plain `0.16.0`, indistinguishable from the qualified release it exists to be distinguishable from. The second-order effect was worse. `doctor` compares `installed_version` against `imported_version` and emitted `installed_version_differs` reading "Two installed copies are shadowing each other; which one runs depends on path order" — on an environment holding exactly one copy. A diagnostic built in #334 to be trustworthy was made to lie by an ordinary install. The fix is not "stamp both files", which is an intention. It is a check on the **artifact**: the build opens the wheel it just produced and requires `METADATA: Version` to equal the `__version__` packaged inside it. That holds however many version sites exist and however the stamping is spelled. Stamping now covers both sites, each proved by its own round trip over a closed file set, and a companion test pins the committed tree to exactly two sites that already agree — so a third cannot be added and silently left unstamped. A four-case perturbation sweep confirms each guard bites, including a direct replay of the shipped defect. `docs/release-evidence-policy-decision.md` § Amendment 2 records the field note, because the amendment asserted this property before anything enforced it. The general lesson is in there too: C1's index refusal was structural and held; C1's self-description was prose, and did not. Release on a cadence, admit an unqualified preview channel, and prepare v0.16.0 (#491) (#503) 81% of `CHANGELOG.md` had never reached a user: 5,039 of 6,242 lines sat under `## Unreleased`, two milestones were complete and untagged, and the newest published build was 56 days old. The pipeline was not the defect — the cadence was. Releases were gated on a single expensive artifact (#456), so they happened when that artifact happened. **The cadence is a policy with a number.** `docs/release-runbook.md` § Cadence fixes a 30-day interval and a 45-day defect threshold, names the release owner, and says what happens when the interval is missed. `scripts/release_cadence.py` counts days since the newest release tag — `v` plus a complete PEP 440 version, so `preview-*`, `wip-sectiond` and `m3-pre-rebase` are correctly not releases — and CI prints it into every job summary, warning once the interval lapses. It warns rather than fails: a red check would fail whichever unrelated change arrived after the interval lapsed, and that author cannot cut a release. **The preview channel is admissible, and the finding says why.** `docs/release-evidence-policy-decision.md` § Amendment 2 records the owner's answer to the question #491 required be answered first, including the cases that would have reversed it. `release-preview.yml` publishes a GitHub pre-release at `preview-` carrying one wheel, built from a commit CI has already accepted using the release's own hash-locked toolchain. Five properties keep it from being read as a release, each with a negative control: 1. a PEP 440 **local** version segment, which a public index must refuse — so an unqualified build can never consume the `0.16.0` that PyPI's immutability would then make permanent; 2. a ref outside the `v*` trigger namespace; 3. a separate workflow file, so `stage` cannot find it when it looks for the mandatory rehearsal — structural, not a matter of an `if:` condition; 4. no qualification artifact, so `verify_safety_qualification_release.py` rejects it on both the missing artifact and `tag == v`; 5. `--prerelease`, so it is never `Latest`. Nothing in the release path was widened to accommodate it: `build_manifest` still requires `tag == v`, and a preview simply produces no candidate manifest. A 10-case perturbation sweep confirms each control fails when broken. **The release note is separated from the record.** `## Unreleased` had reached 338,932 characters — 2.7x the 125,000-character limit the release body is checked against — so no tag could have been cut from it whatever else was green. `CHANGELOG.md` now carries one line per change (15,755 characters), and the full reviewed prose moved verbatim to `docs/changelog/0.16.0.md`. All 165 entries are byte-identical in the record and all 165 have a line in the note. **v0.16.0 is prepared, not tagged.** `pyproject` moves 0.16.0b7 → 0.16.0 and the version propagates to every surface the guards bind to it; STABILITY.md's 16 pending migration notes are stamped. The tag itself still needs the signed qualification artifact (#456), the four `SAFETY_QUALIFICATION_*` variables, a real `signer_identity`, and a green rehearsal — and a preview is not evidence toward any of them. build(deps): complete the cryptography 50.0.1 and hypothesis 6.165.10 bumps with a recompiled dev lock (#452, #454) (#490) Dependabot raised the two declared floors in pyproject.toml but left constraints/dev.txt compiled from the old declarations, so both PRs failed the lock-consistency step ("constraints/dev.txt was not compiled from the current declarations"). This completes the bumps the way the lock is designed to move: the floors are raised, and constraints/dev.txt alone is recompiled with the pinned resolver (uv 0.12.5, matching constraints/release-publish.in), so the other three locks are untouched. Recompiling refreshes the dev closure: 19 pins move (build, coverage, cyclonedx-python-lib, hypothesis to 6.167.1, pydantic/pydantic-core 2.13.5/2.46.5 together, ruff 0.16.5, typer 0.27.2, ...), all patch or minor within their declared caps. scripts/verify_dependency_lock.py passes for all four locks; no environment markers dropped and no extras-qualified pins introduced. Build(deps-dev): Update mcp requirement from <3,>=2.0.0 to >=2.1.1,<3 (#453) Updates the requirements on [mcp](https://github.com/modelcontextprotocol/python-sdk) to permit the latest version. - [Release notes](https://github.com/modelcontextprotocol/python-sdk/releases) - [Changelog](https://github.com/modelcontextprotocol/python-sdk/blob/main/RELEASE.md) - [Commits](https://github.com/modelcontextprotocol/python-sdk/compare/v2.0.0...v2.1.1) --- updated-dependencies: - dependency-name: mcp dependency-version: 2.1.1 dependency-type: direct:development ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Build(deps-dev): Bump lxml from 6.1.1 to 6.1.2 (#450) Bumps [lxml](https://github.com/lxml/lxml) from 6.1.1 to 6.1.2. - [Release notes](https://github.com/lxml/lxml/releases) - [Changelog](https://github.com/lxml/lxml/blob/master/CHANGES.txt) - [Commits](https://github.com/lxml/lxml/compare/lxml-6.1.1...lxml-6.1.2) --- updated-dependencies: - dependency-name: lxml dependency-version: 6.1.2 dependency-type: direct:development update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Staff Cut B across three sessions and prepare Cut C for the pre-1.0 corpus (#456) (#489) * feat(benchmark): settle the Cut B sourcing contract — constructions, pins, and claims (#456) Cut B is staffed as three parallel sessions coordinating through the inventory CSV. Before they start, the contract they claim slots against has to exist: where a corpus-built synthetic lives (never samples/), what makes it a case (base/ and head/ that differ, CASE.md beside them, no engine output inside), how real history enters (a sweep, then a miner label, so the slot stays holdout-eligible), and the pin form for closed-unmerged and reverted PRs, which the merge-commit convention cannot express. The guard gains a constructed_design basis, a by-name engine_tests detector for constructions, and reads n8n as the top-level section it is rather than a tool-source type it is not. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): enumerate closed-unmerged and reverted PRs in the miner (#456) The safety-qualification corpus needs a rejected_or_reverted vein that merged history cannot supply. `mine --state closed` enumerates PRs closed without merging, pinned at the fork point and the PR head; `--state reverted` finds merged PRs a later Revert PR undid, pinned like any merged PR with the revert recorded in notes; `--pr N` mines one named PR in whichever state it is in. Merged enumeration is untouched. Co-Authored-By: Claude Fable 5.1 * docs(benchmark): bring the labeling guide to the four-way corpus vocabulary (#456) Restructure benchmark/miner/LABELING.md so the top of the file is a self-contained rater rubric for passed / review_required / insufficient_evidence / blocked, with a first-draft rule for the review_required vs insufficient_evidence line, constructed illustrations that name no corpus case, a mapping to the miner's three labels, and the JSON output contract. The miner process moves below a 'not a rater input' heading and the real-PR anchor that stated a corpus case's label and verifier verdicts is removed. A guard test keeps the guide free of inventory candidate refs (both spellings), slot samples, and SHIP- check IDs, and requires the four decisions as headings above the miner section. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): add the rater packet builder and session harnesses for Cut C (#456) build_packet.py turns an external (clone + base/head SHAs) or constructed (base/ + head/ trees) case into the three inputs Amendment 1 condition 2 allows: repo/ (the head tree, no .git, verifier output and CASE.md excluded), diff.patch, a byte copy of LABELING.md, plus a role-specific TASK.md and a MANIFEST.json hashing every file. It refuses a source tree that carries the strata inventory. run_rater.py runs one fresh read-only session per family (claude -p with Read/Grep/Glob only and no MCP; codex exec --sandbox read-only), archives the complete transcript content-addressed, and parses the final message into an IndependentHumanLabelV1 with reviewer_id ::. Anything but exactly one valid JSON object with a known decision fails closed. --home-mode isolated runs --bare in an empty HOME (needs ANTHROPIC_API_KEY); shared keeps the caller's HOME under file-level memory checks. --dry-run prints the command and environment. Co-Authored-By: Claude Fable 5.1 * docs(benchmark): pick the five calibration cases (#456) Name the five non-corpus cases for the Amendment 1 calibration round: four merged reserve PRs pinned as merge commit + first parent from a clone (mongodb-js/mongodb-mcp-server#1417, awslabs/mcp#4489, stripe/ai#353, openai/openai-agents-python#3518) and one constructed blocked-shaped LangChain case under calibration/cal-5/, since no reserve PR has that shape. Each entry records why the diff tests the guide and where two raters are expected to diverge; no label or verdict is recorded. The file discloses that the chooser had read the inventory, which is why none of these can ever become a corpus case. Co-Authored-By: Claude Fable 5.1 * fix(benchmark): skip a revert title that names something that is not a PR (#456) aaif-goose/goose has a merged revert whose quoted title carries a truncated `(#65…` reference; `gh pr view 65` fails and the whole reverted enumeration aborted. One bad reference now skips that revert. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): construct the MCP declared-binding synthetic case for Cut B (#456) mcp_export_adds_undeclared_tool: a tool server with a complete root declaration publishes a fourth tool the declaration omits (the #432 shape), built outside samples/ so the slot stays holdout-eligible. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): construct the OpenAI Agents SDK synthetic cases for Cut B (#456) sdk_agent_adds_ticket_update_tool adds a scoped Zendesk write with no approval policy; sdk_agent_loads_tools_from_registry replaces the literal tools list with a runtime registry so the bound surface is not enumerable. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): construct the LangChain and CrewAI synthetic cases for Cut B (#456) crewai_tools_from_factory builds the crew's tools from a profile factory; langchain_agent_adds_refund_tool adds a Stripe refund tool with neither an approval policy nor idempotency evidence. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): construct the Google ADK synthetic cases for Cut B (#456) adk_agent_adds_calendar_toolset adds a domain-scoped calendar write through a resolved McpToolset (the adk-samples#1975 shape); adk_billing_sub_agent_refund puts a Stripe refund on a sub-agent (the adk-samples#1745 shape) with no approval policy. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): construct the n8n synthetic cases and reserves for Cut B (#456) Five slot cases on one support-orders workflow family: a POST tool with no approval, a Code Tool whose request target is built at run time, a Call Workflow Tool with an expression target, a DELETE tool behind a public webhook, and a disabled send-and-wait approval node. Two reserves: a read-only GET tool and an allowlisted MCP Client Tool write. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): construct the multi-agent handoff synthetic cases and reserve for Cut B (#456) A triage root with an Agents SDK handoff to a billing sub-agent: a write added on the sub-agent, handoffs resolved from an environment route list, and an approvals sub-agent that decides the refund requests the root submits. One reserve adds a scoped read on the sub-agent. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): claim the Session A constructions in the strata inventory (#456) Fifteen synthetic gap slots move to pinned with target_basis constructed_design, exposure none and their CASE.md as evidence; the register gains one row per construction and three reserve rows; the summary tables and the shape prose are recomputed from the CSV. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): sweep n8n, unwalked MCP servers and the rejected vein for Cut B (#456) 2026-W36-cutb: 912 rows over 21 repositories, the first post-#403 run. Every n8n repository this project had never mined, eight unwalked MCP servers, and closed-unmerged / reverted / named PRs from the new miner states. Labeled from the PR diffs as one session's Cut B cell-targeting labels, not adjudicated. Named off the *-mined glob because it is a selected, capability-dense sample (trigger-skip 0.71), not the unselected history the cross-run noise-bound guard measures; the README says why. Co-Authored-By: Claude Fable 5.1 * feat(benchmark): claim the slots sourced by the W36 sweep (#456) Eighteen slots pinned from 2026-W36-cutb: seventeen of the eighteen gaps Cut B session B owned, plus a third google_adk x insufficient_evidence slot because adk-python#6605 turned out to be engine_tests exposure. Every rejected_or_reverted slot is filled from five repositories; n8n has its first two real-history candidates. The guard gains a `reverted` state (landed history whose rejection came afterwards) mapping to rejected_or_reverted only. langchain_crewai.insufficient_evidence.1 stays a gap with the re-mine's findings as its lead. Summary tables recomputed. Co-Authored-By: Claude Fable 5.1 * docs(benchmark): strike the calibration cases from the reserve and reread the shape after Cut B (#456) Co-Authored-By: Claude Fable 5.1 * fix(benchmark): close three holes in the blind-rater boundary (#456) Review findings on #489, all in the Cut C rater harness. 1. The packet is re-hashed at launch. _check_packet() proved five paths existed and parsed the manifest; it never compared the packet against it. Any edit landing after the build -- a note under repo/, a rewritten guide, a doctored diff -- reached the rater while the label recorded the hash of a manifest describing a packet that no longer existed. prepare() now calls verify_manifest() and refuses before a session is launched. 2. Isolated mode gives codex its own home. HOME was replaced and then CODEX_HOME pointed back at the caller's real profile, which is where codex reads global AGENTS.md, AGENTS.override.md, and a config.toml that can mount MCP servers and enable web search -- so the mode named isolated was neither blind nor packet-limited. It now builds a fresh Codex home per run holding one written config (no MCP servers, web search off, read-only sandbox, no approvals, no history) and authenticates through OPENAI_API_KEY, symmetric with the Claude path; shared mode keeps the real profile only after proving it is silent. 3. Symlinks cannot carry unmanifested content. The copy preserved links verbatim and the hasher skipped them, so a link out of the tree was content the rater's read tools resolve and the manifest cannot describe -- and the packet still verified. A link that escapes the tree or dangles now refuses the build; one that stays inside is materialised so its bytes are hashed; a link found in a packet is tamper, not a skip. Each fix is covered by tests that fail without it. Add action declaration review evidence (#488) * Add local review, declaration review, and MCP code discovery * Address expert review findings * Address expert review and schema synchronization * Reduce PR to declaration review and advance contract v29 * Sync protected surfaces to report v0.43 / packet v0.18 The schema bump left AGENTS.md and the canonical SKILL.md naming report v0.42 / packet v0.17, which failed nine public-surface and docs-link contract tests. Both are protected release surfaces, so `preflight` routes them to a human; the repository owner approved this edit. Bumping them alone is not sufficient, and the cascade is the reason this was left for a person: - `skills/agents-shipgate/SKILL.md` has two byte-identical mirrors under `adoption-kits/` and `plugins/`, enforced by `test_plugin_skill_is_byte_identical_to_canonical`. All three move together or the build trades nine red tests for a different one. - Editing the canonical SKILL.md changes the rendered adoption kit, so `EXPECTED_CLAUDE_CODE_SKILL_RENDER_SHA256` moves to b09ffc59… and the outgoing 6aedbd71… is appended to `prior_render_sha256`, keeping already-adopted copies recognisable as unmodified. - A line linking the current schema must also name the version it supersedes, so the frozen-reference pointers move v0.41 -> v0.42. - `llms-full.txt` is generated from AGENTS.md and is regenerated here. The `Report schema (v0.17 frozen reference)` row keeps its `0.17` version column: it is a frozen pointer, not the current schema, and no test covers that cell. Full suite green. Publish the capability delta as a standalone in-toto attestation (#470) (#487) * feat(verify): publish the capability delta as a standalone in-toto attestation (#470) `verify` already computed the capability delta and bound it to a receipt, and its only consumers were our own renderers. It now writes `agents-shipgate-reports/capability-delta-attestation.json`: an in-toto Statement whose `predicateType` is `https://threemoonslab.com/agents-shipgate/capability-delta/v1` and whose predicate carries the frozen `shipgate.capability_payload/v1` delta view unchanged. Runtime contract 27 -> 28 registers the predicate type, both schema versions, both schema paths and the artifact — the half #469 deliberately declined while nothing emitted the payload. No second payload shape and no second computation. The predicate embeds `project_capability_delta`'s output verbatim, and `diff_capability_locks` — what the PR comment renders — is the same engine, so the capability-change count on both surfaces is one value by construction (#433). The attested subject is the reviewed tree, and `delta.head.ref` must equal it: without that join, four edited characters would move a valid delta onto a commit it never reviewed. Only a committed-tree subject is attestable; a worktree run withholds the file and says why on `verifier.base_notes[]`. `predicate.verification` chains to the receipt through `input_set_id` and `subject_id` and is `bound` or `unbound` — only `bound` may name identities. `request_id`, `engine_requirement_id` and `decision_id` stay out: they mix in the engine build and the platform, so an interchange format carrying them would emit different bytes for an identical review. `analysis_coverage` is populated per side from the run's own conservation law, observed catalog minus analysed subjects, so a tool that arrives and is never bound reaches the attestation as a named `newly_outside_analysis` row instead of being silently absent (#437). A side the run could not establish publishes `unavailable`, which is not a claim that nothing was left out. `tools/verify-capability-delta.py` is the reference consumer: stdlib-only, one file, importing nothing of ours, applying 20 published rules and rejecting a relabelled subject, an inflated summary, a forged subject key, a downgraded coverage status, a forged binding, and a silently escalated effect. Nothing gates: no verdict, no severity, no release impact. `release_decision.decision` remains the only release gate. Co-Authored-By: Claude Opus 5 * fix(#470 review): close four holes the self-review found The reference verifier answered a malformed `subject` with a `TypeError` traceback instead of the rule list a caller reads: the `--expect-*` checks indexed the subject list after `verify` had already collected the real problems. Subject extraction is now total, a broken `--schema` file is a result rather than a crash, and six malformed shapes are pinned. A membership entry carrying *neither* record satisfied every remaining rule — no changed dimensions, direction equal to the transition, no explanations — so it would have published an `added` change naming no capability at all. Both sides are now checked, not only the one the transition forbids. Dropping a catalog row with no canonical id silently removed a subject from `observed`, and an unobserved subject can never be reported as outside analysis — the #437 defect reached by a different route. Both readers now return `None` for the whole side, which publishes `unavailable`. `coverage_from_scan` refuses a fact set under a different agent than the observed rows are keyed under. Sharing no keys would have published every analysed tool as outside analysis: over-reporting is the safe direction, but it is still a false statement about a named subject. Plus: the spec page's sample output now matches what the script prints, and `base.ref` says what it names when the base facts come from the committed reviewed envelope rather than a fresh scan. Co-Authored-By: Claude Opus 5 * test(#470): stop the new tests costing six minutes of CI The `test` job went 484s -> 847s and was cancelled at its 15-minute cap during "Build package". Every substantive step had passed, but a cancelled check is a red check and the job has no headroom, so the cost is the defect. Three causes, all mine. Five tests each ran a full base-and-head `verify` on the same fixture to read five fields out of the same reports directory. They were never five scenarios, they were one — so they are now one test with the assertions grouped and labelled, and the run happens once. `test_the_emitter_needs_nothing_from_a_source_checkout` cost another whole verify to make a claim the wheel test makes more strongly. Deleted. Eight tests each ran two full adapter passes over the shipped sample to get capability facts to project. Facts are frozen values and a pure function of that sample, so they are built once per worker now. The one test that is *about* two builds agreeing still asks for uncached facts, or it would be comparing a value with itself. And the wheel test extracts `built_wheel` instead of consuming `installed_wheel_site`. Session fixtures are per *worker* under xdist, so a second consumer of the install fixture can add a whole wheel build and pip install to another worker's run — for a test that needs the packaged file tree, not a resolved distribution. Co-Authored-By: Claude Opus 5 * fix(#470 review): make the standalone verifier enforce the closed contract Three findings from review, all in the published consumer path. **Stage one now runs natively, before anything derived.** The membership-change branch returned before it looked at the record it carried, so `effect: "harmless"` — outside the closed v1 enum — reached `valid: true` on the documented 30-second command, which omits `--schema`. An external policy engine could act on an out-of-contract safety value. `S1`-`S7` implement the closed object shapes, declared types, vocabularies, id and digest patterns, I-JSON integer bounds, array constraints, and the permission profiles the lattice can actually produce. Every count check is `type(v) is int`: Python's `bool` *is* an `int`, so `isinstance` accepted `added_subjects: true` and then `True == 1` satisfied every arithmetic rule downstream. Order is the fix, not the checks alone. The derived rules read the document as though it conforms, so they now run only when the structure holds and `S0` says when they did not. **`--require-receipt-binding` checked a self-declared string.** Any statement could carry `bound` and two well-formed content ids and pass, while the page presented the option as the gate check for the receipt chain. `--receipt ` is the check: `R1` joins `input_set_id` and `subject_id` against the receipt supplied, and `R2` requires that receipt's artifact manifest to bind these exact bytes. `R2` is the load-bearing one — identities alone would accept an attestation from a different run over the same inputs. The old flag survives, documented and helped as the shape-only check it is. **Malformed documents escaped as tracebacks.** Thirty shapes did, all nested values the derived rules reached before anything checked them; the structural pass removes the class. A 2,770-case mutation sweep over the shipped example is committed and produces rule rows and no stack trace. A top-level backstop turns any unanticipated exception into an `internal` rule row so the published exit contract holds even for a shape nobody foresaw. The sweep found two gaps of my own: the two state refs could declare a different `capability_standard_version` than the payload (`P12`), and a permission profile could state "measured and harmless" (`S7`), the combination the payload spec says must be unrepresentable. 20 rules -> 31. Every vocabulary, pattern and bound the script restates is pinned against the package, the same way the rank tables already were: a drifting copy widens the language accepted, which for a safety vocabulary is the failure that matters. Separately, the CI `test` job is split. It ran 456-839s against a 900s cap on `main` and was cancelled there on a commit unrelated to this branch; this PR adds ~16s of 2694s, so raising the number would have been treating the symptom. The suite runs as three parallel `suite` jobs and the coverage gate is combined and enforced once in a `coverage` job. `conftest.py` assigns whole test *files* deterministically from the collection alone — whole files because module- and session-scoped fixtures are per file — and `tests/test_shard_partition.py` asserts the union of the shards is the suite, because a partition that drops a file leaves every job green while a test stops running. Co-Authored-By: Claude Opus 5 * fix(ci): the coverage job needs coverage, not an editable install `pip install -e . --no-deps --no-build-isolation` needs the backend closure from `constraints/build-backend.txt`, and this job installs only `constraints/dev.txt`, where `hatchling` is not pinned. It would have failed on every run. The install was unnecessary in the first place: the job never imports the package. `coverage combine` and `coverage report` read the data files and resolve sources by the absolute paths recorded in them, which are this checkout's — verified by running both with the project off `sys.path`. Add side-effect-contained local review flow (#486) * Add side-effect-contained local review flow * Deny authorization for provisional manifests * Address local review feedback Read an MCP server's tool surface out of its own source (#431) (#483) * feat(inputs): read an MCP server's tool surface out of its own source (#431) `detect` reported the official MongoDB and Grafana MCP servers as not an agent project. Both publish dozens of tools — `drop-database`, `delete-many`, `update_incident` — and neither commits an export, so every route discovery had was filename-shaped and found nothing. The discriminator was never the language (two of the three servers walked are Go); it was whether the repository happens to emit its tool list. `tool_sources[].type: mcp_server_source` reads the tool name where every one of them writes it: as a string literal at a registration site. It runs no code, evaluates no schema library, infers no type, and reads no annotation the source declares about itself. Which sites count comes from a built-in, versioned registry of named registration idioms, never from a configured pattern — at `detect` time there is no manifest, and post-adoption a configured regex would hand the definition of evidence to configuration shipping in the same pull request as the tool it could hide (#268). Five idioms ship, chosen by a 30-server survey published at docs/mcp-registration-idioms.md. Python's FastMCP decorator is the largest measured shape and is deliberately deferred: a different extraction mechanism, and #393 requires each one to bring its own probe list. Reading is done over a masked copy of the source in which comments and string bodies are overwritten, so a registration can never be found inside a comment or another string. Extraction is `medium`; completeness is measured per file. A name built at runtime is reported as an unenumerated subject in the exclusion ledger rather than dropped, and a file that could not be read holds the whole source at `partial`. `detect` offers the route only where a declared MCP dependency and a resolved registration hold together, and stands down for a committed export, which stays `high` and is named in `excluded_sources`. `TRIGGER-MCP-TOOL-REGISTRATION-SOURCE` routes such a diff on tokens derived from the same registry. Neither the trigger catalog's `schema_version` nor `report_schema_version` changes. Co-Authored-By: Claude Opus 5 * fix(inputs): close two fail-opens the review found in the MCP source reader Both were silent misses of the kind this input exists to end. A class body was carried as the `ts_static_tool_name` site's span, because the sibling `operationType` and `description` literals are looked up in it. The containment rule then read *any* registration written inside that class as "the same registration", so a class whose `toolName` is built at runtime lost its omission the moment its body also called `.registerTool(` — the tool was neither enumerated nor recorded. A span is now always the construct that registers; only a call's argument list is a wrapper. Discovery's `max_source_files` slice was dropped when the loop was rewritten, while the `truncated` flag beside it was still computed and published. `detect` therefore walked an unbounded number of source files on an unknown workspace, and told the reader "this count is a lower bound" for a count that was exact. Also: `this.description = "…"` could be read as the tool's description and reach the questionnaire, because the lookbehind admitted a leading dot; the detection score awarded the dependency point by list index rather than by matching the dependency reason itself; the dependency probe used a narrower skip set than the source walk, so a manifest under `coverage/` could grant a language gate whose files that walk skips; and `_relative_source_path` took a `root` argument it never used, under a docstring describing a fallback the body does not implement. Each has a regression test naming the failure it prevents. Co-Authored-By: Claude Opus 5 * docs: link the two follow-ups the #431 survey and review produced are named where a reader meets the gap: the survey document, the script's own simplifications list, the test that pins the omission, and the changelog. Co-Authored-By: Claude Opus 5 * test(mcp-idioms): make the prefilter test observe the shortcut it names A perturbation sweep over the new invariants caught nine of ten. The tenth — deleting the prefilter from `scan_source` entirely — changed no observable behaviour, because the test compared the result of scanning a token-free file to an empty scan and both are empty either way. The test was about nothing. It now monkeypatches `mask_source` to fail, so the claim it makes — that such a file is answered *without masking it* — is the thing asserted, and it reasserts soundness by scanning a positive sample with the patch lifted. The sweep now catches all ten. Co-Authored-By: Claude Opus 5 * fix(inputs): address PR review — decode per language, and bind the export to the surface it displaces Four review findings, two of them actions silently lost. **Go escapes were decoded with JavaScript's grammar.** One decoder shared between two languages is a mistranslation rather than a parse error: Go writes an octal escape as three digits, so `MustTool("delete\137all", …)` registers `delete_all` and this reader produced `delete137all` — the real action absent from the catalog and an action id nobody serves standing in its place. Each language now decodes with its own grammar, and anything neither grammar defines is refused rather than guessed. A refusal reaches `_resolve_name` as `name_not_literal`, so the tool becomes a recorded omission; guessing is the one outcome a reader of a *name* cannot afford. **An unrelated or partial export erased the source route.** Withholding on the mere existence of an export anywhere in the workspace meant that in a repository holding two servers, an export committed for one deleted every source-only registration of the other, and a partial export deleted the rest of a single server's surface. The test is now containment: the export must name every tool the source route resolved. Where it does not, both routes are suggested and the shortfall is named — two sources describing one server are reconciled by a reviewed `tool_identity` binding (#386), never by dropping one. A wildcard export enumerates nothing, so it can never be shown to contain anything. **A regex could begin a control statement's body.** The slash heuristic read `)` as always ending a call, so `if (ok) /\.registerTool("fake", h)/.test(v)` had its pattern scanned as code and reported a `fake` tool — a registration invented out of a regex body, which is exactly what the masking pass exists to make impossible. The keyword in front of the matching `(` now decides, so `foo(a) / 2` still divides. **Discovery decoded more leniently than the adapter.** `errors="replace"` let `detect` resolve a registration from a file `load_mcp_server_source` then refuses as `unreadable_file`, so the route it named enumerated fewer tools than it promised. Discovery now reads through the adapter's own `load_text_file`. Each has a regression test naming the failure it prevents, and the perturbation sweep over the module's invariants now runs 15 for 15. Add replayable incident fixture suite (#481) * Add replayable incident fixture suite * Harden incident replay expectations * Keep incident replays within CI budget * Address incident fixture review feedback Clarify verifier causes, evidence gaps, and preview routing (#482) * Fix verifier explanation and routing gaps * Address PR review feedback Freeze one capability schema for the exported delta and the committed state (#469) (#479) * Freeze one capability schema for the delta and the committed state (#469) Two planned public surfaces serialize the same internal truth: the exported capability delta published as a standalone attestation (#470) and the committed capability state (#474). Each defining its own serialization would ship two divergent schemas of one structure. `shipgate.capability_payload/v1` is now the one payload both consume, frozen before either exists. - `docs/capability-payload-schema.v1.json` — the JSON Schema, generated from `agents_shipgate.schemas.capability_payload` and gated on drift. - `docs/capability-payload.md` — the spec: identity keys, the required and optional split, the deliberate exclusions, and the evolution policy. - `agents_shipgate.core.capability_payload` — the one projection that fills it, over the existing capability-fact and capability-delta engines. One document, two views discriminated on `view`. Subject identity follows the subject-based counting rules: `subject.key` is derived from agent, provider and tool id and deliberately not from the subject kind, so a tool and its action are one row carrying two changes and the +2-for-one-tool shape (#439) cannot be expressed. `summary` and each subject's `transition` are recomputed from the rows on parse, and a payload that disagrees with them is rejected rather than repaired. `analysis_coverage` carries the subjects the analysed surface left out — an added-but-unbound tool produces no capability fact, so without it the first surface that had to report one (#437) would have needed a second payload shape. Neither `not_requested` nor `unavailable` means zero, and only `complete` may name subjects. The published field set is closed. Every model forbids unknown properties, each excluded internal field is recorded with its reason, and a test asserts those maps cover every `CapabilityFactV1` field, so a new internal field cannot reach either surface — or be dropped from both — without a written decision. `classify_semantic_permission` is split out of the permission lattice so the payload and `mcp audit` share one classifier. Nothing emits the payload yet, by design: no command, no artifact, no check, and no change to `contract_version`, `report_schema_version`, `.well-known`, the capability lock, or any existing schema. `release_decision.decision` remains the only release gate. Co-Authored-By: Claude Opus 5 * Address review: unconditional identity guard, empty-delta invariant, type parity Five findings from the review pass on this branch. - `_merge_subject_refs` returned early when the key *and* the display name matched, so the identity check below it never ran on that path. Two subjects colliding on the 64-bit `subject.key` and sharing a name would have merged silently — the conflation the schema exists to prevent. The identity check is now first and unconditional; a matching name only shortcuts the rename reconciliation after it. - A delta with no subject rows could declare `base` and `head` digests that differ: an attestation saying "nothing changed" between two provably different states, which every validator accepted. No rows means every fact matched on every dimension, so both digests are equal by construction and the invariant is free to enforce. It is now enforced and published in `STABILITY.md` and the spec. - The payload re-spells two closed vocabularies it cannot import without inverting the schemas/core layering. Nothing pinned them, so widening `PermissionClass` or the `reversibility` literal internally would have left CI green and raised a `ValidationError` from inside the projection on the first adopter repository that produced the new value. Two tests pin them: the permission vocabulary against the lattice, and — more generally — every published field's annotation against the internal field it projects. - `canonical_payload_json` passed `default=str`, so the digest an external consumer verifies could be computed over a lossy `str()` rendering, and two distinct values whose text form coincides would digest identically. It now raises instead. - `test_subject_key_is_derivable_from_published_fields` compared the helper against itself, so a change to the digest input or truncation would have moved both sides together and left the recipe published for external consumers silently wrong. It now recomputes the key from stdlib exactly as the spec states, and a companion test fails if the spec stops stating it. Co-Authored-By: Claude Opus 5 * Address second review pass: self-verifying state digests, global capability id Two findings from a second pass over the diff. - `evidence_set_digest` is keyed by `capability_id`, but uniqueness was enforced only *within* a subject row. A repeated id across rows silently dropped one capability's provenance from the digest, so two states with genuinely different provenance would digest alike. `capability_id` is now unique across the whole payload, on both views, and the digest builder refuses a collision rather than hashing a set that is one entry short. - A `state` payload carries every row its digests are taken over, and the spec tells consumers those digests are recomputable from the payload alone — but nothing checked them, so a state could declare digests that did not describe its own rows and a digest-comparing consumer would see no change. The state payload now recomputes and rejects a mismatch, exactly as it already does for `summary`. A delta's `base`/`head` refs name states it does not carry and stay on trust; the state payload for each side is what proves them. The digest recipe moves from the projection into the frozen format (`schemas.capability_payload.state_digests`), so the projection and the validator that runs on a consumer's parse cannot disagree about what a state hashes to. `STABILITY.md` and the spec page carry both guarantees. Co-Authored-By: Claude Opus 5 * Address third review pass: a subject's transition is its presence, not its changes `subject_transition` derived `added`/`removed`/`modified` from the *kinds of a subject's changes*, but a delta row carries only the capabilities that moved. So a tool with two operations that loses one produced a single `removed` change, rolled up to `transition: "removed"` and `summary.removed_subjects: 1` — for a tool still fully present in head. Reproduced against samples/ai_generated_refund_pr. The mirror case reported a tool that gained a second operation as an added subject. That is the #439 defect in the other direction: a question about subjects answered in changes. And the row shape could not express the truth, because unchanged capabilities are absent from a delta row by design. `CapabilityDeltaSubject` now carries `present_in_base` and `present_in_head` explicitly, read off the two fact sets and never off the changes. `transition` is recomputed from that pair on parse — `added` when the subject is not in base, `removed` when not in head, `modified` otherwise — and presence bounds the changes in the other direction too: a subject absent from base can only carry `added` ones, and one absent from head only `removed` ones. A subject present on neither side is not a row at all. Freezing the schema is the point of this issue, so this had to be settled before the format ships rather than after. The spec page, `STABILITY.md` and the CHANGELOG state the rule; the superseded rollup test is replaced by three that pin presence, and the generated delta example is regenerated. Co-Authored-By: Claude Opus 5 * Verify the published digest recipe, and state the verdict-free rule Fourth pass, both about what the spec promises an external consumer. - The digest recipe was documented as recomputable from the payload alone but only ever checked against `state_digests` itself, so a change to what the digest covers would have moved both sides together and left the published recipe silently wrong. It is now recomputed from the serialized payload with stdlib only, exactly as the page states it, and a companion test fails if the page stops stating it — the same treatment the subject key already had. - The payload carries no release impact, severity, or verdict for a subject, and now says so and why: publishing a per-subject impact in an interchange format invites a consumer to gate on it, which is a second verdict by another name. The evolution policy also names the mechanism that decides a version bump — the type-parity test against the fact models — so a widened internal vocabulary is a decision someone makes, never a ValidationError an adopter discovers. Co-Authored-By: Claude Opus 5 * Cover the reidentified projection path and the sorted-row guard The reidentified branch of `project_capability_delta` had no test: no shipped sample produces a scope change, so the one reviewer-relevant case — a tool whose authority scope widened — went through an untested path. It is also the case where getting the pairing wrong is most visible: a published `+1 added, -1 removed` for one broadened scope reads as a new tool and a lost one. The test pins that it stays one subject, one `reidentified` change, direction `broadened`, with both capability ids carried. Also covers the sorted-row guard, which enforces a published determinism guarantee and was only ever exercised in the passing direction. New-module coverage is now 99% and 97%; every remaining line is a fail-closed raise that is unreachable by construction. Co-Authored-By: Claude Opus 5 * docs: correct the rule count and the capability-id scope on the spec page The structural-rules list grew to five while the lead-in still said three, and the identity section still scoped `capability_id` uniqueness to a subject row after it became a payload-wide invariant. Co-Authored-By: Claude Opus 5 * Address the blocking review: ten contract-level fixes before v1 is frozen All ten P1 findings reproduced and fixed, plus the three P2 follow-ups. **Determinism.** `financial` and `production` share a permission rank, so a rank-only sort key inherited hash-randomized set iteration: the same assessment produced `('production','financial')` under one hash seed and the reverse under another, changing the published bytes and `capability_set_digest`. The sort key is now total, with a cross-process regression test over six seeds. **Interoperable canonical bytes.** `json.dumps` escapes non-ASCII by default, so a Python producer hashed `café` where a JavaScript consumer hashes `café` — two digests for one identity, in a format whose whole point is independent consumers. Canonicalization is now UTF-8 and unescaped, with sorted ASCII keys, integers only and no non-finite numbers, agreeing with RFC 8785 within those constraints, fully specified in the spec, and pinned by Unicode test vectors. **A published permission expansion was labelled provenance-only.** The fact layer folds the semantic assessment into `evidence_hash` alone, while this payload publishes a `permission` block derived from it. A structural-to-inferred change moved the published permission from `('write',)` to `('write','unknown')` with side effects newly unknown, and the delta called it `evidence_only`. The projection now classifies what it publishes and emits a `permission_changed` explanation, and the model rejects `evidence_only` whenever the two published records differ semantically. **A delta could not represent a newly unbound subject.** One coverage snapshot cannot distinguish a tool added and unbound from one unbound since before the change, and only the first is something a reviewer of that diff must act on — the two produced byte-identical payloads. A delta now carries both sides plus the recomputed `newly_outside_analysis` / `no_longer_outside_analysis`, with a comparison only as established as its weaker side. **Promised wire fields are required.** Defaults kept `capability_payload_schema_version`, `view`, `analysis_coverage` and `subjects` out of the schema's `required` arrays, so the committed schema accepted a payload with no version and no discriminator. Every field of every object is required now, and producers pass the constants explicitly. **The subject key is validated against its identity material.** Uniqueness alone did not make "one subject, one row" true: splitting one tool across two rows under invented keys validated after recomputing the digest, so `+2` was still expressible. `CapabilitySubjectRef` recomputes the published recipe. **Every published field is bound by a state ref.** `evidence_hash` sat in neither digest and coverage in neither, so two different valid payloads could share a `CapabilityStateRef`. The evidence digest now covers the record's evidence hash, and a third `analysis_coverage_digest` binds coverage; a delta checks each side's coverage digest against the coverage it carries. **Transitions are validated against the records they carry.** `changed` with identical records, `changed` across different capability ids, arbitrary dimension names, and `added` claiming direction `removed` were all accepted. `changed_dimensions` is now derived from the two records' digests, `changed` requires the same identity and `reidentified` requires a different one, and a membership change must carry the matching direction. **The generated schema no longer overclaims.** Pydantic validators do not reach `model_json_schema()`, so a consumer handed the one file got far less than the spec promised. Everything JSON Schema can express is now in the published file — required keys, patterns, non-negative counts, non-empty lists, the transition/sides/direction coupling, the presence/transition coupling, the coverage naming rule — pinned by fourteen negative tests run against the artifact rather than the models. The rest is named as stage two in the schema's own description and enumerated in the spec. **The evolution policy now matches the validators.** Promising additive-within- v1 while shipping `additionalProperties: false` and closed enums was promising something a v1 consumer does not do. `v1` is closed: any addition is `/v2`, old versions stay published for at least a minor cycle, and a producer migrates by emitting both. P2 follow-ups: the stale "rolled up from changes" wording is gone from the generated schema; a rename now publishes the head spelling instead of the lexical minimum, which was as likely to be the name no longer in the tree; and the permission and state-ref models reject classifier-impossible tuples and inconsistent counts. Co-Authored-By: Claude Opus 5 * Close the trust-boundary drift: eight release-contract gaps before v1 freezes All eight P1s from the follow-up review reproduced and fixed. The common thread was the one the review named: the JSON Schema, the reference parser, the producer and the spec each accepted or described a different set of payloads. **Snapshot both sides at API entry.** `project_capability_delta` walks each side in the diff, subject-ref, presence and state-ref passes, and `_state_subjects` walks it again. A caller-owned list answering differently on a later pass produced a valid payload whose delta row named one revision while `head` digested another — reproduced with a `list` subclass. Both entrypoints now take one tuple snapshot per side and read nothing else. **Enforce the canonicalization domain instead of asserting it.** `capability_id` becomes a dynamic object key in the evidence preimage, and Python orders keys by code point where RFC 8785 orders by UTF-16 code unit — `""` and `"😀"` sort opposite ways. It is now constrained to the producer's canonical `cap_[0-9a-f]{16}`, so every key in the payload is ASCII by enforcement. Every published integer is bounded to the I-JSON safe range, because `9007199254740993` reads back as `…92` in JavaScript. The spec states the Unicode and integer rules and publishes cross-language vectors, pinned by a test so the table cannot drift from the code. **Stop reusing the open internal change model as a frozen wire type.** `CapabilitySemanticChange` defaults `before`/`after` to `None` and types them `Any`: omitting both validated and was silently repaired, despite the every-field-required contract. `CapabilityChangeFact` is payload-owned, with a closed `kind`, required nullable values inside the canonical domain, and a per-dimension `direction`. An exhaustive test now asserts every object in `$defs` has `required == properties`; that model was the only miss. **Derive the semantic fields rather than trusting them.** Relabelling a row `broadened`, or deleting its explanations, both validated. `semantic_direction` and `semantic_changes` are now computed from the two carried records — dimension by dimension, over the published content only — and the parser rejects anything else. The direction is therefore the direction of *what this payload publishes*, which is the distinction that produced the earlier `evidence_only` mislabel; `evidence_only` now means exactly "equal apart from provenance" and cannot be claimed. This also replaces the projection's local classifier, so there is one implementation. **Reconcile the state refs with the membership rows.** `head.subject_count - base.subject_count` must equal added minus removed subjects, and the same equation holds for capability counts over added/removed record transitions; a one-added-tool delta claiming head counts of 100/100 validated before. It also makes an empty delta require equal counts. **Make the reference parser schema-strict.** It coerced `summary.subjects="2"`, `present_in_base="false"` and `side_effect_unknown="false"` where the published schema refuses them, so the entrypoint the spec calls the reference implemented a larger language than the file it ships. Every scalar is strict now, while JSON arrays still populate tuple fields. **Move more rules into stage one.** The artifact now also rejects a subject absent from one side carrying changes that side cannot have, classifier- impossible permission tuples (including the reciprocal `unknown` rule the parser itself was missing), membership changes carrying explanations, non-canonical ids, unsafe integers, and stringified scalars. Twenty-six negative cases run against the committed file rather than the models, and the stage-two table now lists the two rules that genuinely need a computation. **Correct the normative document.** It described two digests and the wrong provenance preimage, showed only the state's coverage shape, and said an empty delta needs all digests to agree — which would have rejected the coverage-only exactly, both coverage shapes are shown, and the empty-delta rule is narrowed to the capability and evidence digests with the reason. Publish the determinism boundary: a generated coverage matrix of what shipgate can prove (#473) (#478) * feat(boundary): publish the determinism boundary, generated from the adapters (#473) `insufficient_evidence` outside the boundary reads as "this tool cannot analyse my repository" unless the boundary itself is a published specification. `docs/determinism-boundary.{md,json}` is that specification: per built-in input and per declaration shape — export artifact, literal registration, factory, dynamic construction — what the scan reads, the `Tool.source_type` it produces, the extraction-confidence ceiling that route reaches, and what that ceiling means for a release verdict. Nothing on the page is hand-written. Each adapter declares its coverage beside the code that mints its confidences, and the consequence column is derived by asking the engine's own predicates about a probe tool built from those declared facts. A restated rule would be the #433 route-table mistake, and it would already be wrong: a wildcard MCP export reaches a `high` ceiling and still cannot be pass-eligible, which only the engine can say. Generation is fail-closed both ways. An adapter registered without coverage raises; a source type added to AST_ONLY_SOURCE_TYPES or MCP_SOURCE_TYPES that no cell emits raises; a coverage leaving a shape unanswered, or claiming a ceiling the engine caps, is refused at construction. Four committed negative controls hold that closed. `extraction_is_complete()` is now the one definition of "the adapter read this tool's contract with full confidence", shared by the semantic resolver, `low_confidence_tool_count`, and the generator. Every `insufficient_evidence` verdict links to the page — `scan` stdout, the GitHub step summary, and `report.md` — from one string, so the three surfaces cannot word it differently. Co-Authored-By: Claude Opus 5 * fix(boundary): correct the published threshold and close three review gaps (#473 review) The page's central quantitative claim was wrong. It said `insufficient_evidence` arrives once low-confidence actions reach half the analysed surface — the `_LOW_CONFIDENCE_TOOL_RATIO` clause. That clause is real and never fires first: every action below `high` also raises an `incomplete_surface` semantic issue, and `evidence_below_ie_threshold` treats `semantic_coverage.gap_count > 0` as sufficient on its own. One medium action in ten already withholds a verdict, and the page told that reader they were under the bar. `surface_flags` was a fail-open inside a fail-closed generator. A flag the completeness predicate does not read is inert on the probe, and inert reads as "surface complete" — so a singular typo would have published `proven` for a Codex server that names no tools at all. `SURFACE_INCOMPLETE_ANNOTATIONS` now holds the three keys once, `surface_is_complete` reads them from it, and a cell declaring anything else is refused. `configured_as` published the five per-scan framework inputs as `tool_sources[]`-only, so the ADK row said "a `tool_sources[]` entry" beside a sentence naming `google_adk.tool_inventories[]` — reachable only through the section the page never mentioned, and the one route to a `proven` row for those inputs. The section is now declared per adapter and validated against `AgentsShipgateManifest`: deriving it from `scope` instead would have published a `conductor:` key that does not exist. Also: `verify` prints the reference on `insufficient_evidence` — it is the command AGENTS.md sends agents to for PR evidence and was the one such surface without it — and the committed-equals-generated test loads the generator through `importlib.util` rather than putting `scripts/` on `sys.path`. Co-Authored-By: Claude Opus 5 * fix(boundary): pin the threshold claim to the predicate, not the prose (#473 review 2) The `ConfigurationRoute` comment written for the previous fix repeated the error that fix corrects: it said every `per_scan` framework adapter answers to both routes. `conductor` is `per_scan` and answers only to `tool_sources[]`, and four inputs answer only to their section — which is why the route is declared rather than read off `scope`. The threshold guard asserted on the page's wording. It now asserts the mechanism: one semantic gap among ten actions already satisfies `evidence_below_ie_threshold`, while `_low_confidence_tool_threshold(10)` is greater than one — so the ratio clause demonstrably is not what binds, which is the whole reason publishing it was wrong. Co-Authored-By: Claude Opus 5 * test: merge the duplicate release_decision import in the threshold guard Co-Authored-By: Claude Opus 5 * fix(boundary): correct six routes the matrix described wrongly (#478 review) Every finding reproduced against the real loaders, and each fix carries the probe that found it. **Wildcard inventories were published as `proven`.** An inventory declaring `wildcard: true` is a reviewed file that names nothing: it loads at `high`, carries `wildcard_tools`, and proves no surface. That is the inventory an adopter is most likely to write first, and the page told them it was the way out. All five inventory-bearing inputs now publish a `wildcard inventory` variant at `set_unproven`. **The ADK module-wide downgrade stopped at `google_adk_function`.** `_resolve_extraction_evidence` also lowers the `mcp` and `openapi` actions a resolved toolset contributed; a probe with one unresolved expression returns that MCP action at `medium`. Both the `factory` and `dynamic_construction` shapes now publish that cross-cell downgrade as its own route. **Conductor's two dynamic fields do not behave alike.** Only a non-literal `method` withholds the action; a literal method with an expression-backed `mcpServer` still creates `conductor_mcp_call` at `medium`. Split into two variants, so the page no longer advertises a false case. **n8n's dynamic route was missing entirely.** A `toolName` expression still emits `n8n_workflow_tool` at `medium` while recording the unresolved name — the action is not hidden, its identity is. Added beside the MCP-client wildcard, which is now a named variant rather than the whole shape. **`codex_plugins:` cannot activate its adapter.** `load_codex_plugin_artifacts` returns nothing without a `tool_sources[]` row, so the section supplements what that row finds. `manifest_section_role` distinguishes the two, and only an activating section counts as a configuration route. **"No check runs on it" was false for `not_extracted`.** A dynamic Conductor call emits no `Tool` and still raises a HIGH `SHIP-CONDUCTOR-DYNAMIC-TOOL-SURFACE-NOT-ENUMERABLE`, which withholds the verdict. Cells now carry `raises`, validated against the check catalog at generation time, and the outcome sentence says so. Also, both P2s. The remedy is now per-source, derived from `inventory_manifest_key()` — there is no inventory route for `sdk_function`, `conductor_mcp_call`, or `codex_config_mcp`, and the blanket promise sent those adopters after a manifest key that does not exist. And the matrix stamps the release it describes, so a link stored in an archived report can be checked against the scanner that produced it. Co-Authored-By: Claude Opus 5 * fix(boundary): apply the `raises` correction to every input, not one (#478 review 2) Reviewing my own fix for the P1 "no check runs on it" finding: the review said the claim was untrue *globally*, and I had made it true for Conductor only. LangChain, CrewAI, n8n, and Google ADK all feed a `SHIP-*-DYNAMIC-*` check from the exact construct their factory/dynamic rows describe, so four inputs kept the wrong impression while the sentence pointing at the `Raises` column implied the column was authoritative. All five now declare what they feed, pinned by a test that enumerates the pairs. Two defects in the fix itself: - `_inventory_key` returned the first non-`None` match over cells, so an input emitting source types from two frameworks would publish whichever cell came first — a remedy that moves with an unrelated edit. It is a set now, and two keys fails generation rather than picking one. - `manifest_section_role="supplements"` with no section rendered "The top-level `None:` section supplements ...". Refused at construction. `raises` is also documented for what it is: the checks a route feeds directly, not every finding an action from it might later attract. Co-Authored-By: Claude Opus 5 * fix(boundary): stop denying an inventory route that exists (#478 review 3) Two more of the same class, both introduced by the previous fix. The per-source remedy said `codex_plugin` "has no `tool_inventories[]` key", but its `proven` route *is* a reviewed inventory, declared at `codex_plugins.mcp_tool_inventories[]`. The derivation reads `inventory_manifest_key()`, which is the engine's *gap-prescription* table and does not know that key — so the fallback branch turned "the engine prescribes nothing here" into "no such file exists", which is the P2-1 error inverted. The line now claims only the former. And a third route feeding a check with an empty `Raises`: the OpenAI Agents SDK factory records a toolkit scope bound, which is exactly what `SHIP-SCOPE-TOOLKIT-UNBOUNDED` reads. Added to the enumerating test rather than fixed in isolation, since "true for the input I was looking at, false one input over" has been the recurring shape of this whole review. Co-Authored-By: Claude Opus 5 * fix(boundary): derive the wildcard check instead of declaring it eight times (#478 review 4) A sweep for routes that feed a check with an empty `Raises` found the recurring shape once more, and this time at scale: `checks.inventory` raises `SHIP-INVENTORY-WILDCARD-TOOLS` on *any* tool annotated `wildcard_tools`, which is eight cells across six inputs — every wildcard export, every wildcard inventory, the wildcard host-config server, and the n8n MCP client wildcard. Declaring it per cell would have been a second table for a relationship the engine already owns, and on this review's record the cell that forgot it would have been the one that mattered. It is derived from `surface_flags` now, so no wildcard route can omit it. That also sharpens the wildcard-inventory rows: they read as `set_unproven` *and* a HIGH finding, which is the pairing that makes the row actionable rather than merely discouraging. The same sweep cleared the rest. ADK's unresolved *tools expression* feeds an evidence gap rather than a check — its `factory` row is the one wired to `SHIP-ADK-DYNAMIC-TOOLSET-NOT-ENUMERABLE`, and that is already published. The SDK's dynamic binding produces a source warning. CrewAI's prebuilt route only attracts a finding conditional on `environment.target`, which is an action-level consequence `raises` deliberately excludes. Co-Authored-By: Claude Opus 5 * docs(boundary): lead with the reader's question, not the taxonomy (#478 feedback) The page was a reference document. Someone arriving from an `insufficient_evidence` verdict passed 78 lines of vocabulary — `set_unproven`, `sdk_function`, `incomplete_surface` — before reaching anything about their own repository, then met a seven-column table keyed on internal source-type tokens. It now opens on their question. The summary table's columns ask how *their* repository declares tools ("…written out in code") instead of naming the taxonomy, and its cells read ✅ proven / ⚠️ not proven / ✖ not read. Each framework then gets a plain answer — "Read, but not proven from your source alone: shipgate reads the tools you write out and can say what each one does; what it cannot do is prove that list is all of them" — followed by the fix that applies to it. Nothing was dropped. The full seven-column matrix is still published per input, inside a collapsed block, and the shape and outcome tokens are defined once in a reference section at the bottom. The JSON companion is unchanged, so the machine contract did not move. The answers are derived from the outcomes the engine already computed — four branches over (has a proven route, what the code route reaches) — because a hand-written paragraph per adapter is exactly the drift this page exists to prevent. Two defects in the rewrite, both now guarded: the summary anchors did not follow GitHub's rule, so `#openai-agents-sdk-(python)` — the one row a Python adopter would click — resolved to nothing; and the first ordering answered "yes, from a contract file" for LangChain, whose reader wrote Python and needed to hear that their code *is* read, just never proven. Co-Authored-By: Claude Opus 5 * docs(boundary): give the printed reference the same treatment as the page (#478 feedback) "Coverage boundary: what a scan can establish per input and declaration shape" named the page's internal axes. A reader who does not already know what a "declaration shape" is cannot tell whether the link is worth a click — and that reader is the entire audience for an abstention. It is now "What Agents Shipgate can prove, per framework", which says what they get and mirrors the page's own title, so clicking lands somewhere recognisable. Two guards, because this line has four printers and one of them will be edited alone eventually. The reference may not contain `declaration shape`, `coverage boundary`, `extraction`, `source type`, or `ceiling`; it must be one line ending in the URL; and it must promise what the page's title delivers. Separately, `scan`, `verify`, the step summary, and `report.md` must all reference the constant and none may retype the literal. Four sample goldens regenerated, and the conductor control pointer rebound to the new `report.md` bytes. feat(benchmark): commit the Cut A strata inventory for the pre-1.0 corpus (#456) (#477) * feat(benchmark): commit the Cut A strata inventory for the pre-1.0 corpus (#456) Maps the known candidate pool onto all 28 profile x decision cells the `pre_1_0` policy requires, so Cut B mines the empty cells instead of re-finding the full ones. 56 slots, two per cell: each is a pinned candidate, an identified-but-unpinned one, or a gap carrying the lead that would close it. It is a sourcing plan, not evidence. No label, no verdict, no receipt; nothing in it reaches the qualification runner. It is also not an admissible rater input -- it names a target decision for every slot, so Amendment 1's blindness condition makes a session that has read it unable to produce an admissible label. `target_basis` is a closed vocabulary with no verifier-derived member, because a corpus assembled to match the engine's own verdicts cannot measure the engine. `tests/test_strata_inventory.py` derives the grid from `pre_release_safety_requirements()` rather than restating it, closing the one definition site the decision document says no gate can detect. It also re-reads every cited miner label and pinned SHA from its source, and refuses the two escapes: restating a labeled subject's basis as `diff_substance`, or dropping the pins off a subject an existing sweep already resolved. The plan's shape is the finding: 28 of 56 slots have a candidate, the origin floor (23 qualifying cases) binds well before the case count, `n8n` has one sourced slot of eight, and the sweeps behind the pool ran before #403 -- so their zero-trigger repositories are unexplored rather than empty. Refs #456 Co-Authored-By: Claude Opus 5 * fix(benchmark): close the origin-floor rename and reground the candidate notes (#456) Review of the inventory guards found four issues, all reproduced by perturbing the CSV and confirmed fixed the same way. 1. Nothing bound an in-tree candidate to `synthetic`, so relabeling the twelve `samples/` slots `real_history` reported 41 of 56 qualifying origins against a floor of 23 while adding no real evidence -- the origin floor satisfied by renaming. What a candidate *is* now decides both its origin and its split, checked against the path, and the holdout test consumes that decision instead of restating it. 2. `sample_design` was the one basis whose evidence was never cross-checked: a row could cite any directory that happened to exist. It must now cite the sample it is about. 3. Slot numbering was validated in file order, so re-sorting the CSV to review origin coverage failed a test on byte-identical content. Cell membership and 1..n numbering are now checked as sets; naming another cell, or numbering 1,4, still fails. 4. The `.template.csv` skip in the label reader was unreachable -- the `*.labels.csv` glob never yields one. Replaced with a comment saying why the pattern, not the branch, is what excludes them. Also regrounds all fifteen `human_label` notes in the miner rationales they cite. Several paraphrases were wrong in a way that matters for a sourcing artifact: #3451 is a least-privilege tightening release, not a "runtime handling fix"; #232 removes the toolkit's action and permission bounds rather than merely migrating to MCP; #312 rewires the skill supply chain from an authenticated endpoint to an unauthenticated fetch, which is why it is a trust-root case at all. Refs #456 Co-Authored-By: Claude Opus 5 * test(benchmark): bind the inventory's reading to the plan it describes (#456) Second review round. `strata-inventory.md` restates the CSV in three tables -- 28 of 56 sourced, the per-profile shortfall, the per-outcome scarcity -- and those numbers are what a corpus owner reads to decide where to mine. Nothing kept them in sync, so the first added candidate would leave the reading behind and send the next person to a cell the plan calls empty and the file calls full. Every number on the page is now recomputed from the CSV. The per-outcome figures move from prose into a table so one parser covers all three, and a duplicate numeric row label fails rather than silently shadowing the table nobody then checks. Refs #456 Co-Authored-By: Claude Opus 5 * test(benchmark): scope the register check to the register (#456) Third review round, two narrow fixes. The `diff_substance` check searched the whole document, so a candidate listed only under **Reserve** -- the section for candidates deliberately *not* placed in a cell -- satisfied a row claiming a register entry. It now reads only the register section, and a per-profile row missing from the reading fails with a message instead of a KeyError. Refs #456 Co-Authored-By: Claude Opus 5 * fix(benchmark): record exposure per slot, and bind profile and origin to the PR (#456) Addresses the five review findings. Each was reproduced before the fix and re-checked after; the guard sweep is nine perturbations plus two policy moves. **[P1] Exposure is now recorded, and it decides the split.** Forcing every external PR to `either` was wrong: holdout means the engine was never tuned on the case, which is a fact about this project's history, not about where the bytes live. Six candidates are engine-development inputs -- #3020, `test_init_scaffold_disclosure` and `test_declaration_questionnaire`; stripe/ai#232 through a committed fixture tree; and openai-agents-python#3392, which produced `test_capability_change_schema_hash_parity.py` and the fix behind it. A new `exposure` column carries them, `split_eligibility` is derived from it, and a detector re-reads the tree as a *floor*, so a row may declare more than is findable but never less. `maintainer_walk` has no detector -- mcp-grafana#1080 appears nowhere in this repository and still drove `tool_sources[].binding` -- so `diff_substance` implies it, closing the one exposure anybody could drop. Three `mcp_openapi_declared_binding` cells gain a third slot whose only job is to be holdout-eligible. **[P2] The holdout floor is computed.** `ceil(count * fraction)` from the policy, not a hard-coded one. At 0.60 the inventory now fails in 14 cells. **[P2] Profile and merge state are bound to the register.** Every sourced candidate has a register entry; swapping two candidates' profiles now fails. For the five source-type profiles a sample's declared manifest type is checked too; `coding_agent_trust_roots` and `multi_agent_handoffs` are scenario profiles with nothing mechanical to check against, which the register says. **[P1] awslabs/mcp#4489 was misdescribed and misplaced.** It adds two literal FastMCP entrypoints, `budget-actions` and `budget-notifications`, taking the budget surface 1 -> 3; both read-only. Static literals are exactly what a grep-the-literal extraction reads, so this never supported an `insufficient_evidence` hypothesis. Moved to the reserve as a `passed` alternate; the cell is now a gap leading to a genuinely non-enumerable surface. **[P2] An origin is a fact about the PR.** adk-python#6605 was closed without merge, so the gap naming it plans `rejected_or_reverted`. Checking this also caught adk-samples#1745, which is still **open** -- not history at all -- so it leaves its slot for the reserve. **Recommendation taken.** `human_label` is renamed `miner_label` and disclosed as *not verifier-independent*: the labeling worksheet carries `head_decision`, `verify_verdict` and `verify_can_merge`, and LABELING.md tells the labeler those are enough to label without opening the diff. That biases cell targeting; it does not reach the corpus label, which Amendment 1 raters produce blind. A test fails if the worksheet ever stops exposing verdicts, so the disclosure cannot outlive its cause. One judgment is flagged rather than settled: `benchmark_scored` does not by itself block holdout, on the rule that being measured is not being tuned on. If any W24-W26 scoring drove engine changes not traceable to a test or a walk, those subjects need reclassifying, and only the owner can say. Refs #456 Co-Authored-By: Claude Opus 5 * test(benchmark): read the two register tables apart, and score samples too (#456) Two loose spots in the guards added for the review, found on a self-review pass and each reproduced by perturbation. The register parser read the Candidate register and the Reserve as one table. They share a `State` column, but the Reserve's second column is an *origin*, not a profile — so a reserved candidate was filed with `real_history` as its profile name, and a slot filled from the Reserve would have been "registered" against a profile that is not one. The two sections are now read apart, a Reserve entry carries no profile, and a sourced row that resolves to one fails. The exposure detector applied `benchmark_scored` only to external candidates, so the constructed sweep's scoring of `fixture://` samples went undeclared on four rows. Applied uniformly. It changes no eligibility — those rows are `shipped_sample` already — but a detector that skips a whole class of participation is a detector nobody can reason about. docs(release): record the pre-1.0 labeling protocol and the participant-validation gate (#456) (#476) Amendment 1 to the release evidence policy decision, recorded by the owner: - pre_1_0 primary labels come from two independent agent sessions on different model families, under six mandatory admissibility conditions (mechanical blindness, archived content-addressed transcripts, owner adjudication with walked-case disclosure, calibration round, artifact disclosure block). The beta (1.0) corpus commits to human primary labels; pre_1_0 evidence cannot qualify 1.0 by construction, so the protocol cannot leak upward. - Gate 2, participant validation: after labels freeze and receipts exist, each real-history case's author and reviewer receive a one-page case card and two separated questions. Five pre-registered rules fix what responses can and cannot change - frozen labels do not move; corrections land in the beta corpus; reviewers outrank authors; naming needs consent. - The corpus runbook now points at the amendment where it defines the two labels, and the Gate 2 card format, message template, and response log live in benchmark/safety-qualification/participant-validation.md. Fix capability delta subject and exclusion reporting (#468) * Fix capability delta subject reporting * Harden capability delta review projection * Prefer exact capability change identity Fix reviewed risk override evidence handling (#467) * Fix reviewed risk override evidence handling * Fix reviewed risk-tag replacement routing feat(mcp): report contradictory client annotations (#464) * feat(mcp): report contradictory client annotations * fix(mcp): keep contradiction deltas exact * fix(mcp): address annotation contradiction review Render cold-reader artifacts surface first (#466) * Render cold-reader artifacts surface first * Reuse audited Git probes for cold-reader ordering * Keep packet ordering fail-closed for block-tier findings * Address cold-reader review findings test(samples): pin the #424 repair loop with a committed artifact (#424) (#465) * test(samples): pin the #424 repair loop with a committed artifact (#424) The last open acceptance item on #424: "the repair loop is exercised by a committed artifact, not only unit tests". Everything the fix publishes was guarded, but every guard on the `declare_risk_tags` route ran against tools built in-test — the same blind spot that let the class ship, and the one `test_the_published_repair_closes_the_row_it_is_printed_on` sat inside while passing for the whole time the repair was swapping its row for a blocking conflict. `samples/google_adk_cold_start_agent` could not close it. A cold start has no declaration to challenge, so `declaration_below_inferred_evidence` is the one questionnaire shape it cannot render, and a challenged row is the whole subject of #424. `samples/declaration_repair_agent` is the step after that cold start. Every action is declared, controlled and owned, and two rows are challenged, so the sample sits at `insufficient_evidence` with exactly two open questions. Pasting the two blocks its committed `expected/suggested-declarations.yaml` publishes — each into the action it names, the `override:` alternative deleted as the block's own comments instruct, nothing else — reaches `passed`. Two rows, because the second is the case #424's own repair creates. It already carries a reviewed `risk_tags: [financial_write]`, and `risk_tags` is one YAML key, so a block naming it *replaces* it: a repair publishing only the newly uncovered category would tell that reviewer to delete their own tag and the next scan would reopen the row asking for it back. The committed golden publishes `[financial_write, destructive]`, and the test reads both the declared list and the published list out of committed files rather than restating either. The guards close over both routes a regression can take, each confirmed by perturbation: * revert the fix and `test_repair_scaffold_matches_its_golden` fails on the bytes; * regenerate the golden to make that pass — the natural next move — and `test_pasting_the_committed_repair_blocks_reaches_passed` fails instead, naming the row it reopened, with `test_the_published_repair_keeps_the_tag_the_reviewer_already_wrote` beside it. Every control the fixture declares was checked by dropping it and rescanning; each one blocks or holds back the *repaired* state. Two spellings that would have been redundant are deliberately absent, and the README says which and why. `report.md` gets a byte comparison so the sample has no committed artifact that nothing reads. No engine, schema or check-ID change: sample, goldens, registration and tests. Co-Authored-By: Claude Opus 5 * test(samples): pin the route, not only the row, on the repair fixture (#424 review) Review of the fixture: it asserted that a challenged row is committed and that pasting the published blocks closes it, but nothing said *which route* the blocks take. `declaration_below_inferred_evidence` publishes two, and only `declare_risk_tags` was broken by #424. A change that answered these rows by raising `effect:` instead would still close them, so the paste test would stay green while the fixture quietly stopped exercising the route it exists for — the one committed artifact for that route, silently repurposed. Confirmed by perturbation: force the route to `raise_effect`, regenerate the golden so the byte comparison passes, and `test_the_repair_golden_still_renders_a_challenged_row` is what fails, naming the block that switched. Co-Authored-By: Claude Opus 5 * docs(samples): say why the repair fixture declares no credential (#424 review) Both rows declare `authority: {mode: none}` on tools that delete customer correspondence and cancel billing mail, which reads as a modelling recommendation and is not one. It is what keeps every question this sample asks in the effect dimension: publishing a credential instead would add an `auth_scope` reading, and the readings here are deliberately keyword-only — the shape #424 reproduces on. Comment only; the goldens are byte-identical. Co-Authored-By: Claude Opus 5 * docs(samples): stop calling two different choices "the row's two routes" (#424 review) The README used "two routes" for both `effect_repair`'s kinds (raise the effect, or declare the category as a tag) and the block's two answers (keep the pre-filled tags, or reject them under `override:`). They are unrelated choices and the reader has to pick the right one to follow either paragraph. Co-Authored-By: Claude Opus 5 * test(samples): join the repair sentence to the value, and fix three doc claims (#465 review) Review findings, each reproduced before being addressed. **The one adopter-facing string #424 is about was committed and unread.** #424 reached a reader as a *sentence* — `ci/release_decision.py` templates it into the guidance an adopter sees — and this sample commits it in `evidence_gaps[].next_action.expects` and `agent_summary.first_recommended_action.why`, neither of which any comparison covers: `report.json` is checked only for schema version, decision and the set of field *paths*. That matters here specifically, because `EffectRepair.instruction` interpolates the whole list and the newly added part separately, and this is the first row in the repo where the two can differ. The only other test pinning `expects` for this route uses a row with no pre-existing tags, so `tags == added` there by construction. Confirmed: rewriting the f-string to interpolate `added` — the round-2 defect confined to the sentence, the block left complete — left all four new tests green while an adopter would read `risk_tags: [destructive]` on the release surface and `[financial_write, destructive]` in the block. The value-join now fails on exactly that, naming both surfaces. **The regeneration recipe dropped every normalization its sibling exists to record.** It is the designated recovery path — `test_repair_scaffold_matches_its_golden` names it in its failure message — and it gave a textual `` replace (a silent no-op on Windows, where `json.dumps` escapes the separators), no `generated_reports` normalization, and no `newline="\n"`, which made it the counterexample to `test_sample_expected_goldens_are_committed_with_lf_newlines`'s own docstring. Replaced with the sibling's runnable block, package-level import included; verified by extracting it from the README and running it verbatim against deliberately dirtied goldens. **Two claims were false.** `destructive` does not introduce `confirmation.required` — `external_communication` already obliged it on both rows — and approval is new only on `support.delete_case_message`, since `financial_write` already obliged it on the other; `safeguards.rollback` is the only control new to both. And the cold-start contrast was far too broad: that sibling declares two actions, pre-fills `effect: financial_write`, and prints a challenged *authority* note. What it has no example of is a declaration challenged in the **effect** dimension — zero `risk_tags:` blocks in its golden — which is the narrow claim now made in all four places it appeared. Also: the third test's docstring understated it. It is not a restatement of the byte comparison — canonicalize the reviewer's spelling (`financial_write` to `financial_action`) and regenerate, and the row still closes, the sample still reaches `passed`, the sentence still agrees with the block, and it is the sole failure. It holds "never rewrite what the reviewer wrote", which nothing else asserts. Docstring now says so. Cross-link back from the cold-start sibling, which was one-directional. feat(agent-mode): finish the control-envelope rollout and prove the adoption walk composes (#323) (#459) * feat(agent-mode): finish the control-envelope rollout and prove the walk composes (#323) on stdout. Three of #323's seven acceptance criteria were left open. **The two documented streams were two shapes.** `doctor`'s two failure routes carried `control`; `detect`'s and five of `init`'s did not — so whether a caller that routes on `control` could route at all depended on which setup command had failed and on which of its failures, and the run that most needs a route is the one that printed no payload to carry it. Every agent-mode error line from those three commands now carries the same envelope its `--json` payload would, projected from the same selected route as `next_action` and `next_actions[]`. `setup_failure_routing` builds it in one place, fail-closed: `execution: "failed"`, `decision: "setup_incomplete"`, `permissions` all false. Two lines still carry none, by design and by documentation: the shared `--workspace` refusal fires before a workspace exists, and `environment_error` is emitted before Shipgate is running. `AGENTS.md` said error lines carry no control object, which had not been true since #372. **The walk could not leave stage 2.** `init --write` over a manifest that already exists published `edit shipgate.yaml` with `expects: "The manifest reflects the desired tool sources, agent declared_purpose, and policies"` — a postcondition already satisfied every time the route was reached, because a manifest that does not load is claimed by the repair route above it. On this contract `next_action` *is* the step, so an envelope-only caller opened the file, found nothing to change, re-ran, and got the identical action back. That re-run is exactly the resume available after a person supplies the `declared_purpose` declaration. The route is now the `doctor` invocation for the manifest on disk. The exit code and the sentence a person acts on are unchanged; `next_action.kind` moves from `edit` to `command`. **Proved by walking it.** `tests/test_adoption_walk.py` takes an unadopted repository from `verify --preview` to a release decision as real subprocesses, choosing every step from the envelope alone. It fails a step that hands back the same action for an unchanged `input_id`, which is how the route above was found; reverting only that change fails the walk and nothing else. A second test covers criterion 1 as a table — each of the six commands at the one entry point it documents. A third sweeps the three setup modules' ASTs so a seventh error route cannot be added without the envelope. Contract 26 → 27; both the `AgentControl` union and the envelope schema are byte-identical, so `minimum_control_contract_version` stays at `21`. Co-Authored-By: Claude Opus 5 * fix(agent-mode): close the first review round on the setup failure routes (#323) Eight findings from reading the diff back. **The `--minimal` fallback still dropped the run it was correcting.** The `internal_error` route named `init --minimal` with no workspace, no `--write`, and no `--json`, so following it exactly produced a dry run against the process directory — the defect `_recovery_command` exists to prevent, and this was the third route spelling a recovery by hand. It goes through that builder now. The comment justifying the first pass was also wrong on its face: `NextAction` retargets `command` on construction, so the console-script spelling was never the divergence; the missing flags were. **The adoption-kit recheck could not read the kit.** It repeated `_requested_setup_flags` only, which deliberately excludes `--agent-instructions-kit` because a kit path is meaningful relative to a workspace. This rerun *is* the same workspace, so it now repeats the path along with `--minimal` and the accepted scope boundary: a recheck for a kit config that does not name the kit cannot confirm the edit it asked for. **The walk executed whatever `AGENTS_SHIPGATE_CLI` named.** Every emitted command is spelled for the entry point that produced it, and the walk runs those commands — so an exported override would have sent the fixture into a different install. It is cleared, with `AGENTS_SHIPGATE_ENABLE_PLUGINS`. **And a new test had the same defect from the other side.** The one asserting that `input_id` moves with the entry point distinguished the two by writing them into the command strings — which `NextAction` retargets, so under an exported override both collapse to one string and the assertion inverts. Confirmed by running the old shape with the variable set. The entry points are pinned through the variable now, which is the property under test. **Two failure routes had no test, and the cycle detector had no exercise.** `init`'s failed render and `detect`'s discovery failure are only reachable under a fault, so both are now pinned under monkeypatch — including that the discovery failure is a human route carrying no command. `_route_key` is exercised directly: a guard whose key did not discriminate would never fire at all, in either direction. **The walk checked the model and not the published document.** A payload accepted only in Python is published as valid to every consumer validating against `docs/agent-control-schema.v1.json`. Every step — six commands, four control states, both streams — is now validated against both. The schema rejects a setup operation claiming `release_decision`, a setup envelope granting merge, and a setup `complete`, so the assertion is not vacuous on the properties the walk relies on. **`check` does not publish a control pointer.** It binds its authority to `input_id` and answers through `--format agent-control-json`; three places had it lumped in with `scan` and `verify`. **Two miscounts.** "`doctor`'s two failure routes" is three, written into four files, and "all nineteen commands" was a count of `require_workspace` call sites rather than of commands. Also: one spelling of the `doctor` invocation for a manifest, one spelling of the failed-render message, and `test_errors_json_lists_every_runtime_emitted_kind` now matches all three emitter names — it matched only the base one, so `config_already_exists` dropped out of its scan when `init` moved to the wrapper, and `emit_agent_mode_error_action`'s call sites had always been invisible to it. Co-Authored-By: Claude Opus 5 * fix(agent-mode): a refused run must not lose what it was asked for (#323 review) Four review findings, and all four are the same defect at different depths: a route that reports it advanced the run while dropping part of what the caller asked for. On the shared envelope `next_action` **is** the step, so each of these was published as the authoritative answer. **[P1] A refused write handed the caller onward under the wrong pack.** `init --write --control-pack financial-strict` over a manifest selecting `default` refuses to overwrite and routes onward — to `doctor` on the plain path, to the gate on the `--agent-instructions` refresh path. Both read a manifest that loads fine and advance under the pack that is *there*, so an envelope-only caller proceeded under `default` having asked for `financial-strict`, with only a payload field recording the difference. The route now names `policies.control_pack` and the exact value, and reaches the onward step only when the manifest already matches. "Asked for" is `--control-pack` at other than its default — the same rule `_requested_setup_flags` already applies when deciding whether to repeat the flag, because Typer cannot tell an omitted option from one passed at its default and two rules for that would disagree at the boundary. The pack is the only manifest-scoped request a refused write can be compared against, and the docstring says why each of the others cannot be. **[P1] Every recovery command was built from the wrong list.** `_requested_setup_flags` answers what a rerun in a *different* workspace may repeat; every recovery `init` publishes reruns the same one. So `init --write --minimal --control-pack ` emitted a recovery with no `--minimal`, and following it wrote a detected manifest where the legacy template was explicitly asked for; an explicit kit became the default kit, and a raised `--max-python-files` was dropped so the rerun refused again on a truncation the caller had already answered. One `_invocation_flags` builder now feeds all five sites, and correcting a value means passing the corrected value rather than dropping the flag. **[P2] `input_id` did not cover every fact that selects the route.** `action_kind` is not carried by the action and decides `next_action.kind` and `verify_required`; diagnostics were reduced to their rank-1 actions, dropping the id that decides ownership and the severity that decides precedence. Both are hashed now, with the same fix applied to `detect`'s and `init`'s own `advance_kind`. `doctor`'s success identity is deliberately untouched: `STABILITY.md` promises it is the same on every machine. **[P1] The published `execution` guarantee was false.** `STABILITY.md` said a setup error envelope always reports `"failed"`. The most common one does not — `init --write` over an existing manifest reached an answer and refused, so it exits 2 with `"succeeded"`, and that line is the resume of the adoption walk itself. A consumer taking the contract literally would have rejected it. The contract now states what actually holds on every setup error line (`decision_source: "setup"`, a setup-vocabulary `decision`, `permissions` all false, never `complete`) and describes `execution` as what it is: whether the command reached an answer, not whether the line is an error. Corrected in `STABILITY.md`, `docs/agent-contract-current.md`, `AGENTS.md`, and `docs/errors.json`, and pinned by a test that asserts both shapes occur. Also from the review: `docs/passed-verdict-contract.md` still stamped contract v26, and the two command-value migrations — the recovery flags and the `internal_error` fallback — are now stated explicitly rather than left under "nothing is removed". **Coverage is of the routes, not their shapes.** The walk fixture now runs the emitted recovery and inspects the manifest it produced, with the no-`--minimal` case as the control so the assertion cannot pass vacuously; it follows the unapplied-pack edit and requires the obligation to be gone afterwards. The `advance_kind` identity guard is structural and says why: no branch varies the kind without varying the action today, so a behavioural test would pass for a reason unrelated to the fix. **And one defect found reviewing the fix.** `_unapplied_control_pack`'s docstring claimed it guarded a `None` on-disk pack; the code did not, so a manifest that loads while its pack does not resolve would have published an edit whose `why` read "the manifest still selects None". Today no such manifest is reachable — an unknown pack id fails schema validation and the defect route claims it first — but that is `_manifest_defect` and `_manifest_control_pack` happening to agree, and they run different amounts of the loader. The guard is what makes a disagreement fall through to `doctor` instead of becoming a route. Co-Authored-By: Claude Opus 5 * fix(agent-mode): the direction owns the pack change, not the caller (#323 review 2) Seven findings from executing the routes the last round emitted. Three are the same defect the round before them was about, one directory down or one flag over; one is a security boundary the last round got wrong. **[P1] `--control-pack` on argv is not a human approval.** The reconciliation route was published as a coding-agent `edit` on the grounds that the value came from the command line rather than from inference. That is exactly backwards for the case that matters: a *governed* coding agent composes its own argv, so `init --write --control-pack read-only-agent` over a `financial-strict` manifest is a policy weakening an agent can request for itself — it drops `write` and `production_operation` obligations. The **direction** decides the owner now. A transition that keeps at least every obligation the manifest has today is an agent `edit`; one that drops any, or that names a pack this build cannot resolve, is `human_review_required` with no command, naming the effects and controls it would remove. `_weakened_pack_rules` moved to `core.control_packs.weakened_pack_obligations` beside the packs it compares, and `verify_policy` delegates: two implementations of "did this get weaker?" is how one of them stops seeing a downgrade (#410 §F). **[P1] The same-workspace capped retry dropped the invocation.** It is the one scope route that reruns *this* workspace, and it was built from the list a rerun in a **different** one may repeat — so `init --write --allow-unresolved-scope --max-python-files 1` emitted a retry without `--allow-unresolved-scope`, and following it returned the refusal it was issued to resolve, writing nothing. There are two lists now with one derivation: `_portable_setup_flags` is what transfers to another workspace, and `_invocation_flags` is that plus the two that only make sense in this one. The retry replaces the cap and nothing else. **[P2] Candidate commands dropped a raised parse cap.** It bounds how much is read, not which directory, so it transfers: without it a candidate inherited the root's truncation, refused again, and wrote no manifest while its `expects` promised one. It is in the portable list. **[P1] An adopted candidate lost the request one directory down.** `scope_candidate_actions` routes an adopted candidate to `doctor`, which reads a manifest that loads fine and advances to the gate under the pack that is *there*. Candidate routing now reconciles first, through the same shared route and the same ownership rule. **[P2] An explicit `--control-pack default` is a request.** Inferring "was this asked for?" by comparing against the default cannot see the one transition that can *only* weaken an existing manifest. Read from the parser instead — and note where the trap is: **typer vendors its own click**, so the returned `ParameterSource` is not the `click.core` enum and `is`/`==` against it are both false, silently, in the direction where every option reads as unasked. Compared by member name, with that written down. **[P2] The changelog contradicted itself** — one paragraph promised `execution: "failed"` on every failure envelope while another correctly documented the succeeded-refusal shape. It states only the schema-enforced invariants now. **[P3] A Windows-only test failure** — a hardcoded POSIX kit path in the parametrized argv expectation, derived from the `Path` now. Also: `docs/errors.json`'s `config_already_exists` hint still told consumers to hand-edit or delete the manifest, which is no longer the route that line carries; it describes the conditional routes and says why deleting is a person's decision rather than a recovery. Rebased onto main. **And one found while fixing.** The weakening route's prose ran past the 400-byte envelope cap, and the clause that got truncated was the one saying why a person owns the decision — leaving a route that read like a formatting quirk. The reason leads now and the effect list is fitted to what is left, the shape the placeholder review already uses. **Coverage executes the emitted commands**, as asked: the capped retry is followed and the manifest it promised has to exist; candidate commands are checked for the bound the root run was given; the adopted-candidate route is required not to be a `doctor` handoff while a request is outstanding; and the explicit-default case is paired with the omitted-flag control so neither can pass vacuously. The strengthening direction is pinned beside the weakening one for the same reason. test(samples): a cold-start sample whose declaration questionnaire is committed (#425) (#458) * test(samples): a cold-start sample whose declaration questionnaire is committed (#425) The declaration questionnaire is the primary cold-start surface -- what an adopter at `insufficient_evidence` reads, in order, to reach a verdict -- and no shipped sample exercised it. Every sample answered every question it was asked, so `open_questions` was `[]` in all five goldens, `suggested-declarations.yaml` had no golden at all, and the progress sentence was only ever rendered at zero. That is how #419 shipped an ordering that put a structurally proven OpenAPI `GET` named `delete_account` at the top of the file ahead of a genuinely unknown effect, against a fully green suite. `samples/google_adk_cold_start_agent` is the sibling of `samples/google_adk_agent` that deliberately stops partway: ten open questions across the rungs the ordering distinguishes. - `update_case_index`, `assemble_case_timeline`, `list_case_attachments` -- nothing read at all, ordered against alphabetical order in both directions by the name band alone. - `ops.queue_backfill` -- only the MCP protocol default stands in for evidence. - `ops.export_case_bundle` -- a declared authority that contradicts the MCP export, so the counter counts a question the questionnaire prints a note for rather than a block, in its numbered place between two blocks. - `issue_goodwill_refund` -- read as a financial write, so its block arrives with a proposed answer. - `support.get_update_history` -- a structurally proven read whose name bands as a write, which is #419's own defect inverted: it ranks last, and ranks first again the moment `_reach` goes back to asking whether a side effect was measured rather than bounded. - `ops.append_case_note` is never asked about and `record_case_outcome` is answered, which is what makes the counter read `2 of 12 answered; 10 open`. `expected/suggested-declarations.yaml` is byte-compared the way `report.md` already was; `open_questions[].*` now reaches `test_sample_expected_report_json_has_no_structural_drift`; and the gap order the decision reason and `first_recommended_action` project is asserted against the questionnaire's own numbering. Verified by perturbation: ranking `_reach` by `conservative_effect`, ranking it by whether a side effect was measured, and flattening `name_shape_band` each fail two committed-artifact tests. Adding a field to `DeclarationQuestionRow` fails the structural-drift test for this sample and no other. Test evidence only -- no CLI command, schema version, report block, discovery surface, or adapter changes. Co-Authored-By: Claude Opus 5 * fix(samples): address review of the cold-start questionnaire fixture (#425) Seven findings from a review pass over the branch. Six were claims that had gone stale as the fixture grew, each of which would send a reader to the wrong place: - the manifest header said "nothing bounds the first four" while listing five unbounded actions, so a reader counting them off would conclude `list_case_attachments` is one the scan read for itself -- the opposite of why it is in the fixture; - the sample README's "the first four are what pin the band" named table rows 1-4, when the band is pinned by rows 1, 2 and 5; - both files described `samples/google_adk_agent`'s questionnaire as `0 of 0`. It is `2 of 2 answered`: what distinguishes the two fixtures is that nothing is open, not that nothing was asked; - a test docstring still said the fixture renders six blocks, which was true before it grew to ten questions. Replaced with a pointer to `COLD_START_QUESTION_ORDER`, which cannot go stale without failing; - the rung guard indexed evidence gaps as one row per (subject, dimension). Five kinds map to the `effect` dimension, so an action carrying two of them would keep whichever was emitted last and evaluate every assertion against a row nobody named. Indexed to lists instead. The seventh was in the shipped artifact. The README's regeneration recipe scanned into a temporary directory and copied the files back, so the committed `report.json` carried `"generated_reports": {"markdown": "/private/var/folders/.../tmp.../report.md"}` -- the generating machine's temp path, in a file meant to be stable. The `` rewrite only replaces the working directory, so it never reached it, and no test compared the value, which is exactly why it would have sat there churning on every regeneration. Fixed at the source: the recipe now uses the relative `output_dir="expected"` that `run_scan` resolves under the manifest dir, matching the other five samples, and deletes the `current-control.json` a scan also publishes. And closed as a class rather than an instance -- `test_sample_expected_report_json_uses_repo_placeholder_for_manifest_dir` now rejects an absolute `generated_reports` path in any sample golden, verified with a negative control. Co-Authored-By: Claude Opus 5 * test(reports): the absolute-path guard must not depend on the platform it runs on (#425 review) `Path(written).is_absolute()` is the *running* platform's `Path`, and the leak this guard exists to catch is a POSIX temp path: `WindowsPath('/private/var/folders/…/report.md').is_absolute()` is False for want of a drive. So on Windows the check would have passed on exactly the golden it was added to reject, while the same golden failed on Linux -- a portability check that is itself unportable. Both spellings now, via `PurePosixPath` and `PureWindowsPath`. Verified with a negative control for each: a `/private/var/...` value and a `C:\Users\...` one both fail the assertion, and a relative `expected/report.md` passes. `.github/workflows/ci.yml` runs a `windows-latest` job that does not invoke pytest today, so this was latent rather than live -- but a contributor running the suite there got a silently weaker check than the failure message promises. Co-Authored-By: Claude Opus 5 * docs(samples): name what could move the fixture's verdict, and fail loudly when the scaffold is gone (#425 review) Two findings from a third pass. The fixture's `insufficient_evidence` verdict is what two tests pin, and it is not reached by the evidence gaps alone: a blocker, or an active high/critical review concern on a proven-reachable capability, both outrank it. So the verdict also depends on this sample's one finding -- `SHIP-ADK-EVAL-COVERAGE-MISSING`, currently `medium` -- staying below that tier, and nothing said so. Raising that check would flip this sample to `review_required` and leave whoever did it reading a decision-mismatch message that names neither the check nor the precedence rule. Documented rather than silenced with an eval file. A repository that has not declared its eval coverage is exactly what a cold start looks like, and shaping the fixture around its own test would hide a real part of the report. Second, `test_cold_start_scaffold_matches_its_golden` read the fresh `suggested-declarations.yaml` unconditionally. `_write_suggested_declarations` *deletes* that file rather than writing one when no gap carries a template, so the likeliest regression the test guards against surfaced as a bare `FileNotFoundError` instead of the message it carefully spells out. Now named before it is read, and verified with a negative control that forces `scaffold_for_report` to return `None`. Co-Authored-By: Claude Opus 5 * docs(samples): name both tests the fixture's verdict is pinned by (#425 review) The paragraph said `insufficient_evidence` is "what two tests pin" and then named one. The other is `test_cold_start_markdown_report_matches_golden`, whose golden carries `Decision: insufficient_evidence` as line 9 — so a severity bump that flips this sample reads there as unexplained golden drift, and invites a regeneration that quietly changes what the fixture claims. Co-Authored-By: Claude Opus 5 * test(samples): address the three reproducibility and test-contract gaps (#425 review) **The regeneration recipe was POSIX-only.** `text.replace(os.getcwd(), "")` works where the separator needs no escaping and is a *no-op* on Windows: `json.dumps` writes `C:\\repo\\samples\\…` while `os.getcwd()` returns `C:\repo\samples\…`, so the two never match, the golden keeps an absolute `manifest_dir`, and the machine that produced it fails `test_sample_expected_report_json_uses_repo_placeholder_for_manifest_dir`. `generated_reports` had the same problem one level down: `expected\report.json` is relative, passes every check, and differs from the committed bytes everywhere else. The recipe now normalizes structurally — `payload["manifest_dir"] = f"/{sample.as_posix()}"` and `Path(written).as_posix()` — and serializes with `json.dumps(payload, indent=2)`, which is exactly what `write_json_report` uses (no `sort_keys`, no trailing newline), so the round trip is byte-identical and nothing but the two normalized fields moves. Verified: the new recipe reproduces the committed `report.json` byte for byte. The guard grew the matching half — no backslash in `manifest_dir` or in any `generated_reports` value, in any sample golden. **The cross-surface test never scanned.** It took `tmp_path` and then read only the committed `report.json`, so it proved the golden agrees with itself rather than that the producer still projects the leading question into `agent_summary.first_recommended_action` — a field `report.md` does not render, the drift test compares by path only, and the scaffold does not carry. It now asserts gap order, reason and first action on a **live** scan, and separately that the committed rows carry the same order. Demonstrated: making `primary_evidence_gap` return the last addressable row instead of the first — which moves nothing but the reason and that field — fails this test and no other cold-start check. **A gap row is not a question.** The comparison had one entry per `EvidenceGap` against one per blank, so a fixture where two rows ask one question — a `declaration_drift` beside a `declaration_below_inferred_evidence` on one effect — would fail it while being correctly ordered, contradicting the list-valued indexing this same file already uses. Contiguous equal keys now collapse, which is sound because `_in_question_order` sorts by `(question, original position)`; a key that recurs after another intervened survives and fails the comparison, because that is the permutation broken. Both halves are pinned directly, since this fixture raises one row per blank and could not show either. Co-Authored-By: Claude Opus 5 * test(samples): force LF on regenerated goldens, and count questions rather than mapped kinds (#425 review 2) **Windows regeneration wrote CRLF, and nothing could see it.** `Path.write_text(..., encoding="utf-8")` opens in text mode with `newline=None`, so on Windows every `\n` becomes `\r\n` — from the recipe, and from `write_json_report`, the Markdown writer and the questionnaire writer alike. `.gitattributes` pins `samples/**/expected/** -text` so Git stores those bytes verbatim, and the golden is committed with different bytes from everyone else's. No existing test could catch it: `read_text()` applies universal newlines and normalizes CRLF back to LF on the way in, so every byte comparison in this repo passes on a file whose committed bytes moved. Confirmed rather than assumed — with a CRLF `report.md`, `test_cold_start_markdown_report_matches_golden` passes. The recipe now forces `newline="\n"` on all three goldens, and `test_sample_expected_goldens_are_committed_with_lf_newlines` reads the **raw bytes** of every `samples/*/expected/*` file so the guard does not share the blindness. All committed goldens are LF today; the guard keeps them that way. **A mapped kind is not a question.** `DIMENSION_BY_GAP_KIND` answers which dimension a kind rides on, not whether a declaration can close it: a `conflicting_effect_evidence` the resolver blames on the *source* maps to `effect` and is deliberately excluded by `is_declaration_answerable`, so the oracle invented a key `open_questions` never carries and would have failed a correct report. `_in_question_order` avoids this by filtering its permutation through the open-question rank rather than the kind map, and this now filters the same way. Both reductions are pinned directly in `test_gap_rows_reduce_to_the_questions_they_ask`, including the branch that tells the two `conflicting_effect_evidence` cases apart, because this fixture raises one answerable row per blank and no source-owned conflict — so it could not exercise either. Verified non-vacuous: dropping either reduction fails that test. Also merges `origin/main`, resolving the `CHANGELOG.md` conflict with both Unreleased entries preserved (#424 and #341 alongside #425). The merged semantics leave this fixture's goldens byte-identical. feat(release): approve and enforce an explicit pre-1.0 evidence bar (#341) (#455) * feat(release): approve and enforce an explicit pre-1.0 evidence bar (#341) The repository had exactly one release policy — 100 adjudicated, receipt-bound cases, written as the claim 1.0 should make — enforced for every tag. No corpus met it, so nothing published and evaluators kept installing v0.15.0: an older, less-verified build than the one being withheld. Route 2 of #341, recorded by Pengfei Hu (product/security) on 2026-08-29 in docs/release-evidence-policy-decision.md. A second named policy, `pre_1_0`, governs 0.x tags: 56 cases, two in each of the same 28 profile x decision strata, origin floor at the same 40% share. Everything that decides whether evidence is believed is byte-identical to production — zero unsafe auto-passes per profile and overall, a unique terminal verifier receipt per case, the 20% holdout fraction, the kappa floor, static_only, and full re-derivation by the verifier. Every exact-match floor is the production rate rounded up, so three of four land on 100% at this size: a smaller corpus buys less tolerance for error, not more. The version decides the policy and the artifact never does. Epoch 0 with major 0 admits `pre_1_0` or the stronger `beta`; anything else — including a version that will not parse — admits `beta` only. Both gates derive this independently, and an artifact naming a tier its version does not admit is rejected *and* then measured against the production counts, so a bad tier cannot shrink what is checked. `production_qualified` keeps meaning "met the 100-case bar". The sealing gate was the sixth definition site, not the fifth: the standard-library verify_qualification_binding.py hard-coded REQUIRED_CASE_COUNT = 100 and tier == "beta" of its own, so a change that moved the documented five would have failed at the last step before publication with a bare case-count error. It is per-tier now, bound to the real constructors by test. The exhaustive verifier stopped hard-coding 100 in six places and derives every count, metric denominator and confusion-matrix profile from the governing policy. Co-Authored-By: Claude Opus 5 * fix(release): close five review findings on the pre-1.0 evidence bar (#341) Self-review of the PR, plus a perturbation sweep over the new guards. - The corpus owner's runbook (benchmark/safety-qualification/README.md) was a *seventh* definition site, and the one that decides what actually gets built. It documented only the 100-case policy and claimed the CLI has no policy selection. Whoever took the corpus-delivery issue would have built the wrong artifact, which no gate can detect and no verifier error would explain. It now documents both policies, both exit-code meanings, and what --policy-tier will and will not do. Its pre-existing claim that receipts must carry report schema 0.40 (the gate pins 0.42) is corrected in passing. - The `requirements=` keyword bypassed the version rule entirely: a caller could score a 1.0 wheel against the pre-1.0 policy and exit 0, with the mismatch found only after signing. `require_tier_governs_version` is now applied to the *computed* tier, so it binds however the requirements were chosen, and `select_release_requirements` routes through the same rule. - docs/INDEX.md still described the decision as open with no route selected, contradicting the document it links to. - The sealer's production-count fallback for an unadmitted tier was untested; the test now asserts both the tier error and the 100-case shortfall. - A perturbation sweep found the one gap the suite could not see: widening the runner's `production_qualified` to any named tier left every test green, because no test produced a *passing* named-policy artifact. Two fixes. `SafetyQualificationResultV1` now refuses to construct an artifact whose `production_qualified` disagrees with `qualified` and the tier, so the producer cannot emit the inconsistency at all — and the now-unreachable restatement of that check is removed from the exhaustive verifier, leaving it in the standard-library sealer where it still parses raw JSON. And a conforming 56-case corpus fixture proves end to end that the approved bar is satisfiable, exits 0, and reports production_qualified false. Co-Authored-By: Claude Opus 5 * docs(release): link the corpus-delivery issue and correct the definition-site framing (#341) - #456 is the corpus-delivery follow-up; the decision doc's last-but-one acceptance box is closed against it. - The "Where the bar is defined" preamble claimed all the sites are cross-checked. Five are. The corpus runbook and the docs index are not, which makes them the more dangerous omission: nothing fails, and the corpus simply gets built to the wrong shape. Co-Authored-By: Claude Opus 5 * fix(release): address review — sealer policy, PEP 440 parsing, envelope v5 (#341) Three blocking findings from review, all reproduced before fixing. **The stdlib sealer restated a case count, not a policy.** It checked count, receipts, runtime failures and unsafe auto-passes, but never the strata or the exact-match floors — so a schema-valid 56-case `pre_1_0` artifact with 12/14 safe passes passed the sealing gate while the exhaustive gate rejected it, and 56 identical rows in no stratum at all passed too. That made the dependency-compromise boundary this file exists to hold decorative. The restated table is now a full `QualificationPolicy` per tier — strata, per-outcome exact-match floors, per-stratum holdout, origin and kappa floors — re-derived from the raw cases, with every field bound to the real constructors by `test_the_stdlib_policy_table_matches_the_named_policies`. The two pre-existing sealer tests used 100 identical rows and now use conforming corpora, which is what makes them meaningful. **The version rule was anchored only at the start.** `0garbage`, `0.16.0garbage`, `0..1`, `0-` and `0x1` all parsed as major 0 and bought the *cheaper* policy — the exact inversion of the documented fallback, in a helper both gates share, so it opened both at once. It now requires a complete PEP 440 parse; leading zeros still normalize (`00.1` is pre-1.0, `01.0.0` is not). **New grammar under a frozen envelope id.** `qualification_tier: pre_1_0` and the production_qualified/tier invariant are both grammar changes a v4 reader rejects, so a genuine pre-1.0 artifact must not claim v4. `shipgate.safety_qualification` advances v4 -> v5. v4 stays readable — its vocabulary is a strict subset of v5's and its producer always satisfied the new invariant — and the corpus and receipt-index envelopes deliberately do not move, because their grammar did not change. STABILITY.md carries the migration note. Seven perturbations of the new guards were run; all seven fail the suite. Co-Authored-By: Claude Opus 5 * fix(release): close the second review round on received-artifact validation (#341) Rebased onto main (CHANGELOG conflict resolved, keeping both Unreleased entries). All five findings reproduced before fixing. **[P1] The envelope bump had a hole, in the direction it existed to close.** Upgrading a legacy envelope was a bare `schema_version` field validator, so it rewrote v4 to v5 *before* anything could look at `qualification_tier`: relabelling a conforming pre-1.0 artifact as v4 was accepted by the model, by the exhaustive gate, and by the sealer (which never read the envelope at all). My previous reply claimed "upgrading grants nothing" — that was wrong on exactly this point. A legacy envelope is now read only when the payload uses the vocabulary that envelope can express, as a model-level before-validator, and the sealer enforces the same pairing on raw JSON. **[P1] The sealer accepted non-terminal and unidentified evidence.** The floors count *matches*, so a case with a null `actual_decision` merely failed to count toward its floor — 13 of 14 safe matches, enough — and 56 rows sharing one id looked like 56 cases as long as their receipt digests differed. Case identity (unique, non-blank, string) and terminal expected/actual decisions are now required before any floor is applied. **[P1] Non-finite kappa satisfied the floor.** The JSON literal `1e309` loads as `inf`, which passes any `>=` bound while the exhaustive gate rejects it against `<= 1.0`. Kappa must now be finite and within `[floor, 1.0]`, and `qualified_origin_cases` a genuine integer within `[minimum, case_count]`. **[P1] The report schema requirement was not bound in the dependency-free gate.** It has no representation in `cases`, so `0.42` could be restated as `0.1` and still seal. `QualificationPolicy` now carries it and emits `as_requirements_payload()`; the sealer compares the artifact's whole declared `requirements` block field-for-field, and the binding test asserts that payload equals `SafetyQualificationRequirementsV1.model_dump()` — so a field added to the model breaks the build until the stdlib copy restates it. **[P2] The 1-tuning/1-holdout split was documentation, not policy** — and should stay that way. Enforcing a minimum tuning count is a *maximum* on holdout, and holdout evidence is evidence the engine was never tuned on, so an all-holdout corpus is stronger, not weaker; a gate must not reject a corpus for being more conservative than required. The decision document, schema docstring and corpus runbook now state the holdout floor that is actually enforced and why there is deliberately no tuning floor, with the behaviour pinned by `test_a_corpus_with_more_holdout_than_required_is_accepted`. Seven perturbations of the new guards; all seven fail the suite. fix(semantics): a reviewed risk tag refines the row it sits in, it does not contradict it (#424) (#457) * fix(semantics): a reviewed risk tag refines the row it sits in, it does not contradict it (#424) `declaration_below_inferred_evidence` publishes two routes, and the second is the one the row names: "add `action_surface.actions[].risk_tags: [X]` so the X controls apply to this action". Applying it exactly as instructed replaced the review-tier row with a blocking `conflicting_effect_evidence` whose message blamed the reviewer's own manifest — a published next step that could not close the row it was printed on. Measured over every declared effect against every one-, two-, and three-observation combination: 281 of the 390 pairs that take the tag route. Two branches of `_assess_effect` disagreed about what a reviewed `risk_tags` entry is. `claims_above_declared_effect` treats it as covering, deliberately — a declared tag is policy-eligible, so it both accounts for the observation and applies that category's built-in controls. The `contradictory` filter ran first and read the same claim as source evidence outranking the declaration. It now excludes the two spellings of "declare this category as reviewed": `action_risk_tag_declaration` and `risk_hint:manual`, named once as `REVIEWED_RISK_TAG_CLAIM_SOURCES` and shared with the one other site that had the pair spelled inline. Same class already fixed one branch over in `_source_read_conflict`. Deliberately narrower than `_is_manifest_owned`: that predicate also covers `action_scope`, and #417 made a declared `crm.delete` grant bound the action's effect. A tag refines the effect the same person wrote; a grant asserts an independent fact that bounds it. Protocol annotations, a source's own scopes, and typed provider facts keep contradicting a weaker declaration. The gate verdict this moves: `effect: read` + `risk_tags: [destructive]` goes from blocking to pass-eligible, and the two spellings now reach identical gate answers. The P0 canary pinning the old outcome is replaced in its slot by the boundary the fix draws, `declared_delete_scope_cannot_downgrade_to_read`, and its property is asserted in its new form beside it. Guards: the repair sweep now walks 1-3 observations, requires both published routes, and requires the whole effect dimension clean rather than the absence of one kind; the questionnaire round-trip walks both routes per kind; and a scan-level test pastes the published template into the manifest and rescans. Closes #424 Co-Authored-By: Claude Opus 5 * fix(semantics): one predicate for "the manifest wrote this", at all three sites (#424 review) Reviewing the #424 fix found the same missing second route at two more comparisons. The manifest reaches the effect dimension by two routes — the action row (`DECLARATION_CLAIM_SOURCES`) and `risk_overrides.tags`, which arrives as `risk_hint:manual` carrying a `reviewed_declaration` basis and no declaration source. Both comparisons excluded the first only. * `SHIP-ACTION-EFFECT-DOWNGRADE-DECLARED` derives "the effect Shipgate inferred" by excluding the manifest's own claims. A reviewed `risk_overrides.tags: [destructive]` beside a source that says only `write` came back as `inferred_effect: destructive`, in a recommendation telling the reviewer to declare the value they had already written. * `declaration_below_inferred_evidence` names the evidence that agrees with the declaration, in a sentence reading "source evidence agrees with the declaration (risk_hint:manual)" — the manifest confirming itself, which the comment directly above that filter already forbade. `_is_manifest_owned` is promoted to `is_manifest_owned_effect_claim` and both sites ask it, so a fourth site cannot spell it a fifth way. Each fix has a guard, and each guard was confirmed to fail with only that fix reverted. Co-Authored-By: Claude Opus 5 * test(semantics): pin the trust premise the reviewed-tag exclusion rests on (#424 review) Excluding `risk_hint:manual` from the source-evidence comparison is only safe because tool-published content cannot produce one, and two things make that true: `_validated_hint_basis` grants the `reviewed_declaration` basis to no other hint source, and no adapter writes that basis or a `"manual"` source. Both are one edit away from not being true, and neither was asserted — the same argument `_source_read_conflict` has relied on unstated since it was written. Co-Authored-By: Claude Opus 5 * style(semantics): reflow the #424 comment block Co-Authored-By: Claude Opus 5 * docs(tests): say what the guards assert, in the words they assert it in Co-Authored-By: Claude Opus 5 * fix(semantics): the tag repair writes the whole risk_tags value, and a read claim stops synthesizing read_only (#424 review 2) Three findings from review, all reproduced first. 1. The published tag repair still could not close its row on an already-tagged action. `risk_tags` is one YAML key, so a block naming it replaces it, and the repair named only the newly uncovered category: a row reading `risk_tags: [financial_write]` was asked to delete the tag covering the financial reading, and the next scan reopened the row asking for it back — the #424 defect surviving in the case #424's own repair creates. `EffectRepair.risk_tags` is now the complete value to write and `added_risk_tags` carries what changed, so the sentence and the value cannot disagree. The declared list is taken verbatim from a new `declared_risk_tags` entry on the declared-effect claim; it cannot be rebuilt from the tag claims, because `read_only`, `network_access` and `customer_data` map to no positive effect and produce none — the same lossiness `declaration_drift` already refuses to guess through. 2. The new pass-eligible spelling published a false fact. `build_action` unions a risk tag for every policy-eligible effect claim, and `read` is not a category — it is the assertion that none apply. `effect: read` beside `risk_tags: [financial_write]` therefore synthesized `read_only` on a `financial_write` action, and `derive_side_effect` reads that tag as positive evidence: `reversibility: reversible`, which that helper's own docstring says a declared read must not buy. This is #461, folded in because this branch is what moves the wrong fact into a passing report. 3. The contract overstated the equivalence. A tag *adds* its category to the declared effect; it does not replace it. `effect: external_communication` with a financial tag owes confirmation as well as approval, audit and idempotency, where `effect: financial_write` alone does not — so the two edits are not interchangeable, and the docs now say which is which. Guards: the end-to-end paste test is parametrized over a row that already carries a covering tag and pastes the block as a merge instruction replaces; the equivalence guard compares whole action and capability facts (equal facts are equal against every base) and asserts the claim lists still differ, so the audit trail is not flattened; and a new negative control pins the union of obligations. Each was confirmed to fail with only its own fix reverted. Co-Authored-By: Claude Opus 5 * test(semantics): sweep already-tagged rows through the repair, exhaustively (#424 review 2) The paste test covers one already-tagged fixture. This walks every declared effect against every one-, two-, and three-observation combination twice — bare, and with a category the tool really reads already declared as a reviewed tag — and applies the repair by *replacing* `risk_tags`, which is what pasting a block that names the key does. Reverting the union fails it on 50+ pairs. build(deps): bump uv to 0.12.5 without rewriting the publication lock (#406) (#449) Dependabot's #406 carried the right one-line bump but regenerated `constraints/release-publish.txt` with a different resolver, and the collateral was the part that failed CI: * `pydantic==2.13.4` was rewritten as `pydantic[email]==2.13.4`, which `scripts/verify_dependency_lock.py` rejects — its `_PIN` regex names `[A-Za-z0-9][A-Za-z0-9._-]*`, so it cannot parse an extras-qualified pin and reports a legitimate exact pin as "neither an exact pin nor a direct URL". That was the actual red build. * Two PyPy environment markers were silently dropped, turning `cffi==2.1.1 ; platform_python_implementation != 'PyPy'` and the matching `pycparser` line into unconditional installs. Nobody asked for that, and it would have shipped unreviewed behind a uv bump. Bumps `release-publish.in` and recompiles only that one lock through scripts/update_locks.py. Both markers survive and pydantic keeps its plain spelling. Compiled with uv 0.12.5 rather than the outgoing 0.11.7, so the lock is produced by the resolver version it pins and the next recompile reproduces it. `update_locks.py` warns during the transition because it reads the pinned version from the lock it is about to replace; the warning clears once this lands. Note: recompiling re-resolves the whole sigstore closure, so this also carries nine unrelated upgrades (cryptography, idna, pydantic, pydantic-core, pygments, securesystemslib, platformdirs, charset-normalizer, typing-inspection). `uv pip compile` writes to stdout and never reads the existing lock, so there is no per-package upgrade mode; hand-pinning them back would produce a lock no recompile reproduces. build(deps): stop Dependabot raising transitive pins past their parent's cap (#448) `constraints/*.txt` are hash-pinned closures whose lines are mostly transitive, recorded with a `# via ` comment. Dependabot treats each line as independently bumpable, so it raises pins above the cap of the parent that pulls them in. The result fails with ResolutionImpossible at install time, in roughly 20 seconds, before a single test runs — and no rebase can fix it, because no release of the parent accepts the new version. Two are proven impossible and were closed rather than repaired: * #407 bumped chardet 5.2.0 -> 7.6.0. chardet is `# via cyclonedx-bom`, which requires chardet<6.0,>=5.1; cyclonedx-bom 7.3.1 is the latest release overall, not just the latest 7.x. * #382 bumped pydantic-core 2.46.4 -> 2.48.0. pydantic pins it exactly, not as a range: 2.13.4 needs ==2.46.4 and 2.13.5 needs ==2.46.5. Without an `ignore:` rule both regenerate every week. chardet is ignored for major updates only, so 5.x movement still comes through and the bump reopens on its own if cyclonedx-bom lifts the cap. pydantic-core is ignored outright, since an exact parent pin leaves no update type that could ever resolve. build(deps): complete the hatchling 1.32.0 bump (#381) (#447) Dependabot's #381 moved the version string and left the lock behind, in three separate ways: * `constraints/build-backend.txt` pinned `hatchling==1.32.0` while keeping 1.31.0's two `--hash` digests, so every hash-checked install failed with THESE PACKAGES DO NOT MATCH THE HASHES. The digest pip actually got was the genuine 1.32.0 wheel from PyPI. * `constraints/release-seal.txt` was not touched at all, leaving the sealing job pinned to 1.31.0 while `release-seal.in` declared 1.32.0. * hatchling 1.32.0 adds a dependency 1.31.0 did not have — `tomlkit` — which was locked nowhere. Bumps the three declaration sources (pyproject.toml's build-system floor, constraints/release-build.txt, constraints/release-seal.in) and recompiles only the two locks hatchling feeds, via scripts/update_locks.py, so the change does not drag constraints/dev.txt or release-publish.txt forward with it. Note: recompiling `release-seal.txt` re-resolves that whole closure, so it also carries nine unrelated upgrades (cryptography, idna, pydantic, pygments, securesystemslib, platformdirs, charset-normalizer, typing-inspection, pydantic-core). `uv pip compile` writes to stdout and never reads the existing lock, so there is no per-package upgrade mode to avoid this; hand-pinning them back would produce a lock no recompile reproduces. Build(deps-dev): Bump idna from 3.18 to 3.19 (#405) Bumps [idna](https://github.com/kjd/idna) from 3.18 to 3.19. - [Release notes](https://github.com/kjd/idna/releases) - [Changelog](https://github.com/kjd/idna/blob/master/HISTORY.md) - [Commits](https://github.com/kjd/idna/compare/v3.18...v3.19) --- updated-dependencies: - dependency-name: idna dependency-version: '3.19' dependency-type: direct:development update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> fix(samples): the shipped control pointer hashes the bytes committed beside it (#446) * fix(samples): the shipped control pointer hashes the bytes committed beside it `samples/conductor_agent/expected/current-control.json` bound a `report.json` digest and `size_bytes` that no committed file could ever match, so `read_current_control()` refused Shipgate's own shipped sample with `artifact_mismatch` -- the product rejecting its own evidence set. The cause is an ordering bug in the golden recipe, not in the producer. A regeneration runs `run_scan(...)`, which writes `report.json` carrying the absolute `manifest_dir` and hashes exactly those bytes; only afterwards does the recipe substitute the generating checkout's path with ``. The committed digest therefore describes a pre-substitution file that is never committed, and `size_bytes` moves with the contributor's path length -- which is why the value churned across releases without ever being right. `report_markdown` was correct throughout, because nothing rewrites `report.md` after it is hashed. A half-correct artifact is worse than a wholly broken one: the accurate half is what made the broken half easy to trust. The pointer is rebound through the product's own model and `content_id(current_control_identity_payload(...))` rather than hand-edited, because `current_control_id` covers the artifact refs -- a patched digest without a recomputed identity fails `model_validate` outright. Serializing as `_publish` does (`indent=2`, `sort_keys=True`, trailing newline) keeps the diff to the three lines that actually changed. The guard is what keeps the class closed, since nothing read this file before: `test_sample_current_control_pointers_bind_the_committed_artifacts` walks every `samples/*/expected/current-control.json` and confirms each `artifacts.*` entry against the file it names. It hashes with plain `hashlib`, deliberately not `bind_current_control_artifacts` -- an oracle that reuses the producer passes vacuously when the producer is wrong -- and asserts the glob is non-empty, so moving the sample cannot make the invariant silently disappear, which is how it rotted in the first place. Every branch was shown to fail on a perturbation: the pre-fix pointer, drifted artifact bytes, a bound file that is not committed, a hand-edited digest (caught by the model validator), the `report_markdown` half, and an empty glob. The ordering constraint is recorded in `samples/README.md`, since the recipe otherwise lives only in team notes and no script enforces it. Co-Authored-By: Claude Opus 5 * fix(samples): a hash-bound golden must survive checkout, symlinks, and removal Addresses all three review comments on #446. Each was reproduced at 3b44c462 before being fixed, and each is a way the new guard could stay green while `read_current_control()` refuses the shipped sample -- the exact failure the branch exists to close, left open one level down. [P2] Nothing pinned the line endings of files whose digests are recorded. `git check-attr text eol` reported both `report.json` and `report.md` as unspecified, and the repository had no `.gitattributes` at all. Cloning this branch with `core.autocrlf=true` -- the Git for Windows default, and this repository ships a `windows-launcher` job -- yields a 52,236-byte `report.json` hashing `9ddd4fc0...c763` against a pointer that records 50,809 and `cd87ed0b...d024`; `report.md` moves too. The guard fails and the reader raises `artifact_mismatch`, so the previous commit would have handed Windows contributors a red suite. Normalizing inside the test was rejected: the reader hashes real bytes, so that would hide the breakage rather than fix it. `samples/**/expected/** -text` makes Git hand these files over exactly as committed. The committed blobs contain no CR bytes, so nothing is rewritten. [P2] The per-artifact loop was vacuous on an empty `artifacts` map, and that map is schema-valid for a `scan` pointer -- `CurrentControlPointer` requires a binding only for `verify` and `preview`. Stripping the refs and recomputing `current_control_id` produced a pointer that validated and left the guard with nothing to assert, permitting exactly the alternative the PR argued against. `EXPECTED_CURRENT_CONTROL_ARTIFACTS` now names each fixture's keys, so both dropping a binding and adding an undeclared fixture are failures. [P2] `is_file()` and `read_bytes()` follow symlinks; the reader refuses them. Replacing `report.md` with a relative symlink to its own bytes passed the guard while `read_current_control()` failed with `artifact_unreadable`. The test now rejects a symlinked artifact directly, for a message naming the file, and then runs the reader over the fixture -- the end-to-end property the digests are only a proxy for. The reader call is not redundant with the local checks: a symlinked `current-control.json` satisfies every digest, is caught by neither, and is refused by the reader as `unsafe_pointer`. Verified reachable that way. Perturbation sweep re-run, each failing on its own assertion: all three findings above, plus drifted bytes, a missing bound file, a hand-edited digest (caught by the model validator), a symlinked pointer, and an empty glob. fix(control): the §D loop finishes — reachable route, legal order, and a continuation (#429) (#444) Three separate walls stood between an agent and a declaration a human could review, and the loop needed all three gone. **Reachable.** The `declare_action` route was refused on `capability_review.policy_weakened`, the fail-closed routing flag, which stays raised whenever the direction of a policy change could not be established. Establishing it means proving no file in the tree parses as a manifest under any name — a whole-tree read that one blob past its per-candidate bound ends, and `google/adk-samples` carries 35. So a first adoption, the run with every question open, was the one run that could never be offered the route that answers them. The route now asks a separate, cheaper, sound question: does this diff introduce the gate it is judged by, touching no other resolved policy input, deleting and renaming away nothing at all? The probe, the flag, the verdict and the adoption wording are untouched. **Legal.** `apply-patches` writes into the trust root, so `agent control` refuses with `workspace_changed` the moment it succeeds and the permissions printed beside the route were computed against a manifest that no longer exists. The route now says rerun first and act on what the rerun authorizes, and when that refresh refuses it advertises the producing run's own exact local rerun instead of a fixed PR verify that scanned committed HEAD. **Finishable.** The rerun is a fresh decision, and a declaration that makes a risk judgeable turns it `blocked` — which authorizes nothing, so the proposal Shipgate drafted could never reach the person it was drafted for. A continuation receipt now pins the manifest by byte digest on both sides of the write; the run additionally compares the two manifests and requires the delta to be added `action_surface.actions` rows and nothing else. On that proof the blocked run is publish-only. `merge` and `report_complete` stay denied, and without a receipt a blocked decision authorizes nothing exactly as before. Contract 25 → 26; verifier 0.14 → 0.15, handoff v7 → v8, verify-run v4 → v5. fix(init): a field the tool chose without evidence is marked as one (#441) (#443) * fix(init): a field the tool chose without evidence is marked as one (#441) `detect` on a FastMCP Python MCP server reports "not a Shipgate target". `init` — the command the control loop routes to from `verify --preview` — wrote a manifest for it anyway, and that manifest declared `type: openapi` for a repository containing no OpenAPI spec. `id` and `path` were flagged in `placeholders[]`; `type` was in neither `placeholders[]` nor a comment, so filling in the two flagged blanks yielded a schema-valid manifest describing a source that does not exist. - The fallback `tool_sources` block is now `id`/`type`/`path` all `CHANGE_ME`, all three in `placeholders[]`, under a comment listing every accepted `type` and stating that discovery had no evidence for any of them. - `init --json` gains `tool_surface_origin` (`detected` | `scaffold`), decided by the renderer and reported nowhere else; the same fact is stated in prose in `manifest_message`, on stdout, and in `control.reason`. It is `null` when this run's render reached neither disk nor the payload, the authority rule `placeholders` already follows. - Conventional-dir discovery reads the whole tree, not just the workspace root, so a distribution whose tools sit under the import package contributes its one structural signal. Deduplicated by directory name. Two defects found while fixing those: `AdapterRegistry.require('CHANGE_ME')` told the reader to install a third-party adapter for a value Shipgate wrote, and `init --minimal` on a source-less workspace emitted an empty `openai_api:` block — a manifest the schema rejects — because the guard selecting its fallback tested a dict with fixed keys. Closes #441 Co-Authored-By: Claude Opus 5 * fix(detect): name the conventional directory a reader can open (#441 review) Three findings from reviewing the first commit. `conventional_dirs` still held bare directory names after the check started reading the whole tree, so `detect --json` and the negative-control message reported `tools/` for a repository whose only `tools/` is at `awslabs/billing_cost_management_mcp_server/tools/` — a name that resolves nowhere. It now carries the workspace-relative path. `SHIP-DIAG-PURE-PROMPT-EXPERIMENT` read the widened `has_prompts_dir`, and "only prompts/ is present" is a claim about the shape of the workspace: a thirty-file TypeScript MCP server with `src/prompts/` was reported as a prompt experiment. That control asks about a root `prompts/` again — which is what a located path equal to its bare name means. And the `--minimal` renderer's narrowed `openai_api:` gate is now pinned: a bare `prompts/` is a signal an Anthropic-only project carries too, so neither renderer anchors the block on it. Co-Authored-By: Claude Opus 5 * perf(detect): walk directories, not files, to locate conventional dirs (#441 review) Reading conventional directories from the whole tree with `path.relative_to(workspace).parts` per inventory entry cost 4.4 s on a 120k-file inventory — a whole-workspace pass `detect` already runs for exactly the monorepos #363 and #395 are about. The inventory is a file list whose entries mostly share a parent, so deduplicating parents first and testing the workspace prefix as a string slice brings the same answer in 42 ms. Behaviour is unchanged, including for a resolved symlink pointing out of the tree, which `relative_to` rejected and the prefix test now does. A workspace that is itself a filesystem or drive root no longer builds a doubled separator. Co-Authored-By: Claude Opus 5 * docs(changelog): record the conventional-dir path, root predicate, and scan cost (#441 review) Co-Authored-By: Claude Opus 5 * refactor(scaffold): drop an unused property and a schema rule stated wrong (#441 review) `RenderedManifest.is_scaffold` had no caller, and two comments enumerated the "at least one of" source blocks the schema requires while omitting three of them. Both now say only what they can. Co-Authored-By: Claude Opus 5 * fix(init): address the seven self-review findings on #441 **[P1] The renderers return a usable string again.** `cli/discovery`'s own docstring names `render_manifest_template` among the imports it keeps stable, and returning a dataclass broke every caller outside the tree while every in-tree call site was updated to `.text` — the shape of break that stays invisible until it reaches somebody else's code. `RenderedManifest` is now a `str` subclass, so `yaml.safe_load(render_manifest_template(ws))` works and the provenance still travels with the text. **[P1] The provenance is bound into `control.input_id`.** Two `--minimal` dry runs — one in an empty directory, one after an OpenAPI spec was added — published `scaffold` and `detected` with different `control.reason` under a byte-identical identity, because `manifest_status`, the placeholder list, and the advance are the same in both. An identity that does not cover the answer is a cache that serves the wrong one. **[P2] `--minimal` no longer claims discovery it never ran.** It probes artifact globs and stops, so "discovery found no tool surface" and "no framework import was read" are claims it has no observation behind: `detect` reports `langchain` for `samples/simple_langchain_agent` while that renderer scaffolds. Each renderer carries its own summary, and the manifest, `manifest_message`, and `control.reason` all quote the one belonging to the mode that ran. **[P2] The scaffold no longer closes an open set.** `ToolSourceConfig.type` is deliberately open to third-party adapters, and a repository only a custom adapter can read is *more* likely to reach the scaffold. The comment, the control postcondition, the diagnostic, and the registry message all name the built-ins as built-ins. **[P2] The `CHANGE_ME` diagnostic title matches its route** — it said "install/enable the adapter, or fix a typo", the two remedies its own branch says do not apply. **[P2] The zero-install detector locates nested conventional dirs too.** It still read only the workspace root, and the parity test compared `workspace_signals` *keys*, so `has_tools_dir` was `true` on one side and `false` on the other with every test green. Ported, plus value-level parity on the three fields and a nested-directory fixture the samples cannot provide. **[P3] The docs describe the states actually emitted** — `control.reason` carries the scaffold fact only where init's own reason is the envelope's, and `tool_surface_origin` is `null` when the render reached neither disk nor the payload. Co-Authored-By: Claude Opus 5 * fix(scaffold): restore the whole string contract, and stop denying evidence we read (#441 review 2) **[P2] copy/deepcopy/pickle.** `str`'s inherited reduction rebuilds a subclass as `cls(text)` and `RenderedManifest.__new__` takes its provenance keyword-only, so all three raised `TypeError` on a value that was a plain `str` before this field existed. `__getnewargs_ex__` carries both fields through. Restoring the string contract meant restoring all of it: caches, multiprocessing, and generic copying code do this without knowing what a manifest is. **[P2] "No framework import ... was read" was false.** Auto mode can read and classify one and still need the scaffold: `import anthropic` alone scores a full `anthropic` detection, and Anthropic is artifact-based, so with no `tools/anthropic-tools.json` there is nothing for the manifest to point at. The scaffold decision was right; the evidence sentence was not. It now says no tool export, API spec, or framework *source file* was found to point at — and that a framework import on its own is not one. **[P2] "and none matched" was false for `--minimal`.** A prompts/ directory populates `prompt_files`, and response-format-only and trace-only workspaces do the same. Those matches are real; they are deliberately not anchors, because an Anthropic-only project has a prompts/ directory too. It now says none that *anchors* a tool surface matched. **[P3]** The quickstart documents the `null` origin, and the CHANGELOG no longer calls `BUILTIN_TOOL_SOURCE_TYPES` every accepted type. fix(inputs): a codex_config source is one loaded source per MCP server (#442) * fix(inputs): a codex_config source is one loaded source per MCP server A `tool_sources[]` row of `type: codex_config` failed outright with `InputParseError` — "A tool source loader reported tool 't_codex' as belonging to a different source than the one it was read from" — whenever the workspace contained a `.mcp.json` or `.codex/config.toml` declaring a server. Since a server with no enumerable tools still mints a `.*` wildcard, *any* server at all reached it: the source type worked only on workspaces that had no MCP declarations to read, which is the opposite of what it is for. `inputs.mcp_manifest` minted a per-server id for every tool (`mcp_json:`, `codex_config_mcp:`) and then wrapped a whole file's servers in one `LoadedToolSource` named for the *file* (`codex_config_mcp:`). `core.tool_identity` requires a loaded source to name the id its tools carry — a tool arriving under another source's name is a loader defect — so the two spellings could never agree. The loader now emits one `LoadedToolSource` per MCP server, built from the server, so there is no second place that spells the id. The minted ids themselves are unchanged and stay deliberately path-free: an MCP capability is its server and tool, so moving a `.mcp.json` remains a pure rename rather than a capability addition (`mcp audit` pins that). One source per *file* was the other candidate repair and is wrong: `_native_locator` is the bare tool name for MCP-like sources, so grouping a file's servers together made a legal `.mcp.json` whose two servers both expose `query` fail with "defines the tool 'query' more than once". Why no test caught it: the three existing `type: codex_config` fixtures (`test_verify_orchestrator.py`, `test_preflight.py`, `test_codex_boundary_check.py`) all point at workspaces where the loader mints no tools, so nothing reached the check. The loader-contract failure also gains a way forward. "That is a defect in the loader, not in this repository's configuration" was accurate and terminal: the reader cannot edit an adapter they do not own, and had just been told their own configuration was not at fault. When the dispatcher can prove which `tool_sources` entry produced the read, the message now adds the workaround after the diagnosis, named as the workaround it is. `details` gains `configured_source_id`. No version moves and no schema changes. Nothing that previously produced a report changes shape, because no workspace with an MCP server could produce one. Co-Authored-By: Claude Opus 5 * fix(inputs): name every id shape codex_config mints, and pin all three Review follow-ups on the loaded-source-per-server fix. Two comments illustrated the codex minted-id shape as `mcp_json:` alone. That is one of three: `.codex/config.toml` mints `codex_config_mcp:` and a plugin-nested server mints `codex_plugin_config_mcp::`. The `configured_source_id` comment is the reference for why declarations must not be joined on `source_id` (#410 review), so a reader reasoning from one prefix handles the minority shape and misses the two that Codex adopters actually produce. Both now name all three, and the domain comment no longer states "matches no configured row" as a guarantee: a short server name is something an adopter could plausibly reuse as a `tool_sources[].id`, which the old path-based spelling made near-impossible. Two tests claimed coverage they did not assert. `test_every_minted_codex_mcp_tool_names_the_source_it_was_read_from` named a plugin-nested server and both file kinds in its docstring but asserted only that no tool mismatched its source, which a loader that stopped emitting a whole branch satisfies vacuously. It now pins the exact minted source ids and their tools. `test_the_minted_server_id_stays_free_of_the_path_it_was_read_from` guarded the rejected file-qualification for `mcp_json` only, so a cross-file fix that qualified just the TOML id would regress `mcp audit`'s pure-rename invariant with every guard green. It is now parametrized over all three prefixes. Both perturbations confirmed to fail the intended tests: qualifying only the TOML id fails the two new parametrizations, and dropping the plugin branch fails the exact-id assertion. Neither was caught before this commit. Co-Authored-By: Claude Opus 5 * docs(inputs): name all three minted prefixes in the loader comments The two remaining partial enumerations, for consistency with the ones corrected in the previous commit: the path-free rationale applies to a `.codex/config.toml` as much as to a `.mcp.json`, and the mismatch the per-server split fixes covered the plugin-nested prefix too. Comments only; no behaviour change. fix(mcp): one server name in two files is one capability, reconciled (#445) A `codex_config` row over a workspace where two packages each carry a `.mcp.json` naming `github` aborted the scan with `'pkg_b/.mcp.json' was read twice as one tool source [...] Remove the repeated shipgate.yaml entry naming 'pkg_b/.mcp.json'` — false on both counts, and naming an edit nobody could make: the manifest names `path: .` once, and that file was read once. Two deliberate rules collided. An MCP capability is identified by `(server, tool)` and by nothing else, so that moving a `.mcp.json` is not a capability change (`mcp audit` pins this); and one identity may be observed only once. Qualifying the minted id with the path resolves the collision and breaks the first rule, so it stays rejected — a named test keeps it rejected. The answer is that the two files declare one capability twice, and the reader reconciles them before the catalog sees them. Identical declarations are silent. Nothing was dropped, and a source warning is a gating input; the ordinary monorepo layout gains no evidence gap. Sameness is decided on the declaration as written, not on the fields that reach a catalog row, so two files agreeing on every tool and disagreeing on the command that serves them are not identical. Disagreeing declarations are merged conservatively, never by picking one. The merged server carries the union of the tools; a claim that raises risk survives from any one declaration (`destructiveHint`, an unenumerated `server.*` remainder, secret environment names, an external URL, an authority scope, `approval_mode: approve` — which reaches even a tool whose own file said nothing about approval); a reassuring claim survives only when every declaration makes it (`readOnlyHint`, a local-documentation server, a transport); and two files describing one tool with different schemas leave the interface unknown rather than publishing either. A declaration the auth parser refuses stays refused rather than being rebuilt as a valid one. Each merged tool names the file that declares it, and each declaration that did not enter the catalog on its own becomes a `SourceSurfaceOmission` — an `adapter_parse` row in the #403 exclusion ledger, accounted for by the source warning that reports it. Two defects found alongside it: - A `codex_config` row over *any* config naming an MCP server aborted the scan before reaching the above: the loader returned one file-level tool source holding tools stamped per server, so every tool was reported as belonging to a source other than the one it was read from. It returns one source per server. Broken since #191; no test scanned such a workspace end to end. - Two `tool_sources` entries reading one server from two files is the one duplicate no reader can reconcile — neither entry's declarations speak for both — and it was reported as a repeated manifest entry, which it is not. A third `details.cause`, `duplicate_across_artifacts`, names both files and the two repairs that exist. Documented in `docs/errors.json`. No verdict, count, or version moves: `report_schema_version`, `contract_version` and every published schema document are unchanged. fix(reporting): the Root agent line names the agent, not its digest (#329) (#438) * fix(reporting): the Root agent line names the agent, not its digest (#329) `_append_binding_surface` printed `Root agent: agent_v1:7205d836...` at the head of the Agent Binding Surface section, and every shipped sample carried it. That is the #329 invariant: proceeding must not require understanding the internal identity model, and a derived digest appears in no file the adopter has. It now resolves through `core.surface_exclusions.agent_label_index` -- the same index a binding gap uses, so the section header and a finding about that agent cannot spell it two different ways -- and reads `Root agent: durable_order_agent [conductor_workflows]`. `unresolved` covers every way the graph can fail to name a root: no root, the `legacy_direct` sentinel, a root id no node carries, and nodes that disagree with the root. Chaining back to `root_agent_id` would restore the digest on exactly the graphs that already read worst; the id stays in `binding_surface_facts.root_agent_id` for a bug report. The guard that should have caught this could not see it. The renderer escapes every value it prints, so the line reached `test_sample_markdown_speaks_the_adopters_vocabulary` as `agent\_v1:7205d836` -- and every shape in `INTERNAL_ID_SHAPES` and every term in `MANIFEST_SPELLED_TERMS` contains an underscore, which made that sweep close to vacuous over the whole half of the report that goes through `_safe_markdown_text`. The sweep now un-escapes each line first, against a `unescape_markdown_text` defined as the escaper's inverse beside it. Running the un-escaped sweep over all five samples before fixing anything found `Root agent:` as the only violation, so the class is contained. Tests: the root line names the agent, asserted on the un-escaped text; the four unnameable-root shapes never fall back to the id; a negative control asserting `internal_vocabulary()` returns `()` on the raw escaped string while the sweep rejects it (without it the guard can go vacuous again silently); and an escape/un-escape round trip over every `MARKDOWN_ESCAPE_CHARS` entry including a literal backslash, which is what makes the escaper's ordering load-bearing. Markdown-only, so no schema moves. The five sample `report.md` goldens are regenerated; no `report.json` changed. `samples/conductor_agent/expected/ current-control.json` moves with them because its `report_markdown` digest matched the old golden exactly and now matches the new one. Co-Authored-By: Claude Opus 5 * fix(reporting): `unresolved` is only ever said about an agent (review) Review of this PR found one real regression it introduced, plus a docstring that asserted the regression could not happen. `legacy_direct` is the `root_agent_id` of the compatibility assessment a report carries when no agent graph was resolved at all — `core.findings.report_builder` builds it whenever its `binding_surface_facts` argument is `None`, and it is the schema default for a `report.json` written before the field existed. It is truthy and no node carries it, so resolving it through the label index printed `Root agent: unresolved` directly between `Status: structural` and `Pass eligible: true`, describing a failure the report was not reporting. Confirmed by rendering a bare `ReadinessReport`. This is the same mistake `_root_agent_line` already warned about for the tool-source state, one branch below the warning, so it now gets the same shape of answer: `Root agent: none (tools bound directly, no agent graph)`. `_agent_display`'s docstring claimed the `unresolved` fallback "is a fail-safe rather than a state the scanner can reach". True of a graph the walk built, where the root and every entry point have nodes; false of the sentinel, which lands on that fallback on its first call. A claim in prose with nothing holding it up is how the branch got written. Rewritten to say what callers owe the function, and the sentinel is answered before the index is consulted. The parametrized `legacy_sentinel` case had pinned the wrong output: it sat under "every way the graph can fail to name a root", which `legacy_direct` is not. Split into two tests — one on the assessment, one reaching it through the schema default the way production does, so the coupling cannot drift. Also from the review: `unescape_markdown_text` was new public module surface whose only caller is the test sweep, while the escaper it inverts is private. Renamed `_unescape_markdown_text`. Not changed, flagged instead: `Entry points:` is now a verbatim repeat of `Root agent:` for a rooted graph, since both resolve the same id through the same index. The redundancy is #432's rendering decision and predates this branch; it was only masked because the two lines used to spell the same agent differently. Full CI-equivalent run green; sample goldens unchanged, since every sample has a real root. Co-Authored-By: Claude Opus 5 * fix(reporting): a single declared surface is the same state as several (review) Addresses both review comments on #438. [P2] `_select_root` returns the sole surface node when a repository declares exactly one reviewed `tool_sources[].binding` and nothing observed an agent (core/agent_bindings.py:1025-1026). That graph therefore has a *truthy* `root_agent_id` pointing at a node whose `kind` is `tool_source`, unlike the two-surface form, which leaves the root unset. Keying the line on "is a root id set" printed `Root agent: github_mcp` beside `Pass eligible: true`, announcing an agent object the repository does not have, while the two-source form of the same repository printed the correct sentence. Reproduced at this head before fixing, exactly as reported. Whether a repository has an agent object cannot depend on how many sources it declared, so the answer is now read from the root node's `kind`, and the sentence itself is a single constant rather than a literal in each of the two branches that reach it — the drift was the defect, so one spelling is the fix. Both forms now render `none (graph rooted by declared tool sources)`, and the `Entry points:` line still names the surface that roots each. The renderer was chosen over the producer deliberately: changing which node `_select_root` returns would move `root_agent_id` in `report.json`, a published field that `ci/release_decision` reads, which is a contract change rather than a rendering fix. Regression test added beside its two-source twin in `test_source_binding.py`, asserting the producer's shape first — truthy root, one `tool_source` node, `entry_point_agent_ids == [root_agent_id]` — so it cannot pass by accident if that shape changes. [P3] The changelog entry listed `legacy_direct` under what `unresolved` covers while the paragraph below it, and the code, render `none (tools bound directly, no agent graph)`. Removed from the list, and the new single-source state recorded alongside it. Also corrected the helper's name there, which was privatised in the previous commit. Full CI-equivalent run green. Sample goldens unchanged: every sample has a real agent root, so neither state occurs in them. feat(binding): a published tool surface is one declaration, not one row per tool (#432) (#435) * feat(binding): a published tool surface is one declaration, not one row per tool (#432) Binding an MCP server's own surface required naming every tool individually under `agent_bindings.declarations[].tools` — 116 selector rows for `github/github-mcp-server` to state a fact that is structurally true of the source, and the point at which an adopter stops. The two shorter spellings a reader reaches for both dead-ended: `agent_bindings.root` naming the source reported `ambiguous_root_agent` as though a better selector existed, and a `"*"` selector was matched as a literal tool name. Until one of them was written `reachable_tools` was 0, nothing downstream ran, and the verdict was `insufficient_evidence` whatever the change did. * **A new `tool_sources[].binding` block** — `{complete: true, reason}` — states once that a source's published surface *is* the surface under review. It sits where `tool_sources[].authority` sits and makes the same shape of argument: binding, like authority, is a fact about the source rather than about each function it exposes. For an agent that is not true — a catalog may hold 63 OpenAPI operations of which the agent wires 5, and #385 drew that boundary deliberately — so the block is opt-in per source and changes nothing where it is not written. * **Additive and widening.** It can only move tools *into* the analysed surface, where every check then judges them. On a `verify --base` run over an MCP server that adds a `destructiveHint` tool, that tool is now judged and blocks on its missing controls instead of leaving the surface before any check could see it. * **The dead end is a route.** When nothing observed an agent object, the root gap no longer prescribes a selector that cannot exist. It says so, routes to `shipgate.yaml#tool_sources[].binding`, and the declaration scaffold writes one block per configured source carrying the ids read off the surface and `` for both halves of the judgement. A catalog whose tools come from a per-scan adapter has no row to declare on, so it keeps the closed-world `agent_bindings.declarations` route rather than being handed a remedy the schema rejects. Once a source is declared, `root: {object: , source_id: }` resolves to it. * **It stays a human declaration.** Inferring "this source binds everything" from the source's own content is the #268 attack; the block is refused in agent-authored `tool_sources` proposals and is a human-owned placeholder. A reviewed declaration that binds no tool fails closed rather than proving a graph over an empty analysed surface. No published schema changes: the manifest field is additive, and a CLI that predates it rejects the key with a routable `ConfigError` (exit 2) rather than ignoring a reviewed claim. Co-Authored-By: Claude Opus 5 * fix(binding): pair every published path with the block it names, and close the pattern spelling (#432 review) Review pass on this branch. Three findings, and the first is a real path/template contradiction the change introduced. **One failure, two blocks.** A reviewed source binding that reached no tool is raised against its surface node — which is the graph's *root* whenever it is the only declared surface. That put it inside `_binding_declaration_template`'s root-scoped branch, so a gap whose own `path` named `shipgate.yaml#tool_sources[].binding` shipped an `agent_bindings.declarations` block listing an unrelated source's tools under an `agent:` naming a tool source, which resolves to nothing. Asserting the `path` alone passed; asserting the template alone passed. It is now checked as a pair, by a rule over every shape that raises a binding gap rather than at the site that was wrong. **The other spelling the issue reported.** `tools: [{tool: "*", source_id: …}]` is a reader trying to say "all of this source's tools" in one row. It was matched as a literal name and reported as matching none — true, and no help. It now names the statement it was reaching for, at the binding site only: a pattern in an `action_surface` row is a different misunderstanding with a different repair. **An explicit selector still means the agent object.** `root.object` searched every node, so adding a `binding` block to a source whose id happened to match an observed agent turned a selector that used to resolve into `ambiguous_root_agent`. An observed agent now wins the name outright; a declared surface answers the selector only where nothing observed offers that name, which is the artifact-only case the selector could not be satisfied in at all. **And the guard that was guarding nothing.** A perturbation sweep over thirteen mechanisms found three that failed no test: the surface's disjoint id namespace, its exclusion from name resolution, and its invisibility to root heuristics — all three the #385 "a prescribed fix whose own side effect breaks the graph" class. The collision test could not reach them because its fixture collided on a name whose source contributed no agents at all. Replaced with a Google ADK source named `front_desk` holding an agent named `front_desk` — the observations agree on both keys the graph dedupes on — and a root/sub-agent shape where a surface counted among the candidates makes the picked root ambiguous. Every mechanism now fails at least one test. Co-Authored-By: Claude Opus 5 * fix(reporting): name a source-scoped binding row as a source, in every field (#432 review 2) Second review pass. One sentence under the verdict was wrong three ways at once, and each is a class this repository has been bitten by before. **It spoke in the voice of an action about a subject that has no agent.** `_GAP_PHRASE` is keyed by kind alone, so the row a reviewed `tool_sources[].binding` raises read "the agent's tool bindings are unproven (empty_source)" — which sends the reader looking for an agent that does not exist. Increment 3 already built the mechanism for exactly this (`subject_kind: tool_source` selecting a source-scoped phrase); the row now carries that kind, and `SOURCE_SCOPED_GAP_KINDS` states the covered set beside the table it governs rather than reading it off one of the two producers and silently missing the other. **The subject repeated its own qualifier.** `agent_subject` qualifies a name with the source it was read from, to separate two sources defining an agent of the same name — and a declared surface is named by its own source id, so it rendered `empty_source [empty_source]`. A qualifier that *is* the name separates nothing and is dropped. **And it published a source id in a field two joins read as a tool id.** `surface_exclusions` and `semantic_consistency` join a binding gap's `subject_id` against canonical tool ids, and neither checked which namespace it was in — the exact hazard the schema's own comment names. Both now scope to `subject_kind: "action"`, which is what every existing row already carries, so nothing else moves. Guarded by the collision constructed rather than argued about: a `tool_sources[].id` spelled as the canonical id of the tool that is genuinely excluded. Without the guard the ledger reports that tool as `evidence_gap`-accounted, gated by a row about a different subject — the ledger claiming gating it does not have (#404). The perturbation sweep now covers seventeen mechanisms and every one of them fails at least one test. Two of the new guards were vacuous on the first attempt for opposite reasons: one because its fixture could not reach the code it was written for, the other because the assertion re-implemented the filter it was checking instead of observing its product. Co-Authored-By: Claude Opus 5 * fix(binding): one punctuation rule for the selector sentence, and pin the status pairing (#432 review 3) Two small ones from a read-through of the whole diff. The pattern hint appended `. A binding selector …` to the resolver's own sentence, which carries no terminator — but to the fallback sentence, which does, producing `exactly.. A binding selector`. One rule for both. And a reviewed source binding that reaches no tool leaves `status: declared` beside `pass_eligible: false`, because a reviewed declaration does exist and the gate is carried by eligibility — exactly as it already is for an `invalid_binding_annotation` on an otherwise declared graph. Pinned so the pairing is a decision rather than an accident. Co-Authored-By: Claude Opus 5 * test(binding): move the shared gap fixtures beside the other helpers (#432) The path/template table and the one-workspace-per-gap-shape fixtures were defined halfway down the file and used above it. Python resolves that fine; a reader does not. Co-Authored-By: Claude Opus 5 * docs(binding): the documented example is one a scan actually loads (#432 review 4) The manifest reference example wrote `path: __toolsnaps__` — the directory `github/github-mcp-server` keeps its tool snapshots in, which is where the 116 tools in the issue came from. An `mcp` source names one protocol artifact, so a reader copying the block got `static input is not a regular file` before any of the prose about it could be true. A documentation example that does not run is worse than no example: it is a claim about behaviour nobody checked. Guarded by running it — the block is lifted out of the section, its artifact materialised, and the outcome the section claims asserted. The first version of that guard was vacuous, and instructively so: it created the file at *whatever path the doc chose*, so `__toolsnaps__` became a file named `__toolsnaps__` and loaded perfectly. No scan of an example can catch the shape of its own path. That half is a stated rule instead, taken from where the codebase already states it (`_FILE_SOURCE_TYPES` — those types name a single protocol artifact, so the documented path carries a suffix), and verified by putting the original defect back and watching it fail. Co-Authored-By: Claude Opus 5 * fix(binding): address the five review findings and publish the entry-point state (#432 review) All five inline findings reproduced before fixing, and both contract points addressed. Report schema 0.41 → 0.42; v0.41 frozen and hash-pinned. **[P1] A per-scan adapter must record which configured row produced a result.** The dispatcher cannot pass a per-scan adapter the config object, so it attributes pass-2 results by the `(adapter source type, minted source id)` pair. A Codex plugin MCP inventory is minted as `codex_plugin:/:inventory` with source type `codex_plugin_mcp_inventory` — *neither* half — so its tools carried no configured provenance and adding `binding` reported that the source contributed nothing while its tool sat reachable in the catalog. The adapter now stamps the association where it is known, and `_absorb` no longer overwrites an id an adapter recorded. This also silently fixed `tool_sources[].authority`, which had the same exposure and failed quietly. Checked as a class rather than an instance: a probe across all nine built-in `tool_sources` types found exactly one broken (`codex_plugin`), and a parametrised sweep now asserts the additive guarantee per type. `codex_config` is excluded and filed separately — a row of that type plus a `.mcp.json` fails every scan on `main`, before this block exists. **[P1] Two reviewed closed-world statements about one node conflict.** A declared surface selected by `agent_bindings.root` answers to `agent: root`, so "this source publishes a and b" and "this root reaches only a" could both be written and were silently unioned into a `declared` / `pass_eligible` graph. They now raise `conflicting_binding_evidence`, beside the structural rule they mirror. Two statements that agree are still not a conflict. **[P2] A merged tool identity is declarable by every contributor.** `binding` is a widening claim, so a tool two configured sources contribute is covered as soon as any of them declares — the helper's own docstring said so while the code returned "undeclarable" and withheld the route. It returns `None` only where a catalog tool has no configured source at all. That is deliberately not the rule `authority` obeys: authority *replaces* published evidence. **[P2] Runtime and published schema now accept the same manifests.** Eleven shapes, zero mismatches. `complete: 1` is refused in both (a `strict` constraint cannot be applied to a `Literal`, so the rule is stated as a before-validator); `reason` carries `VISIBLE_CONTENT_PATTERN`, so the schema rejects the blank the CLI rejects and both refuse a reason made only of U+200B. `agent_bindings.declarations[].complete` is deliberately untouched: tightening a field that has shipped is a compatibility decision of its own. **[P2] Remediation values that can close the gap.** The binds-nothing row advertised `complete:true` / `reason` — the fields the failing declaration already carries, rendered verbatim into `fix_task`. It now publishes `correct_source_path` / `remove_binding` and targets the exact row (`tool_sources[id='…'].binding`, spelled as `action_surface.actions[tool='…']` already is). **The rootless entry-point state is published as itself.** `pass_eligible: true` beside `root_agent_id: null` was a new state no contract described and `report.md` rendered as `Root agent: unresolved`. `binding_surface_facts.entry_point_agent_ids` is every node the reachability walk started from — empty exactly when nothing rooted the graph, equal to `[root_agent_id]` for every graph a prior release could produce — and `agents[].kind` (`agent` | `tool_source`) says what each node in that array is, which the passed-verdict contract needed anyway for the skill-only Codex plugin package root it was already hand-waving about. **And the release note no longer says "no published schema changes"** beside a `docs/manifest-v0.1.json` that gained a field. Co-Authored-By: Claude Opus 5 * docs(binding): state the `passed` criterion where a rootless graph can meet it (#432 review) The review found `docs/passed-verdict-contract.md` describing inventory and actions as root-reachable. The same sentence is stated in four more places a reader is more likely to open — README, AGENTS.md, the Claude Code skill, and docs/examples.md — and my own manifest-reference section had added a fifth. A repository that publishes tool surfaces has no root agent by construction, so "`passed` requires a complete root-reachable static binding graph" told exactly the adopter this change is for that they cannot get there. It reads "from its entry points" now, which is the same requirement stated in terms the new `entry_point_agent_ids` names — and is unchanged for an agent application, whose only entry point is its root. fix(reporting): a new evidence gap names the subject that left the surface (#433) (#434) * fix(reporting): a new evidence gap names the subject that left the surface (#433) The exclusion ledger from #403 records precisely which subject each stage removed from the analysed surface — `("binding", "find_duplicate [github_mcp]", "evidence_gap")` — and no human-facing surface carried it. A reviewer of `github/github-mcp-server#3020`, a PR that adds exactly one tool, was told "1 of 83 evidence gap(s) are new in this diff" and never *what* the one was. The blockers there are pre-existing debt about the other 115 tools, so the `Most severe:` clause that carried the subject in the cases where this looked fine was about something unrelated to the change. That is #403's own thesis — a stage computed the right signal, stored it, and did not connect it to the decision — standing at the ledger's own output, and it is the epic's last open box. `verify`'s headline, and with it `control.reason`, `control.next_action.why` and the PR comment's `Summary:` / `Next action:` lines, now continue: Excluded from analysis: find_duplicate [github_mcp] — added by this diff and not bound to the root agent. Rendering only. No verdict, count, gap, finding or permission moves, and no version does either. **Selected by the ledger's own pointer.** A row is named when its `accounted_by` gap is one of the identities the "N of M are new" count was computed from, so the clause and the number in front of it cannot describe different sets. `not_claimed` rows carry no pointer at all, which is why a settled workspace gains no clause. The subject printed is the ledger entry's own string, built by `catalog_subject`, so it cannot drift from the gap row it came from (#413) — and `nameable_subject` refuses that function's tool-id fallback, the one spelling that reaches prose carrying a digest because `derived_id_kind` deliberately allows a derived shape in the name position. **Bounded, because it shares a 400-byte envelope.** At most three subjects, each capped as scanned input, grouped by cause with an `and N more` tail. The clause fits itself to its own byte budget by naming fewer and counting more, rather than being cut mid-name by the envelope's tail truncation. Reason tokens render through one table beside the builder that emits them, with an AST-scanned test asserting the two sets are equal in both directions, so a new emitter fails rather than silently printing the generic fallback. Co-Authored-By: Claude Opus 5 * fix(reporting): fit the headline's context by whole sentences (#433 review) Self-review of the clause this PR adds. The evidence-gap context was composed into one string and sliced to fit the 400-byte envelope, which is right for one unbroken run of untrusted text — a blocker title, a manifest path — and wrong the moment that text *names subjects*. `delete_repo…` is not a shortening of `delete_repository` a reader can act on; it is a plausible other tool. And `Excluded from analysis: find_dup…` names nothing at all. Reachable on every composed route: with the trust-root suffix reserved, a three-subject clause is cut at byte 193, inside the subject list. `_gap_provenance_note` now returns ordered whole sentences, most load-bearing first, and `_fit_sentences` takes sentences off the end until the rest fits. The plain `_lead` route is bounded too — it applied no budget at all before, so the envelope's own tail truncation did the cutting, and that is the route a blocked repository takes. The pre-existing "no new evidence gap" note is split the same way, so a tight budget drops the declaration remedy and keeps the fact rather than losing both. Byte-identical wherever the whole note fitted. A `str` satisfies `Sequence[str]` and iterating one yields characters, which no type checker sees and which would render the note with a space between every letter, so the parameter refuses one outright. Also from the review: `test_a_settled_workspace_adds_no_exclusion_clause` was passing for the wrong reason. Its fixture's unnamed MCP entry is itself a gap-backed exclusion, so the clause *did* fire and only the assertion ("delete_repository not in note") missed it. The fixture now declares the new tool, leaving one pre-existing `not_claimed` row and `gated: 0` — the acceptance box as written — and the adapter-omission case it used to be is kept as its own positive test, because the clause is not binding-only. Co-Authored-By: Claude Opus 5 * test(reporting): keep the title axis live, and re-state the clause cap's reason (#433 review 2) Two consequences of the round-1 fixes. The budget parametrization passed a prebuilt report for two of its four routes, so the short/crowded title axis ran twice with identical inputs there. Each case now supplies a report *factory* taking the title. And `_EXCLUSION_CLAUSE_MAX_BYTES` was justified by a truncation that `_fit_sentences` now prevents. Its real job is the opposite one: a clause that does not fit is dropped whole, so an unbounded clause is one that never survives a route with a reserved governance suffix. Stated as such, because a constant defended by a reason that no longer holds is the next person's deletion. Co-Authored-By: Claude Opus 5 * fix(reporting): the clause's lead-in must be true of every stage under it (#433 review 3) "Excluded from analysis" is the ledger's own framing, and it is false of one of the ledger's own stages. `_surface_completeness_exclusions` says so explicitly: an `incomplete_surface` tool *is* analysed, as far as its own surface could be read, and the excluded subject is the unread remainder, which has no name — so the tool names it. Reading "Excluded from analysis: charge_card [billing]" a reviewer concludes no check saw that tool, which is the opposite of what happened. One lead-in covers a grouped list, so it has to hold for every stage that can appear under it. "Not fully analysed" does, for an unbound tool and for a partly-read one alike, and the phrase after the dash still says which case a row is. Also adds the two tests the grouped renderer never had: two causes under one lead-in with a tail of its own, and the byte cap naming two subjects instead of three rather than overrunning — the case where the whole clause would otherwise be dropped by the headline composition. Co-Authored-By: Claude Opus 5 * docs(reporting): finish the lead-in rename in prose (#433 review 3) Co-Authored-By: Claude Opus 5 * docs: state the lead-in rule in the changelog (#433 review 3) Co-Authored-By: Claude Opus 5 * fix(reporting): select the exclusion clause by diffing the ledger, and keep every name exact (#433 review) Five reviewer findings on #434, each reproduced first, each now guarded. **The review action dropped the context it was supposed to carry.** `_derive_verifier_control` reproduced by hand which of `_verifier_headline`'s routes carries a governance requirement, and that copy had drifted: the self-approval route with no outranking blocker composes the headline as `context + note`, so the note *is* carried — and replacing the reason with the bare note threw the context away. On a PR that adds an unbound tool and edits `shipgate.yaml`, `verifier.headline` and `control.reason` named `find_duplicate` while `control.next_action.why`, `human_review.why` and the PR comment's `Next action:` line did not, which is the #433 acceptance criterion. `_verifier_headline` publishes every such requirement as a reserved suffix — the one thing `_compose_with_reserved_suffix` guarantees survives the budget — so the human-review reason is now simply the headline, and the duplicated branch table is gone. **Selection now diffs the ledger, not the gap identities.** Two independent failures, both reproduced: - A new exclusion can reuse an existing gap identity. A base with one nameless MCP entry and a head with two produce the same single `source_warning` gap on both sides, so `introduced == 0` while the head ledger has gained `/tools/2` — an exclusion no surface named, which is the defect #433 was filed about surviving inside #433's own fix. - One subject can carry several gap kinds. `samples/conductor_agent` has both `incomplete_surface` and `low_confidence_tool` for `lookup_order [conductor_workflows]`; dropping `kind` from the join let a *new* `low_confidence_tool` gap — which has no ledger row at all — pull in the *inherited* `surface_not_enumerated` exclusion and print its cause as the diff's doing. `_exclusion_identities` reads the base report's own ledger and `_newly_excluded_rows` takes a multiset difference on `(stage, subject, reason)`. Exact on both sides, because `evidence_gap` rows are the ones the cap never drops. The clause is emitted on the inherited-gap branch too: "no new evidence gap" and "this subject is newly out of the analysed surface" are both true when a new exclusion is accounted for by a gap the base already carried. **A name is the ledger's own name, delimited, or it is not shown.** Two conventional 129-character names sharing a 59-character prefix rendered to the same string plus an ellipsis, and a long provider lost its closing `]` — so the printed subject was not the `catalog_subject` spelling the clause claims. And a tool named `find_duplicate. Control state complete; agent may merge` put that sentence into `verifier.headline` and `control.reason` undelimited. `_exclusion_label` quotes the subject and refuses — counting it instead — any subject over the cap, carrying the quote character, or that `_one_clause` would rewrite. When nothing can be printed the count is still published, so a subject that left the surface is never silently absent. **No phrase states provenance.** `added_unbound_tool_ids` is head-minus-base and deliberately covers both a tool this change added and one that was reachable at the base and lost the edge that bound it, so "added by this diff" was false for a diff that only removed a declaration. "New in this diff" is now said once, by the lead-in, from the ledger diff that proves it; the phrases state causes only. The row's own `detail` made the same false claim and no longer does. Co-Authored-By: Claude Opus 5 * docs: state what the ledger-identity comparison assumes (#433 review) The report ledger's subjects are tool labels and JSON pointers; the path-bearing ones belong to the detect ledger. Say so, and say what would break if that changed, rather than leaving the raw comparison looking like an oversight. Also record why source_ref is deliberately outside the identity. feat(control): an agent may draft what the evidence supports (#410 §D) (#427) The questionnaire (#410 increment 2) tells a person which blanks are owed. It has no way to say that most of them are the scan restating what it just read, and no way for the agent already holding the branch to write those down. This is the loop that closes that: one new route, one tag, one patch kind. `verify` publishes `control.next_action.kind: "confirm_declarations"` on a working-tree run whose verdict is `insufficient_evidence` and at least one open question is one the scan can answer from its own evidence. It carries the exact `apply-patches` command that writes those answers plus every open question, each tagged `authorable_by`. The agent applies what it may, commits it to the branch, and stops at the rest by name. Authorship is decided by content, never by who is running. A row is `coding_agent` only where the scan filled every blank in its template — an effect, from the closed vocabulary, never weaker than any reading observed — and only where the question is not one asking a person to look again: a `declaration_drift` row restates a confirmed answer beside a moved pin and stays human. The rules are enforced on the models, and the schema binds the patch to the row that published it: the patch is exactly the template, split into the keys that name the action and the fields that are written. What the agent gains is a pen, not a decision. `declare_action` is outside the default `--kinds`, writes only into fields the manifest leaves silent, and refuses (exit 5, nothing written) on a conflicting answer, two equally compatible rows, or a manifest that moved. Same-name tools from two providers are matched on the qualifiers they share rather than refused. `requires_human_review` stays true and the route's permissions are publish-only: writing a declaration touches the trust root, so a person still merges it. `patch.target_path` is relative to `manifest_dir`. The row is embedded by the packet, the SARIF file and cached base scans, all of which travel; the absolute form named a temp archive that no longer existed and changed every run, moving digests meant to be reproducible. Three fixes fell out. `VerifierArtifact` demanded an agent repair route's command equal the fix task's *rerun* command, forbidding the one thing such a route exists to publish. `apply-patches` now hashes the bytes on disk, so an untouched CRLF manifest no longer reports drift for ever, and preserves the newline style. And it pins YAML sequence indentation to the style `init` writes, so a one-line declaration is not buried in whitespace. Report 0.40 -> 0.41, packet 0.16 -> 0.17, verifier 0.13 -> 0.14, contract 24 -> 25 (minimum_control_contract_version stays 21). feat(policy): the rules layer is a pack, chosen once at init (#410 §F) (#426) * feat(policy): the rules layer is a pack, chosen once at init (#410 §F) Shipgate has always had a rule saying a financial write needs approval, an audit log, and idempotency. It was written four times in the engine — as `BUILTIN_EFFECT_OBLIGATIONS`, as literals inside three branches of the action lens, and twice more as the effect sets in the tool-level `SHIP-POLICY-*` checks — and it was not selectable, not named, and stated nowhere the adopter reads. Seven findings said "lacks a declared approval policy" and none of them said *what* requires one. `policies.control_pack` now names that rule set for the repository: `default` (today's rules exactly), `financial-strict`, or `read-only-agent`. `shipgate init --control-pack ` writes the selection with a glossary of the alternatives above it; `init --json` reports the choice and every answer it takes; `doctor` reads it back. Omitting the key means `default`, so every existing manifest keeps its verdict — findings, severities, `blocks_release`, and release decisions are unchanged across the shipped samples. **Selecting a pack can only tighten the gate.** Every built-in pack requires at least what `default` requires — asserted at import, and pinned by a real scan over every (pack × effect) pair plus a whole-finding-set comparison, because a table read by eye is not the property that matters. That invariant is also why the choice needs no new report field: a report that passes under any pack would have passed under `default`. **A pack decides which control findings fire, never what a declaration means.** The obligation lattice deciding whether a declared effect covers an inferred one stays the built-in table, so a pack requiring identical controls for two effects still cannot let a declaration of one discharge the other (#413). The engine reads the one table rather than mirroring it. Effects with no control check of their own report a pack obligation through `SHIP-ACTION-POLICY-VIOLATION` at `high`, naming the rule in `evidence.policy_id` as `control-pack::` — one rule, one finding, however many effects share it — and, like the four dedicated families, a `checks.ignore` entry records the exception without waiving the blocker. `scan` stdout and `report.md` now name the rule that wanted each missing control, once per rule instead of once per tool: "financial write requires approval.required, safeguards.audit_log, and safeguards.idempotency — 3 actions short". Both render one projection, so they cannot report different counts from one report. `SHIP-ACTION-POLICY-VIOLATION`'s built-in title says "without required controls" rather than "without approval", which under a stricter pack was telling an action that *has* approval that it did not. No new check ids, no report-schema change, no new CLI command. Findings carry `evidence.control_pack`, excluded from the fingerprint so baselines recorded before the field keep matching; `run_id` does move, because which rules produced a report is part of what that report is. Co-Authored-By: Claude Opus 5 * fix(policy): guard the mechanisms a perturbation sweep found unguarded (#410 §F review) Review pass on this branch: eleven perturbations, each expected to break tests. **Three returned zero failures**, and each was a real gap. 1. **The tool-level checks' pack-awareness was completely unguarded.** Reverting `subject_requires_approval_review` / `subject_requires_confirmation_review` to their hand-written effect sets broke nothing — the mechanism worked and nothing proved it. Now walked per (pack × effect) through a real scan, asserting each check fires exactly where `effects_obliging` says it should. 2. **The console line was asserted only through its projection.** Deleting it broke nothing, in a test named "the report and the console name the same rules" which read only the report. Both halves now go through the real command, and a second test pins the truncation counter. 3. **The rule-row discriminator was an unreachable check-id branch.** `finding_control_rule` excluded the tool-level `SHIP-POLICY-*` findings by name — a guard scoped to one id shape, vacuous for the next check that carries a pack, and unreachable today because those findings have no `policy_id` either. A rule row is about an *action*, so the gate is now an `action_id`, and the test feeds it a tool-level finding *with* a `policy_id` so it can only pass for the structural reason. Also from the same pass: - **A rule nobody can satisfy is a trap, not a gate.** All 27 (pack × effect) pairs now declare exactly what the finding tells them to declare and require every control family to go quiet. `confirmation.required` is the one control not writable on the action row, which is exactly where an unsatisfiable instruction would hide (#399). - **The reference table is the table.** `docs/manifest-v0.1.md` gains the full effect × pack matrix, rebuilt from the pack objects and compared row by row — a pack is chosen by reading the docs, so a drifted row is a wrong answer given confidently. The prose table said nothing about `financial-strict` requiring approval on outbound communication or an audit log on code execution and destruction, which the matrix now states. - `financial-strict`'s summary names the posture it actually encodes (recoverability and record) rather than restating `read-only-agent`'s. - `verify --base` exercised under two packs on a real two-commit repository: no `CapabilityFactV1` divergence, and blockers go 2 → 4, so the monotonicity property holds through the base-comparison path too, not only plain `scan`. - `README.md` and `STABILITY.md` record the field, the fingerprint/`run_id` split, and the baseline consequence of moving to a stricter pack. - The console line reads `actions short of: (n); (n)`. Co-Authored-By: Claude Opus 5 * fix(policy): a pack move is a gate weakening, and the comparator has to see it (#410 §F review) Review round 2 found a fail-open this branch introduced. `policies.control_pack` can weaken the release gate — moving `read-only-agent` → `default` drops obligations across eight effects — and it was absent from the effective-policy snapshot `SHIP-VERIFY-POLICY-WEAKENED` compares base against head. The manifest is a trust root, so a human still merges the change; but "a reviewer will spot it in the diff" is exactly the argument that check exists to reject for `ci.mode`, `fail_on`, and severity overrides. Every other field in that snapshot answers *does the same finding still block?*. A control pack answers *does the same action still produce the finding?*, which is the other way a gate gets weaker. `effective_policy.control_pack` publishes the pack in force and `SHIP-VERIFY-POLICY-WEAKENED` gains `kind: control_pack_weakened` — **one finding per pack move**, the shape `fail_on_loosened` already uses. The first draft emitted one per effect: a single changed line produced eight `high` findings repeating one sentence, which is the shape this issue exists to remove. `removed_controls[]` carries `{effect, controls}` for each, the title says how many, and the sentence names three plus how many it is not naming. A base snapshot with no pack is compared *as* `default` rather than skipped. A build without control packs could not have loaded a manifest naming one, so it ran `default`'s rules — and resolving it that way puts the "no pack is weaker than default" invariant in the comparison instead of in a comment beside it. That is a published-schema change, so report schema `0.38 → 0.39` with v0.38 frozen, hash-pinned, and read forward — the bump this branch had deferred, now carrying the auditability it was deferring too: a report says which rules produced it. Full checklist: regenerated schemas, sample goldens, the `.claude` skill render plus its prior-render hash, `llms-full.txt`, the qualification gate's exact-equality pin, and the frozen/current parity rows. Packet goldens are untouched — `effective_policy` is not part of `run_id`. Verified end to end on a two-commit repository: the weakened head is `review_required` with `policy_weakened: true`, and the nine-pair sweep covers both directions including `financial-strict` ↔ `read-only-agent`, which are incomparable — each requires something the other does not, so either direction is a weakening for some effect and asserting only on "strict → default" would have left the sideways move silent. Co-Authored-By: Claude Opus 5 * fix(policy): the weakening sentence agrees at every length (#410 §F review) Round-3 read of the copy this branch adds. No two built-in packs differ on exactly one effect, so the singular reading of the weakening title and sentence is unreachable through a real scan — it would have shipped unread, saying "1 effect require less". Both agreements are now built in two named helpers and exercised at each length directly, which is the only way to reach the branch. Also documents the detection where the adopter chooses the pack: changing `policies.control_pack` to one that requires less is reported by `verify --base`, and changing it to make a scan pass is what that check exists to catch. Co-Authored-By: Claude Opus 5 * test(policy): prove the missing-base rule is not the same as skipping it (#410 §F review) A fourth zero-failure perturbation. Replacing "resolve a `None` base pack to `default` and compare" with an early `return []` broke nothing — and could not, because every pack today requires at least what `default` does, so a `None` base yields no weakening either way. The test named for the rule asserted three things none of which distinguish it. What the branch is for is the case `_assert_packs_extend_default` forbids from existing. Registering a weaker-than-default pack and comparing a `None` base against it now exercises exactly that, which is the difference between an invariant enforced twice and an invariant assumed once. Co-Authored-By: Claude Opus 5 * fix(policy): address the expert review — two P1s and five correctness findings (#410 §F) **[P1] `init` recovery routes drop the selected pack.** A dry run with `--control-pack financial-strict --json` emitted a machine-readable `next_action` that wrote `default` when executed — the agent-mode promise is that following the emitted argv completes the setup that was asked for. `_requested_setup_flags` now carries it, so the dry-run advance, the unresolved-scope candidates, and the scoped-refusal rerun all repeat it; route identity follows, because the identity hashes the selected action. It is repeated only when it is not the default, which is what makes the recovery faithful rather than merely verbose: the manifest written is byte-identical either way, while repeating it unconditionally would rewrite every existing route string, and every route identity, to say what the previous string already meant. Both early-validation routes — the unknown pack and the invalid `--agent-instructions` selector beside it — built a bare `init …` command that dropped `--workspace`, `--write`, and `--json`, so following the recovery ran a dry run against the wrong directory and printed prose to a JSON consumer. One `_recovery_command` builds both. **[P1] `run_id` did not cover the pack.** The identity is hashed before `effective_policy` is assigned, so two manifests selecting `default` and `financial-strict` hashed identically wherever neither produced a control finding — different enforced policy, one run identity, and a `STABILITY.md` claim that was false as written. `_run_id` now hashes the pack's id, version, **and canonical obligations**, resolved through the same single reader every other consumer uses. Obligations rather than the name alone: a release that changes what `default` requires changes what the run means. **[P2] "Cannot compare" is not "no weakening".** A base report naming a pack this build cannot resolve returned an empty delta, so `future-strict → default` read as a clean comparison. It routes to `SHIP-VERIFY-POLICY-BASE-ABSENT` with `kind: control_pack_unrecognized` — the reason code whose whole job is saying the comparison could not be made. **[P2] The high-impact rule invented effects and controls.** One `policy_id` serves `production_operation` and `code_execution`, so recovering the effects from it reported a code-execution action as also operating on production — and the rule row then unioned both obligations and demanded a rollback the actual rule never asked for. Findings now stamp `evidence.control_effects`; nothing is reconstructed from an id. **[P2] A suppressed mandatory control went unexplained.** The release decision keeps it blocking; the summary dropped every suppressed finding, so a report could read BLOCKED with no Control Pack section naming the blocker. One predicate, `is_mandatory_current_control`, now serves both. **[P2] Provenance from a string.** `action_surface.policies[].id` can collide with the engine's grammar, which would make a user's own rule non-waivable and promote their suppression of it into a blocker. Non-waivability is decided from `evidence.control_pack` — which only the engine writes — and the `control-pack:` prefix is reserved at manifest load, the way `SHIP-` already is for check ids. **[P2] Equivalent rules churned fingerprints.** The pack-only `policy_id` embedded the pack name, so switching between two packs that require the same controls for the same effect re-fingerprinted an identical finding and dropped the baseline entry accepting it. The id is now `control-pack:`; where two packs require *different* controls the `missing` rows still differ, so the fingerprint moves for the reason that actually changed. Each fix is guarded by a test that fails when the fix is reverted; the seven perturbations are in the branch's review notes. Report schema renumbered `0.39 → 0.40` after #423 took `0.39`. feat(evidence): pin an answer to its evidence, and name the rung you stand on (#410 increment 4) (#423) * feat(evidence): pin an answer to its evidence, and name the rung you stand on (#410 increment 4) Section E and section G of the evidence-first declaration RFC — the last increment. E is what makes the system safe to run for years; G is what makes every state of an adoption a place with a name and a next step instead of a verdict that reads like a failure. * **`action_surface.actions[].basis` pins an effect answer to the evidence it was given against.** Declarations match by name and nothing ever re-opened one, so a green gate at month twelve could rest on a description of a function that no longer does what it did. Every scan re-derives the digest and compares: equal is complete silence, different re-opens the question as a `declaration_drift` gap that names what the action reads as now and hands over the new pin. The digest covers what the questionnaire already showed the reviewer — the effects the scan *observed* — never the producers that observed them, so a release that adds a heuristic cannot re-open every pinned declaration on every adopter at once. It is also stable across the arrival of the answer, which is the one property that would make pinning worse than not pinning. Additive: unpinned manifests behave as before. * **`environment.target: template`** answers the authority dimension once for a repository that ships to be copied, where asking each action which credential it runs with asks a question nothing can answer. A default and never an override — an action's `authority`, a source's block, and an action's own `scopes:` each win over it, and a source publishing a real credential is challenged by it. Never silent either: every action it answers for is a review concern, so a template repository reaches `review_required` and never `passed`. * **`SHIP-TRUST-MANIFEST-UNPROTECTED`** says whether changing the gate takes a named human's approval, at the one moment that matters — the manifest's own `ci.mode: strict` — and never claims the half no file in a checkout can see. Guidance only: it never becomes a review item and never moves a verdict. * **`doctor` names the adoption rung and only the conditions actually unmet** to reach the next one, published as `adoption` on `doctor --json`. A workspace with no manifest is told it is on rung 0 rather than only that a file is missing. Also closes a divergence this change would otherwise have created: an action's permission list is resolved at one site, and `environment.target` is a third place a reviewed authority can be written, so it is threaded into that resolver rather than re-derived without it by the action lens. Verified with a real `verify --base` run, because `CapabilityFactV1` enforces the two surfaces' agreement only there. Report schema 0.38 → 0.39, packet 0.15 → 0.16, verifier 0.12 → 0.13, capability lock 0.7 → 0.8, capability-lock diff 0.8 → 0.9; all additive, all prior versions frozen, hash-pinned, and read forward. Co-Authored-By: Claude Opus 5 * fix(evidence): a statement about deployment cannot subtract published evidence (#410 review) Review pass on this branch. Five findings, and the first is the one that matters — it was a fail-open, found by the perturbation sweep returning zero failures for a guard I thought I had written. **`environment.target: template` erased a permission a source published.** The reviewed record supplies the *whole* permission list, so applying the repository-wide "nothing here holds a credential" over a tool publishing `oauth2 + docs:read` emptied that action's `required_scopes` — and `SHIP-AUTH-SCOPE-COVERAGE-MISSING` silently stopped seeing anything to cover. Reproduced end to end: flipping `staging` → `template` on an otherwise unchanged workspace dropped a `high` finding. The guard covered an action row's own `scopes:` and stopped there; it now covers anything the source publishes, which is the same rule stated once — a claim about the absence of a deployment may add an answer where there is none and may never subtract evidence something else proved. Pinned by asserting the *whole* finding set under `local` is a subset of the one under `template`, so any future shape that removes a row fails rather than one named check. **A concern counted where it was not the answer.** The template claim is appended whenever the record exists, so an action whose authority was contested would have been counted — and named in the reason sentence — as one that *takes* its authority from the template. It is now counted only where it is the operative answer, and the implication that makes that guard a no-op today is pinned by a sweep over eleven source-auth shapes rather than assumed. **Two spellings, twice.** `ManifestProtection.reviewed` was an alias for `covered` — and the worse name of the two, because "reviewed" is CODEOWNERS *plus* branch protection and this checkout cannot see the second half. And the `basis`/`risk_tags` re-insertion in the proposal template was a second, partial spelling of the field ordering the scaffold renderer already applies to every block. **A lossy paste offered as a faithful one.** The drift row carried the declared `risk_tags` alongside the effect, read back from the resolved claims — but only tags that map to an effect produce a claim, so `risk_tags: [read_only]` would have been silently dropped from a block described as a faithful replacement. The row carries the declared effect, which *is* recoverable, and says why it carries nothing else. Also: the reason sentence now names the sharper concern first (an action on a known unscoped credential, ahead of a repository that is not deployed), the adoption rung joins `doctor`'s input identity because it is published and can move without the manifest text moving, and `docs/diagnostics.md` documents the `adoption` payload. Co-Authored-By: Claude Opus 5 * test(evidence): guard the two mechanisms the sweep found unguarded (#410 review 2) Second review round, driven by walking the loop the increment exists for rather than by re-reading the diff. Four things came out of it. **The loop closes, and now something says so.** Pin an answer, change the code the answer described, take the value the drift row publishes, re-scan: silent again. Every other test here checked one hop; this walks all four through real scans, because the pin is stamped by one surface and compared by another and a unit test holds both ends itself. It also reaches the "nothing is observed about it any more" branch, which only a disappearing annotation produces. **Two mechanisms had no guard, and the sweep is what said so.** Perturbing every new behaviour and counting failures returned zero twice: the scaffold's manifest field ordering had no test at all — a rendering property, so the guard asserts on the rendered text — and the CODEOWNERS trailing-slash rule was covered only by paths where it could not matter. Fifteen perturbations, fifteen non-zero counts. **A pre-existing defect, filed rather than folded in.** The merged block a drift row produces also carries the `declaration_below_inferred_evidence` repair, and applying that repair exactly as instructed replaces the row with a blocking `conflicting_effect_evidence` — the `contradictory` filter reads the reviewer's own `risk_tags` as source evidence contradicting their own `effect`, one branch over from where the same class was already fixed. Measured: 0 of 92 first-answer proposals affected, 7 repair shapes affected. It predates this branch and the fix turns a blocking verdict into a passing one, so it is [#424](https://github.com/ThreeMoonsLab/agents-shipgate/issues/424) with a reproduction rather than a rider here. Co-Authored-By: Claude Opus 5 * fix(adoption): every rung promises what the product actually does (#410 review 3) Third review round, reading the ladder text against the behaviour rather than against the RFC. Three rungs were describing the RFC's intent instead. **Rung 1 promised a delta verdict.** "A verdict on every pull request, covering what that pull request changed" is the RFC's rung 1 — and it needs a backlog auto-baseline that does not exist, and could not exist for semantic gaps, which are deliberately un-baselinable. What a manifest with nothing declared actually produces is `insufficient_evidence` over the whole surface plus an ordered questionnaire, which is worth having and is what the rung now says. Tied to the behaviour by a test that scans a rung-1 workspace and requires the verdict the sentence names — the two cannot drift apart silently. **Rung 3 promised enforcement it cannot see.** "The manifest takes a named owner's review to change" is CODEOWNERS *plus* branch protection, and the rung condition is only the first — the same line `SHIP-TRUST-MANIFEST-UNPROTECTED` refuses to cross, crossed two files away. It now says a CODEOWNERS rule names who reviews the change. **Rung 2's next step asserted a backlog.** "Clear the remaining declaration questions" tells a repository that has already cleared them to do finished work — and nothing at this layer has run a scan to know. It asks for any the scan still reports. Also: the end-to-end loop test now takes its pin from the claims the scan itself published rather than re-deriving it from a synthetic tool, which was asserting that two derivations agree — the thing under test. Co-Authored-By: Claude Opus 5 * fix(evidence): address review — orphaned locks, an erased credential mode, and rungs that claimed too much (#423 review) Nine review comments: three release blockers and six correctness issues. Each has a guard that fails without the fix. **[P1] Existing v0.7 capability locks stopped loading.** The lock bump froze `0.7` without adding it to `_STRUCTURALLY_CURRENT_LOCK_SCHEMAS`, and the historical standard map lost it too — `0.7` had been covered only by the *current constant* key, so v0.8 taking that key silently sent every committed v0.7 lock to a parse error and would have relabelled it as current-standard on the way. Both rows are literal now. The old guard pinned `0.6` and kept passing throughout; the replacement walks the published `docs/capability-lock-schema.v*.json` set and requires every version describing the *current standard* to load, so the next bump cannot repeat this. **[P1] `environment.target: template` erased a source-published credential mode.** `AuthInfo` accepts `credential_mode="service_account"` with no mode/type/scopes; `_source_authority` never reads that field, graded it `unknown`, and the template record replaced it with `mode: none` and no credential mode — taking the capability out of reach of `credential_modes` policy selectors entirely. The guard is keyed off `ReviewedAuthority`'s own fields: the record replaces four, so a source publishing any of the four disqualifies it, and a fifth field cannot be forgotten. No first-party adapter emits the bare shape (MCP and n8n both set `explicit` beside it), which is why the assertion is at the capability fact rather than through a scan that would pass for the wrong reason. **[P1] Rung 3 claimed an enforcement it cannot see.** `ci.mode: strict` is the manifest's own statement; the workflow that runs Shipgate passes its own `ci_mode`, and the generated one ships `advisory`. Promising "CI fails on a blocking verdict" was the same over-claim `SHIP-TRUST-MANIFEST-UNPROTECTED` refuses to make about branch protection, two files away. Rung 3 now states the declared posture and names both limits. **[P2] Four more ways a rung said more than it knew.** A control-only action row (`approval:` and nothing else) answers neither declaration dimension and no longer advances the rung. Rung 1 no longer names a verdict — a structural surface owes no questions and can pass from there. Rung 0's next step is the rendered `init --workspace … --write`, not a bare `init` that is a dry run pointed at the wrong directory. And adoption reads the manifest the user wrote, not the copy with unresolved sources removed, which reported a repository that had answered as one that had not. Making the rungs honest also un-stranded them: they name their own conditions rather than nesting, so a structurally-resolved repository can reach rung 3 without ever having been asked a question. **[P2] CODEOWNERS now fails closed on both halves.** Outside a git checkout there is no pull request for a rule to gate, so a stray file of that name proved nothing and awarded rung 3 on a filename. And GitHub skips tokens that are not owners, which last-wins makes decisive: a narrower rule with a typo *removes* the ownership a broad rule granted, so counting the typo reported the manifest protected by the very line that unprotected it. **[P2] The contract bump reached the prose.** Lock/diff numbers in `STABILITY.md`, `docs/capability-standard.md`, `README.md`, and every shipped skill copy, plus a skill that labelled the packet v0.15 beside a v0.16 link and named v0.37 as the report predecessor. Three parity rules were extended so this class stops recurring: version literals are recognised in their `v0.16` form (the `\bv` boundary is why the packet label was invisible), a line naming a superseded version must name the one it actually superseded, and the capability-lock schemas are covered at all. A wrapped historical sentence now keeps its marker. Co-Authored-By: Claude Opus 5 * fix(adoption): route rung 0's next step through the helper that already knows where to point (#423 review follow-up) Self-review of the review fix. `_init_command` was building the workspace with `Path(config).parent`, which gets the glob form wrong: `-c '/repo/*/shipgate.yaml'` resolves to a literal `*` directory, so the command a rung-0 adopter is handed would create a manifest somewhere that does not exist. `doctor` already has the routing for this — the missing-manifest diagnostic uses `_missing_manifest_workspace`, which walks to the longest non-glob prefix and falls back to the cwd — so this reads it rather than spelling a second one. Co-Authored-By: Claude Opus 5 * fix(evidence): pin the strength of evidence, not only its reading (#423 review 2) Second review pass: two blockers and five correctness gaps. Each fix has a perturbation that fails without it. **[P1] The pin was blind to replaced authoritative evidence.** It digested the effects a scan observed but not their strength, so a tool published with `readOnlyHint: true` beside a `read_only` keyword hint kept the same pin after the annotation was deleted — it still read `read`, from the heuristic alone. `read` is the worst classification to lose it on: a heuristic may never establish read-only (#357), so the safety-sensitive answer survived on evidence that could not have produced it. The digest now covers `(reading, strongest evidence class)`. Strength is taken per *reading*, not per producer, which is what keeps corroboration quiet while replacement moves the pin — both directions are asserted. `observed_readings[].policy_eligible` publishes that half so a consumer can reproduce the pin from the row it is printed on, and the questionnaire marks a heuristic-only reading as such. **[P1] Rung 1 promised pull-request gating from a manifest alone.** A manifest is not a workflow — `init` installs one only with `--ci` — and nothing in the ladder inspects `.github/workflows`. It now describes a gate that *can run here* and points at `--ci`; rung 0's next step does too. The regression is a real `doctor` run in a workspace with no workflow. **[P2] CODEOWNERS, four ways it credited protection GitHub would not honour.** A rule set that owns the manifest but not itself is protection one edit deep, so self-coverage is now required and the finding says which half is missing. A file of 3 MB or more is ignored by GitHub entirely, with no fallback to a lower-precedence location — parsing one credited rules the forge never loaded. `docs/*` matched further-nested files, which GitHub explicitly documents it does not; the discriminator is whether the last segment is a wildcard. And the three locations are now trust roots, so a pull request that rewrites the file deciding protection is classified as touching one. **[P2] Any risk tag advanced the rung.** The tag vocabulary is wider than the effects it maps to: a manifest carrying only `risk_tags: [network_access]` reached "Answer on touch" while the action still reported both `missing_effect_evidence` and `missing_authority_evidence`. The ladder reads the resolver's own `risk_tag_answers_effect` rather than a second copy of the table. **[P2] Three stale prose claims, and the guard that could not see them.** `docs/architecture.md` said new exports use v0.7 and `docs/passed-verdict-contract.md` — the authoritative one — said report v0.38 and packet v0.15. The existing rules read a same-line link or a quoted `_schema_version`; none of these is either. A new rule reads the natural-language forms these documents use ("report schema vX", "Current packet schema vX", "new exports use vX"), skipping a version that is a link's own label. Its markers are read on the matched line only, deliberately: these phrases declare themselves current, and a `prior` one line above is exactly what hid "new exports use v0.7" for a whole bump. fix(reporting): group human-facing findings by subject; recommend only what is missing (#364) (#422) * fix(reporting): group human-facing findings by subject; recommend only what is missing (#364) Human-facing output listed findings flat, ordered by severity, so a scan of four money-moving tools presented one fact as seventeen rows across five check families — and the three-row summary spent all three on the *same* check on sibling tools, never mentioning scopes, idempotency, owners or guardrails. Severity is not the axis a reader acts along: they open one tool, fix what is wrong with it, and move on. `scan` stdout, `report.md`, and the PR comment now read one projection (`core.findings.subject_rollup`) that groups by subject, with severity and blocking status as attributes of each row, a location hoisted to the heading when every row shares one, and every truncation stating what it hid. One selection rule replaces three; before this each surface had its own, so one scan reported a different "top" three ways. Separately, the finding a reader would open first told them to declare a control the same finding's `evidence.missing` says they already declared: the sentence was a per-check literal naming every control the effect obliges. The three built-in control checks now build the evidence, the sentence, and the predicate row from one `missing` list at one call site, so they agree by construction. Where nothing is declared the sentence is unchanged, which is why no shipped sample's `recommendation` moved. Presentation only: `findings[]`, fingerprints, counts, severities, `blocks_release`, SARIF, and the release decision are untouched, and the sample `report.json` goldens are byte-identical. Co-Authored-By: Claude Opus 5 * fix(reporting): the decision decides what blocks, and a shared title is not an identity (#364) Two defects in the rollup, both found reviewing it. `finding.blocks_release` and "the release decision blocks on this" are two different claims, and a baseline separates them: `ci.release_decision` files a policy finding whose debt a baseline has accepted as a *review item* while the flag stays true. Reading the flag printed BLOCKS RELEASE two lines under a verdict that had accepted it. The decision is now the only authority; the flag is consulted only when there is no decision to contradict. The check-id-and-title fallback was applied to every decision item, not only to items carrying neither an id nor a fingerprint. `samples/conductor_agent` emits the same check twice with the same title, so a decision naming one of them marked both — and pulled the other into the summary as well. Co-Authored-By: Claude Opus 5 * fix(reporting): a fingerprint is not an identity, and a sentence needs no holes (#364) Review round two on the rollup. `_DecisionIndex` held ids and fingerprints in one set, so a decision naming one finding also marked its fingerprint collision partner — the pair `assign_finding_ids` appends a discriminator for, and that `_to_item` copies both values from. The index is now three tiers in descending precision, each holding only the items that could not supply the tier above: id, then fingerprint for an item with no id, then check id and title for an item with neither. `missing_control_recommendation` with an empty list rendered `Declare for this financial write action.` It is unreachable — every branch is inside `if missing:` — so it is a wiring mistake rather than a state, and the useful answer to one is the pre-#364 sentence rather than a hole. Raising would abort a scan over a rendering detail. Also drops a constant left dead by the `note_prefix` refactor and the unused `findings=` parameter on `roll_up_findings`. Co-Authored-By: Claude Opus 5 * test(reporting): move the model-level rollup fixtures beside the scan ones (#364) Co-Authored-By: Claude Opus 5 * fix(reporting): one tool, one heading — and stop saying the missing list twice (#364) Review round three. `checks.baseline_integrity` emits findings with a tool *name* and no tool id, so a tool with one of those and one ordinary finding got two headings — `create_refund` and `create_refund [stripe]` — the second-spelling failure `catalog_label_index` exists to prevent, reached from the other side. A name-only finding now joins the tool its name resolves to, where the name resolves to exactly one; an ambiguous name is left unresolved rather than guessed, because filing a finding under a tool it is not about is worse than a heading the reader has to reconcile. `report.md` annotated every row with its recommendation, including rows whose detail already *is* the `missing` list the recommendation is derived from — one fact twice in different words, which is the reason the compact surfaces skip the sentence, and that reason does not stop applying inside a report. The three-surface test now renders all three, and pins what has to agree: each surface shows the first N subjects of one projection, never a different N. Co-Authored-By: Claude Opus 5 * fix(reporting): a name repeated for one tool id is not ambiguous (#364) Ambiguity is two tools answering to a name, not two catalog rows. Marking a repeated row ambiguous pushed a resolvable name-only finding back into its own heading for no reason. Co-Authored-By: Claude Opus 5 * fix(reporting): address review on #364 — six findings 1. **The default PR comment had none of this.** `render_pr_comment` defaults to `capability-review`; the block landed only in `_render_findings_comment`, which the CLI reaches only under `--pr-comment-style findings` — the legacy style documented as available for one more minor cycle. Both styles now render the same projection at the same budget, and a test exercises `render_pr_comment` without overriding `style`. The 6000-char truncation still preserves the agent instruction block, which is now also pinned. 2. **A blocking row could be truncated away.** Rows sorted by severity alone, so a subject with five equal-severity findings whose only blocker sorted last by check id printed BLOCKS RELEASE above three rows that do not block. Blocking rows now sort first — the same order the groups are in, one level down. A heading has to be able to show its own evidence. 3. **Not every `evidence.missing` entry is an absent thing.** `_missing_requirements` writes `{"path", "expected"}` rows for a path that does not exist *and* for one that exists with the wrong value, the actual going to `evidence.observed`. Flattening those said `missing: safeguards.dry_run` about an action that declares `dry_run`, suppressed the adopter's own recommendation, and rendered two policies requiring one path identically. Only a list of plain strings is read as a missing list now — a structural discriminator, not a check-id allowlist. 4. **A name is not a binding.** Resolving a name-only finding through the *current* catalog was the wrong idea: `SHIP-BASELINE-INTEGRITY-*` copies its name from a historical `BaselineFinding`, so a removed tool whose name a new provider now exposes would file a stale entry under a tool it was never about. Binding it properly means carrying `BaselineFinding.tool_id` through to the finding, which moves that check's fingerprint — a baseline-compatibility change, not a rendering one. 5. **A location fell back to `ref` and skipped `location`.** Most adapters populate `ref="agent.py"` + `location="agent.py:5"` and leave `path` unset, so the line was dropped — and four findings on four functions then *shared* a suffix and had it hoisted to the heading as one place. The `.strip()` calls also rewrote identity-bearing values before `display_literal` could preserve them; emptiness is now decided by `has_visible_content`. 6. **`info` was missing from the severity histogram.** A finding the decision names is selected whatever its severity, so an `info` blocker rendered `BLOCKS RELEASE ()`. The display order is derived from `SEVERITY_ORDER` rather than written out beside it. Also: `report.md` decides "this sentence only repeats the row" through `control_phrase`, the same table that renders the paths — asking for the raw path missed every control the sentence renames, so `confirmation.required` printed twice. fix(evidence): rank questions by the ceiling, not the inferred floor (#419) (#421) * fix(evidence): rank questions by the ceiling, not the inferred floor (#419) The declaration questionnaire promised an order — "by how much answering can move the verdict" — and delivered the inverse of it. Each question was ranked by `effect_evidence_rank(conservative_effect)`, the effect the scan had *already inferred* for the action, and a pre-filled proposal is offered on exactly the same condition (#416). The two mechanisms ran off one signal, so every question that arrived with a draft answer outranked every question that arrived blank: the cheapest questions first, the most valuable ones last. On the fifth `adk-samples#1745` walk that put three already-drafted mail tools at Q2-Q4 and `create_sap_sales_order` — the single question that produces both `critical` blockers the moment it is answered `financial_write` — at Q6, behind three drafts a reader working top to bottom confirms first. The ranking was faithful to observed risk. Observed risk is not the quantity the header names. A question is now ranked by the **ceiling** of what its answer can establish: where the scan measured a side effect that measurement bounds the answer from below and the old rank still applies, and where it measured nothing there is no bound at all, so the question sorts above every measured one. An action nothing was observed about is not a low-risk action; it is an unmeasured one, and it is exactly where a human answer carries new information. On the same walk the money question moves Q6 to Q3 and all three drafts move to the end. "Measured" and "could be drafted" are one fact, so they are one predicate: `effect_is_measured` is the gate `propose_effect_declaration` already applied, now read from both ends. A heuristic reading of `read` is deliberately not a measurement — this resolver refuses to establish a read-only action from a heuristic (#357), so such an action is as unproven as one with no reading. Within the unmeasured group nothing was observed, so the tiebreaker is the shape of the action's name in three coarse bands, read with the keyword vocabulary the scanner already owns rather than a second one. That vocabulary is deliberately gated for *evidence* on most source types — a Python function called `create_sap_sales_order` is not proof that it writes anything — and ordering is given no more trust than it needs: the band is inert for every measured action, so it cannot reorder anything the scan did read, and a source-level test pins its call sites so it can never reach a claim, an issue, or a verdict. Two rendering changes follow. The header states the order the file actually uses, and a test renders the file and checks the two against each other. And a blank with no reading to print now says the scan read nothing about that action: the header explains that the top of the file is the unread half, and a block that printed nothing let its silence read as "nothing to see here". Co-Authored-By: Claude Opus 5 * fix(evidence): a proven read is bounded, not unread (#419 review) Four findings from review of the ordering change. **[P1] `effect_is_measured` was the wrong predicate to rank with.** It is a proposal-safety rule: it returns False for every read-only reading, however authoritative, because a pre-filled `effect: read` is the one direction where a confirmed guess loses safety (#357). An OpenAPI `GET`, a trusted `readOnlyHint: true`, and a reviewed `effect: read` are all *proven* reads, and all three read as "nothing to propose" — so ranking with it sent them to the ceiling and let the name band break the tie among them. A structural `GET` named `delete_account` led the questionnaire ahead of a genuinely unknown effect: this issue's own defect inverted, with a repository-chosen name ordering something the scan had established. Ordering now asks `effect_is_bounded`, which is two clauses because the resolver records the two cases differently: an effect status of `declared`, `structural`, or `conflicting` — which is what catches a reviewed `effect: read`, since a declaration leaves no reading behind — or an observed side effect, which is the `inferred` action whose keyword hint reads `external_communication`. A heuristic reading of `read` is in neither: this resolver may not act on it, so the answer stays open. Three regressions cover it, all named to look as mutating as the band can score, and all three fail against the old predicate. `_PendingQuestion` is now seeded from its first action rather than from a zero floor, which a bounded `read` reaches exactly. **[P2] The published ordering contract said the old thing.** `DeclarationQuestionRow`'s docstring is emitted verbatim into the report, packet, and verifier schemas, and it — with `docs/agent-contract-current.md` — still said "highest-acting action first". Both now describe the ranking and state that position is not severity: the action at the top is the one *least* is known about. Schemas and `llms-full.txt` regenerated. No field shapes change. **[P2] The header still excluded one member of its own tier.** "could read nothing about" is false for an action whose only reading is a heuristic `read` — it prints that reading above its block and is still unbounded. The header now describes the tier as what nothing has pinned down, "no effect evidence at all, or only a reading this scan is not allowed to act on", and the rendered acceptance test carries all three shapes: nothing read, a protocol default, and a heuristic read. **[P3] The block note claimed a position it may not have.** `build_declaration_scaffold(..., questions=None)` is supported for older reports and preserves gap emission order, so a blank can follow a bounded question. The sentence claiming placement is now emitted only when the block is actually numbered by coverage; the rest of the note, which says what the silence means, is unconditional. fix(action-surface): one permission list per action (#418) * fix(action-surface): one permission list per action A manifest row that listed `scopes:` and declared no `authority:` block at either site turned `verify --base` into `Internal error` (exit 4) on a legal manifest. The action's permissions were spelled twice and the two spellings disagreed: the action lens took the row's list when it had one and the source's auth scopes otherwise, while `_assess_authority` took the row's list only where a reviewed authority record existed. `CapabilityFactV1._semantic_projection_is_consistent` requires the two to project one list, so where they disagreed, rebuilding a capability fact from a serialized `ActionFact` raised — and that rebuild happens on exactly one route, the MCP capability comparison against a base scan. A plain `scan` never reaches it, which is why no sample and no scan-level test saw this. Declaring authority once per source (#410 increment 3) closed the reviewed half of that divergence by normalizing both manifest sites into one record. This closes the rest: `resolve_action_scopes` is the single derivation, read by the action lens and the authority dimension alike, rather than teaching the capability builder to paper over a disagreement it would then have to keep tolerating. Ten of the twenty-eight shapes in the new sweep raised before it. Because the shared rule is the one the lens already applied, no shape that resolves today resolves differently — every list that moves belongs to a shape that was exit 4. Publishing the row's list on the authority dimension is also what would let `scopes: [crm.read]` erase a `crm.write` grant the source proves, with a `structural` status and nothing raised. The subset rule a reviewed authority is held to now reads the *resolved* list, so a bare `scopes:` list is held to it too and reports `conflicting_authority_evidence` against `actions[].scopes`. Co-Authored-By: Claude Opus 5 * refactor(action-surface): one normalizer, not a second copy of it Review of the resolver caught it introducing a byte-identical copy of the lens's `_normalize_strings` — a fresh second spelling of a rule, in the change whose whole argument is that a rule with two spellings is a defect. Same body, same output, and nothing forcing them to stay that way. The primitive moves to `core/action_semantics.py`, the leaf module both readers already import, under a name that fits both lists an `action_surface.actions` row carries: `scopes` and `risk_tags` normalize identically because comparing them is what decides whether a declaration matches, broadens, or narrows. feat(evidence): authority follows credentials, not functions (#410 increment 3) (#417) * feat(evidence): authority follows credentials, not functions (#410 increment 3) Every action a tool source contributes normally runs with the same credential, and asking for it once per action asks the same infrastructure question N times. That is not merely tedious — it is what breeds the copy-paste that breeds wrong answers, and a wrong authority declaration is the one that makes an unscoped production credential read as `mode: none`. Section C of the evidence-first declaration RFC moves the claim to where the fact lives. * **A new `tool_sources[].authority` block** — `{mode, auth_type, credential_mode, scopes, reason}`, with exactly the mode co-requirements an action row already obeys. They are now one shared rule rather than two copies, so a manifest one site rejects and the other accepts is not reachable. The only difference is where `scopes` lives: an action row keeps its permission list in the sibling `actions[].scopes` field, and a source, having no such sibling, carries its scopes inside the block. * **Additive, and never a weaker statement.** An `action_surface.actions[]` row declaring its own `authority` still wins for that action; a source with no block resolves exactly as before. Both spellings are normalized into one record before anything judges them, so the source block is held to the same conflict rule: declaring `mode: none` across a source whose actions publish an OAuth scope raises `conflicting_authority_evidence` on each action that disagrees, and it still cannot stand in for authority a source publishes ambiguously. * **One blank is one question.** The questionnaire's unit was `(action, dimension)`, which counted one edit as N things to do. A question is now identified by the manifest block that answers it (`answer_path`), so a source of N actions with no authority evidence is one question, one numbered block, and one `evidence_gaps[]` row naming the source and how many actions wait on it. Nothing above the published rows changed: every action still carries the issue and still fails pass eligibility for it. Conflicts stay per action, because each asks a reader of *that* action which claim is wrong. * **Nothing is prescribed where nothing can be written.** The source route is offered only for a `source_id` the manifest configures; a per-scan adapter's source id keeps being asked on its own row. Also fixes a divergence the change surfaced: two derivations of "which manifest site is operative" gave two answers to what an action is granted. An action declaring `mode: none` under a scoped source published the source's scopes as its `required_scopes` while its assessment reported none. Both surfaces now read `reviewed_authority`. Report schema 0.37 → 0.38, packet 0.14 → 0.15, verifier 0.11 → 0.12; all additive, all prior versions frozen, hash-pinned, and read forward. Co-Authored-By: Claude Opus 5 * fix(evidence): one permission list, published and judged the same way (#410 review) Review pass on this branch. Three findings, and the first is the one that matters: the change had two derivations of *what is this action granted?* **Two answers to one question.** The first draft published a source's grant as each action's `required_scopes` without reading the same list as effect evidence, so a source granting `crm.delete` beside a declared `effect: read` had nothing objecting — the #409 failure mode with the assertion moved four lines up the file. Keeping the grant off the action instead turned out to be worse than a style choice: `CapabilityFactV1` *requires* `authority.scopes` to equal the semantic authority's, and one of its two builders reconstructs the fact from `required_scopes`, so the divergence raised `CapabilityFactV1.authority.scopes must project semantic authority` and turned adding the block into an internal error on the base-vs-head path. Whichever site is operative now supplies the whole record, permission list included, and that one list feeds the action fact, the authority dimension, the capability fact, and the effect evidence — so a write-verb permission the manifest says an action requires bounds that action's effect whichever site asserted it. Resolved once per call and passed to both dimensions rather than derived twice. Guarded by the invariant asserted directly across every reviewed-authority shape, and by a real `verify --base` run over a commit that adds the block, because the unit-level builder raises identically on `main` and would not have distinguished new from pre-existing. **A fallback that could never reach its fallback.** A source-wide row's `source_ref` was meant to name the file the source is read from rather than the JSON pointer of whichever action happened to build the row — but the chain began with the caller's `source_ref`, which is the issue's own per-action pointer and is always set for this kind. The branch was dead on arrival; found by reading the rendered rows, not the code. **Surface with no reader, and two stale references.** The normalizer was exported while nothing outside the module read it; a comment named a test that does not exist under that name; and nothing pinned the source-answerable kinds to be a subset of the kinds a declaration can close at all. Co-Authored-By: Claude Opus 5 * test(evidence): pin the fail-open the resolved permission list closes The adversarial declaration from #409, with the grant moved to the source block: `effect: read` on an action the manifest says requires `crm.delete`. Written on the action row this has always been a blocking conflict, and the whole point of reading one resolved permission list is that moving the grant four lines up the file does not make it go quiet. Verdict `blocked`, `0/1` pass-eligible, `conflicting_effect_evidence` raised. Co-Authored-By: Claude Opus 5 * docs(evidence): the resolver's docstring describes the model that shipped A paragraph survived from a draft this branch abandoned — it said the source's grant is deliberately kept off the action's `required_scopes`, which is the opposite of what the code does and of why. `CapabilityFactV1` requires the two to agree. Co-Authored-By: Claude Opus 5 * fix(evidence): join a declaration to the source that produced the action (#410 review) Four review findings on this branch. The first is the one that matters: the join was wrong, in both directions. **A source id is not a foreign key.** `Tool.source_id` is minted by the adapter, and configured ids share that namespace. A `tool_sources` row of type `mcp` calling itself `openai_api` had its reviewed authority applied to the OpenAI API surface — clearing `missing_authority_evidence` for actions nobody declared anything about, and moving them to eligible on that dimension — while a `codex_config` row, whose adapter emits ids derived from the file it read, applied to nothing at all. The dispatcher now records which configured entry each loaded result was produced *for* (exactly, in pass 1; by the `(adapter type, minted id)` pair in pass 2, where the four manifest-only adapters are rejected in `tool_sources` outright and so cannot collide), identity resolution carries it onto the canonical action, and the declaration joins on that. A reviewed `tool_identity` binding merging observations from two configured sources answers for neither — their credentials are separate facts — and the question stays on the action row. **An omitted optional field is not a claim of absence.** A declaration that leaves out `credential_mode` was overwriting a published `service_account` with nothing, leaving the dimension `declared` and pass-eligible while capability policies matching `credential_modes: [service_account]` silently stopped matching. The published value is preserved where the declaration states none; a *different* stated value is still a conflict. **`mode: none` means no credential, including its mode.** Both declaration sites accepted `{mode: none, credential_mode: service_account}` — a fact about a credential the same block says does not exist — and on a structurally complete read action that pair was pass-eligible. **A bump moves the labels, not only the filenames.** Table cells reading `0.37` beside a v0.38 link, "The packet schema is `0.14`" above a v0.15 link, and a `verifier_schema_version: "0.7"` in README and the Claude Code skill that had been stale for several releases. Two parity tests hold them together now: a line linking the current schema must also name its version, and a quoted `_schema_version` must equal what the engine emits unless the line or its section marks it as history. Co-Authored-By: Claude Opus 5 * refactor(scan): key the configured-source lookup on the pair itself A joined string needs a separator, and the id half is repository-chosen — one more thing that can appear inside it. The tuple is the key. Co-Authored-By: Claude Opus 5 * refactor(identity): two names for two different sets of source ids `configured_source_ids` was a local holding the ids adapters *minted*, which is what a binding member selector names. Since the provenance fix that spelling means something else and more precise — the `tool_sources` entry a result was produced *for* — and the two sets differ wherever an adapter derives its own ids. The local is `loaded_source_ids` now; behaviour is unchanged. fix(agent-mode): adopter-facing output stops naming the internal identity model (#329) (#415) Invariant 5 of #327. Running the tool on your own repository for the first time could produce: Duplicate tool observation identity: source_type='google_adk_function', source_id='google_adk:agent.py', native_locator='agent.py#map_salesforce_account_to_sap_bp' Three internal concepts, none of them in the manifest that person wrote, two of them derived — and the one recoverable fact, that a file was listed twice, unstated. #321 fixed the collision behind it; the message shape was a separate and more general problem, and nothing tracked it. Every string whose purpose is to tell a person what to do next now names a file, a symbol, an agent, or a manifest key: console output, the agent-mode `message` / `next_action` / `next_actions[]`, `agent-handoff.json` prose, `fix_task.instructions[]`, and PR comment text. Internal identifiers move to an additive `details` object on the error envelope, where a machine consumer or a bug report can still read them. `report.json` evidence blocks and the tool catalog are untouched — they are the identity model, and they are supposed to be precise. Two defects the audit found beyond the reported one: - A binding gap whose issue named no tool fell back to the derived agent id, so `samples/conductor_agent` shipped "the agent's tool binding graph is incomplete (agent_v1:7205d836…)" as the sentence under its verdict. The report's conservation invariant, which already refused a raw *tool* id in a gap subject, now refuses any derived id, matched by shape — a guard scoped to one kind of identifier passes vacuously for every other one. - `scan` and `verify` each caught `InputParseError` and wrote their own recovery, so a failure with a precise route on one command got generic advice on the other. One resolver now serves both and the assembly path. Review round (seven findings, all reproduced before fixing): - The emitted `edit` action named the CLI's raw `-c` spelling, so a `scan --workspace ` that discovers a sole nested `services/billing/shipgate.yaml` published `path: "shipgate.yaml"` — an unrelated trust root in the caller's cwd. `run_scan` now records the manifest it read on the way out, once, for every caller. - A tool read twice is either a repeated manifest entry or a duplicate definition inside the artifact, and the action carries one path. The check reports which cause it saw and the action follows it, instead of naming both. - `_source_file` returned `source_ref` verbatim, which several adapters write as a locator (`spec.yaml#/paths/…`); an edit action must name a file. - `tool_sources[].id` accepted blank and padded values, so a manifest error was reported as a defect in Agents Shipgate. It is stripped and refused at load — it is the key `tool_inventories[]` and `tool_identity.bindings[]` join on, and both were already stripped where they are declared. - The derived-id patterns matched inside adopter-controlled names, so an agent named `customer_agent_v1_deadbeef` would abort an otherwise valid scan through the conservation invariant. They now require a token boundary and the producer's exact separator, and live in one module. - `unresolved_agent_binding` carries no `agent_id` *because the agent did not resolve*, so falling through to the root labelled `root -> missing_worker` as `root [sdk]`. The root fallback is scoped to whole-graph kinds. - `incomplete_tool_identity` published `path: shipgate.yaml#tool_identity` while its `expects` named `tool_sources`, sending the agent and the human to different sections. - The guard's own AST sweep resolved names module-wide, so two functions both assigning `guidance` were rendered as the concatenation of both and an unrelated `shipgate.yaml` laundered the offender. Resolution is lexical. `tests/test_adopter_vocabulary.py` is the guard: it enumerates the adopter-facing strings four ways — every evidence-gap kind through the real renderers, every published message builder, every hand-written string at an emit site in the modules that produce this output, and the shipped sample artifacts — plus three end-to-end runs, and it pins its own extractor. Closes #329 feat(evidence): ask only what the scanner cannot prove (#410 increment 2) (#416) * feat(evidence): ask only what the scanner cannot prove (#410 increment 2) Adoption stalls at a wall of blanks. The fourth `adk-samples#1745` walk faced `0/12` pass-eligible actions and a report that described the work as "24 semantic evidence gaps" — a symptom count with no order and no finish line, while the same report already held a derived `financial_write` reading for the one tool that moved money. This turns that surface into a questionnaire. * **Effects the scan observed are pre-filled.** `suggested-declarations.yaml` prints the readings behind each effect question and, where they support one conservative answer, offers it in the `effect:` line instead of a blank. A proposal is never weaker than any reading, comes from the closed `ActionEffect` vocabulary rather than from source content, and is offered only where something was observed — a protocol default and a heuristic reading of `read` both keep the blank. * **The file is numbered and counted.** Blocks carry `Question 3 of 5` banners ordered by how much answering them can move the verdict, and the header, the CLI, and the PR comment all render one shared progress sentence. * **A question is not a declaration.** The denominator counts only the `effect`/`authority` of an `action_surface.actions` row, and `answered` is the exact counterfactual: dimensions that gap when the action is re-resolved without its declaration. Also fixes a resolver defect the exhaustive proposal sweep uncovered: the read/side-effect conflict read the manifest's own `risk_tags` as source evidence, so the `risk_tags` repair could not close the row it was printed on whenever the action carried a `readOnlyHint`. Report schema 0.36 → 0.37, packet 0.13 → 0.14, verifier 0.10 → 0.11; all additive, all prior versions frozen and read forward. Co-Authored-By: Claude Opus 5 * fix(evidence): one routing table, one permutation, one derivation (#410 review) Six findings from the review pass on this branch, all of the same two classes this codebase keeps hitting. *Second implementations.* The `kind -> dimension` routing was inverted independently in `declarations.py` and `release_decision.py`, and the effect gap-kind set was spelled a third time as `_EFFECT_GAP_KINDS`. All three now read `ANSWERABLE_ISSUE_KINDS` / `DIMENSION_BY_GAP_KIND` from the one module that defines them, pinned by a test that the table's entries are real gap kinds and that no kind answers two dimensions. `DeclarationQuestion` also carried `readings` and `proposal` that nothing read — a second derivation of what the gap row already publishes — so they are gone. *Permutations that are not bijections.* `_in_question_order` tie-broke on `gaps.index(gap)`, which resolves by value equality on a pydantic model: two rows that render identically both mapped to the first index, duplicating one and dropping the other. It now sorts `(question, original position)` pairs. `_drop_duplicate_blocks` folded byte-identical blocks together while keeping only the first one's question keys, so a question could be numbered by the counter and answered by no block — it now merges the keys and the readings. Also: a conflict row now prints the readings behind it. It is the row with no blank to fill *because* its sources disagree, so what each one says is exactly what the reviewer has to go and reconcile. Co-Authored-By: Claude Opus 5 * fix(release): move the qualification gate with the schema it demands Second review pass. `production_safety_requirements()` pinned `required_report_schema_version="0.36"`, and `_evaluate_receipt` compares it for **exact equality** — so on a 0.37 engine every qualification receipt fails with "qualification report schema mismatch: expected 0.36, got '0.37'", on every case, for a reason that has nothing to do with safety. No test caught it: `test_production_defaults_pin_the_exact_beta_contract` asserts the strata and the thresholds but never the schema. Pinned now by `test_the_qualification_gate_demands_the_schema_the_engine_emits`, which compares the requirement against `ReadinessReport`'s own default, so the next bump cannot leave the gate behind. `benchmark/safety-qualification/README.md` carried the same stale number. Swept the rest of the tree for the class: the only other version literals in `src/` and `scripts/` are the `explain-finding` and `scenario` *minimum* supported versions, which are floors and correctly unchanged. Co-Authored-By: Claude Opus 5 * fix(scaffold): a questionnaire whose numbers run in order (#410 review) Third review pass, both found by rendering a workspace that exercises every shape at once rather than by reading the code. *The numbers did not run in order.* Blocks were emitted first and comment-only entries after them, so an action whose effect question has no blank to fill — a source conflict — printed a file numbered 2, 3–4, 5–6, 1. Numbering that does not run in order is worse than no numbering. The two renderings are one queue and are now interleaved by question number. *A long tool name ran the banner off the line.* A subject is repository-controlled and unbounded; one 60-character name left a ragged heading with a trailing space and no rule. The banner now elides to fit, which is safe precisely here: it is a heading, and the block directly beneath it carries the exact `tool` and `tool_id` a reader acts on. Pinned by a test that also asserts no generated line ends in whitespace. Co-Authored-By: Claude Opus 5 * test(scaffold): make the assertion guard reach the values it now allows Fourth review pass. `test_no_shipped_template_asserts_on_a_humans_behalf` still passed after this branch introduced pre-filled effects — for the wrong reason. `_shipped_templates` built its tools without a semantic assessment, so `propose_effect_declaration` was never reached and every effect template kept its sentinel. A guard that stops covering the thing it guards is the "a test that can skip in CI is not a guard" class again. The invariant has genuinely changed, so the test now says the new truth and guards it: every scalar is still a sentinel or a selector, except `effect` and `risk_tags`, which may carry a proposal — and a proposal must be a value from the closed `ActionEffect` vocabulary (never source content) and must arrive beside the `observed_readings` that justify it. A resolved tool is in the fixture set, and the test fails if no template exercises the proposal path at all. Verified by perturbing what it guards: writing an off-vocabulary value into the template fails on the vocabulary clause, and suppressing the readings fails on "no shipped template exercised the proposal path" — no readings, no proposal. Also confirmed by inspection that nothing applies a `declaration_template`: its only consumers are the advisory scaffold writer and a count in the verify prose, and `auto_apply=false` / `requires_human_review=true` are unchanged. Co-Authored-By: Claude Opus 5 * docs(stability): state the v0.37 contract for the new fields STABILITY.md is the external contract for what each published field means and what may change under it. The two additive fields were documented in the agent contract but not there: `declaration_questions` (what counts as a question, why `answered` is a counterfactual rather than a declaration count, and that nothing gates on it) and `observed_readings` with the pre-filled template (that a proposal comes from the closed vocabulary, is never weaker than a reading, is offered only where something was observed, and is still applied by nobody). Co-Authored-By: Claude Opus 5 * docs: correct the two published claims this branch made false Fifth review pass, over the prose rather than the code. `docs/engineering/insufficient-evidence-cold-start.md` stated, in the present tense, that "every human-owned value stays ``". That is now false for exactly one field. It is an engineering-history doc with rounds, so the 2026-07 paragraph keeps its account and gains a pointer, and a *Third round* section records what changed, why a proposal is safe mechanically rather than editorially, and the two defects that surfacing the proposals exposed — the manifest allowed to contradict itself, and the guard that stopped guarding. `docs/mental-model.md`'s artifact table promised the file asserts nothing; it now says what a pre-filled `effect:` is and what still makes it inert. The rule the third round adds: a progress counter must be measurable on both halves. Inventory and `agent_bindings` declarations are human answers too and are deliberately outside the denominator — there is no counterfactual for them, and counting what can only be measured on one side is how a progress bar starts lying. Co-Authored-By: Claude Opus 5 * fix(scaffold): a merged block never names one dimension twice Final read of the assembled file. A block that answers two questions of one dimension rendered "Questions 2-3 · effect, effect", which reads as a rendering fault rather than as two questions. Reachable only through the defensive duplicate-block merge, but the banner is the line a reader trusts to tell them what a block is for. Co-Authored-By: Claude Opus 5 * fix(evidence): do not ask a question a declaration cannot close (review) Addresses all five review findings on #416. **1 [P1] `partial_authority_evidence` advertised an unreachable finish line.** The resolver preserves that issue whenever the *source's* authority evidence is ambiguous or incomplete, whatever the manifest declares — "reviewed authority cannot replace ambiguous or incomplete source authority alternatives" is a deliberate safety property, so the fix is not to weaken it. Reproduced exactly as reported: an MCP tool published with scopes and no auth type asks one authority question, and writing the exact scoped block the scaffold requests leaves the counter at `0 of 1 answered`. It is now excluded from `ANSWERABLE_ISSUE_KINDS` and routed to `provide_source` with no declaration template and an instruction naming the source shapes that close it. The same defect one branch deeper, found while writing the round-trip test: `conflicting_effect_evidence` is raised about either surface, and only the branch the resolver attributes to `action_surface_declaration` is answerable — a server publishing both `readOnlyHint: true` and `destructiveHint: true` contradicts itself and no declaration touches that. Added the invariant test the review asked for, over every kind and both branches, and verified it by perturbation: re-adding either kind fails it. **2 [P2] A recognised prior packet version had its verdict rewritten.** The downgrade was gated on `legacy_version` — "is this a version I recognise" — rather than on "is this before v0.8", the set `_upgrade_semantic_coverage_v08` documents itself as being for. This predates the branch: on `main`, a stored v0.8-v0.12 `passed` packet already loaded as `insufficient_evidence`, explained by a claim about history false of it; adding v0.13 to the ladder made it reach the immediately previous version. Now scoped to v0.1-v0.7, with a test over the post-v0.8 versions; the existing v0.1-v0.7 downgrade test is untouched. **3 [P2] Protocol defaults were folded into observed evidence.** `effect_readings` grouped by effect and OR-ed the `observed` bit, so an unannotated MCP tool with a `writes_data` hint produced one `observed=True` row carrying `mcp_protocol_default` among its sources — printed under "what this scan read this action's effect as", which is exactly what a default is not. Grouped by `(effect, observed)` now; the proposal reasons over values and is unchanged. **4 [P2] Non-contiguous numbers rendered as a range.** Ordering put dimension before tool id, so two canonical tools sharing a display subject interleaved and the block merged for one owned questions 1 and 3 — announced as "Questions 1-3", claiming the other tool's. Ordering is tool-major now (the dimension order still holds within an action), and the banner renders exact lists (`Questions 1 and 3`) so the renderer cannot lie if ordering ever changes. **5 [P2] Sample goldens were relabelled, not regenerated.** All four are regenerated with real scans; `report.md` is byte-identical, so this is purely the missing v0.37 members. Added a structural-drift guard comparing the set of field paths in each golden against a fresh scan — values legitimately differ between runs, a missing member never does. Verified it catches exactly the reported case. Plus the documentation cleanup: the Release Evidence Packet heading in `AGENTS.md` (mirrored into `llms-full.txt`), and a stale mixed version list in `docs/agent-contract-current.md` — two of its five numbers were already wrong — replaced by a pointer to the one table that carries them. Co-Authored-By: Claude Opus 5 * fix(evidence): point a source-owned row at the source (re-review) Three executable-remediation defects and a stamp, all the same shape as the first round: a published repair that cannot close the row it is printed on. **1 [P1] The path contradicted the sentence above it.** `partial_authority_evidence` emitted `provide_source` and said "a reviewed action declaration cannot close this row" while `_semantic_gap_path` fell through to `shipgate.yaml#action_surface.actions[...]` — the target coding agents and the `Fix at …` line actually consume. It now resolves the tool's own published evidence (`tools.json#/tools/0`), or `None` when nothing openable is known; the row stays addressable through its rerun command either way. **2 [P1] Provenance was dropped before the action was built.** `_semantic_coverage` passed `issue.source` as a display `source_ref` but not to the action builder, so a `conflicting_effect_evidence` the resolver blames on `tool_source` still published every effect value, the manifest route, and "add a conservative reviewed action declaration" — and adding `effect: destructive` leaves the identical row. `issue_source` is threaded through, and the source-owned branch routes to the source with no template and no effect vocabulary. The declaration-owned branch is unchanged. One predicate now decides both: `is_declaration_answerable(kind, source)`. Counting a row the published repair cannot close, and publishing a declaration for a row the counter knows is unanswerable, are the same defect from two ends, so they may not be judged by two spellings. **3 [P2] The other manifest-owned positive-risk surface.** The source-conflict exclusion covered `action_surface.actions` but not `risk_overrides.tags`, which reaches the effect dimension as `risk_hint:manual` with basis `reviewed_declaration`. A reviewed `code_execution` tag on a tool published with `readOnlyHint: true` was reported as the source contradicting itself, and declaring the matching effect and risk tag could not clear it. Manifest ownership is decided by both routes now — and `_validated_hint_basis` grants `reviewed_declaration` to no source other than `manual`, so tool-published content cannot reach it. **4 [P3]** `docs/report-v1-consolidation-rc.md` runtime stamp. Tests assert the fully rendered action — kind, `path`, `accepted_values`, template, and the projected reason — not just the kind, plus a round trip showing the conservative declaration really does leave the conflict standing. fix(evidence): a declaration cannot discharge a category it does not cover (#409 follow-ups) (#413) * fix(evidence): a declaration cannot discharge a category it does not cover (#409) Three follow-ups to #411, each a defect that shipped with the monotone declaration rule. **Effects are risk-ordered; their obligations are not.** `claims_above_declared_effect` compared `_EFFECT_RANK` alone, so declaring `financial_write` over an inferred `external_communication` read as escalation and stayed silent. `financial_write` obliges approval, audit, and idempotency — but not confirmation, which is exactly what communicating outward requires. Reproduced against `main`: `pass_eligible_actions: 1 of 1`, `gap_count: 0`, and no `SHIP-ACTION-EXTERNAL-COMMUNICATION-AUDIT-MISSING`, while the external-write risk tag sat untouched in the same report. A declaration now accounts for an observation only when it ranks at or above it *and* obliges at least that observation's built-in controls. The obligations move into `BUILTIN_EFFECT_OBLIGATIONS`, pinned to the four inline control branches by a test that walks each entry through a real scan rather than by a refactor of gate-critical code. Coverage also reads every policy-eligible claim instead of the `effect` field alone: `risk_tags: [financial_action]` produces a policy-eligible `financial_write` claim and applies the financial-write controls, so a heuristic reading the same effect is already accounted for — which is also the remedy for the cross-category case, and why the two `escalation is silent` fixtures now carry a `risk_tags` entry. **The two comparators disagreed.** `_non_authoritative_effect_escalation_support` still compared `ACTION_EFFECT_RANK`, which orders `write` and `privileged_data_access` opposite to `_EFFECT_RANK` — so a declaration could read as covered by the declaration rule and raise `mixed_policy_evidence` on the same action, a verdict no override could reach. Both now call `declaration_covers`, which requires *both* orders to agree. Picking one would have loosened an existing gate path or contradicted the #409 rule; the conjunction is a strict superset of the predicate it replaces, so nothing that gated before stops gating. **Four published schema documents were mutated in place.** `declaration_below_inferred_evidence` was written into `packet-schema.v0.12`, `verifier-schema.v0.9`, `capability-lock-schema.v0.6`, and `capability-lock-diff-schema.v0.7` while each kept its version identifier — and the two capability-lock documents had no successor version at all, so a consumer pinned to any of the four rejects artifacts that document is supposed to describe. `generate_schemas.py --check` cannot catch this: it proves committed == generated, never that a content change moved the version. All four are restored byte-for-byte; the capability lock advances 0.6 -> 0.7 and its diff 0.7 -> 0.8; and a lock written under 0.6 is advanced on read rather than rejected, which the bump would otherwise have done to every committed `capabilities.lock.json` (the normalizer handled only 0.1-0.4). Also: the row names every uncovered observation rather than the strongest alone, so a reviewer is not asked to acknowledge evidence the row never showed them; and the published remedy is true of its state — `strongest_effect_above_declaration` becomes `effect_remedy_instruction`, which names the missing controls instead of telling a reviewer to raise an effect that already outranks the observation. Co-Authored-By: Claude Opus 5 * fix(evidence): the published repair closes the row, and every override reaches the reviewer Six review items on #413. **[P1] The emitted repair was not guaranteed to close the row.** It named the strongest uncovered observation, so with both a `financial_write` and an `external_communication` reading a reviewer applied the exact edit the row asked for and got the same row back — `financial_write` does not carry confirmation. It also fell through to "declare the `write` controls" for an effect that obliges no built-in control at all, naming nothing to do. `effect_remedy_instruction` becomes `effect_repair`, which derives the route from the **full** uncovered set. A raise is advertised only when one *observed* effect covers every uncovered observation and the value already declared — drawn from the observations themselves, since rank alone would nominate an unrelated effect that merely sits higher, and raising must not quietly drop the reading the reviewer chose. Otherwise the row publishes the `risk_tags` route: a declared tag is policy-eligible evidence, so it both accounts for the observation and makes that category's built-in controls apply, which is the outcome the uncovered obligation was asking for. The structured action follows: `accepted_values` carries risk-tag values on that route, the scaffold template includes the filled `risk_tags` list beside the override block, and the scaffold's own hint names both routes instead of only `effect`. `_GAP_PHRASE` and the row's `why` drop the "weaker" framing for "does not account for". A new test applies the published repair to every gapped declared/observed combination — 243 of them — and asserts the row is gone. **[P2] A multi-effect override hid its secondary observations.** The claim recorded two `overridden_claim_ids` but one `overridden_effect`/`sources` pair, so the second reading an override waived vanished from `acknowledged_overrides`, the PR projection, and the packet the moment it was acknowledged. The resolver now records `overridden_observations` and the projection emits one reviewer row per suppressed observation. The singular pair stays readable so an older report still projects one row rather than none. **[P2] The frozen-schema byte guard skipped in normal CI.** It shelled out to `git show 0c5f40fc~1`, and the main job checks out with `fetch-depth: 1`, so all five freeze cases skipped exactly where the invariant must hold. Replaced with checked-in sha256 hashes, plus a test asserting every frozen document has one. **[P2] The obligation-table drift test protected one direction.** `issubset` still passed when a branch gained a control the table omits, and skipped an effect deleted from the table — the unsafe direction, since `declaration_covers` can then discharge an observation whose new control never applies. It now iterates every `ActionEffect`, scopes to the built-in control checks, and compares the exact set. Verified by perturbing the table and watching it fail. **[P2] The prior lock schema was tied to a mutable constant.** `"0.6"` mapped to `CAPABILITY_STANDARD_VERSION`, so the next standard bump would have relabelled every v0.6 lock as current and compared it instead of requiring re-export — the opposite of the adjacent comment. Pinned to the literal `"0.5"`. **[P2] Examples and canonical docs disagreed with runtime.** The lock-diff example embedded lock refs at `0.6`, a tuple the CLI cannot emit because the loader advances a v0.6 lock and `_lock_ref` stamps the current version; the lock example attributed v0.7 output to `0.16.0b6`, which emitted v0.6. Both corrected, and `STABILITY.md`, `README.md`, and `docs/architecture.md` — stale since before this change at v0.4/v0.5 and v0.5/v0.6 — now advertise the current pair. Added a runtime-semantic example assertion rather than JSON Schema validation alone. fix(evidence): a declaration weaker than its evidence is never silent (#409) (#411) * fix(evidence): a declaration weaker than its evidence is never silent (#409) Increment 1 of the evidence-first declaration RFC (#410): the monotone declaration rule. Declaring `effect: read` on a tool this scanner itself tagged `external_write` was accepted with zero findings. The pre-existing `inferred_effect_only` gap was closed by the very declaration that contradicted the heuristic which raised it, the action went pass-eligible (0/12 -> 1/12), and the contradicting `risk_tags` stayed in the same report with nothing joining them. The contradiction check already existed and was correct for the claims it could see: `_assess_effect` admits a claim into `contradictory` only when `policy_eligible`, which `domain.py` grants only to typed, high-confidence bases. Heuristic risk hints are deliberately excluded -- a heuristic must never *drive* policy (#357). But one flag governed two different powers. Driving a verdict heuristics rightly cannot. Challenging a human assertion they should: a declaration sitting *below* an observation is not the heuristic gating anything, it is a human statement contradicting something the scan saw. Effect declarations are now monotone. Escalation stays silent. De-escalating past a non-policy-eligible inference raises `declaration_below_inferred_evidence` (report schema 0.35 -> 0.36), a review-level evidence gap naming the declared value, the inferred value, and the hint that produced it. The declaration stays operative and the row never blocks, but the action is not evidence-backed-pass until it is answered -- by raising `effect`, or by the new `action_surface.actions[].override` block naming the `evidence` checked and the `reason` it does not apply. An acknowledged override is accepted (the action is pass-eligible again) and reported as one semantic review concern, so a run carrying one can never read `passed`. Two guards keep the rule from firing where nothing is being asserted. A declaration the source itself corroborates is not challenged (`support.search_kb` declares `read` beside `readOnlyHint: true`, and a keyword reading `financial_write` out of "refund" in its description is the weaker signal). Corroboration never counts the manifest row's own restatements of itself, so a declaration cannot close its own gap. And an override can never silence `conflicting_effect_evidence`: where policy-eligible evidence outranks the declaration, the blocking conflict is unchanged and no acknowledgement attaches. Closes #409. Co-Authored-By: Claude Opus 5 * review: corroborating source evidence names the row, it does not exempt it Self-review round 1 on #411. The first draft exempted a declaration the source itself corroborates: `support.search_kb` declares `read` and carries `readOnlyHint: true`, so why make a reviewer defend a protocol annotation against a keyword? Because this resolver already refuses to pass on that annotation alone -- with no declaration the same tool is `inferred_effect_only` and not pass-eligible, precisely because a hint outranks it. A declaration that merely restates the annotation must not buy what the annotation could not, or #409's hole moves rather than closes. Worse, the corroboration was drawn from content the tool source supplies about itself, which is not conditioned on `tool_sources[].trust`: an MCP server could assert `readOnlyHint: true` about a destructive tool and a `read` declaration would sail through with zero rows. Corroboration is now named in the row instead -- "source evidence agrees with the declaration (mcp_annotation)" -- which is what makes the override one line to write rather than an investigation to open. Measured cost of dropping the exemption across the whole sample and fixture corpus: two rows, both true, both now carried as overrides in `samples/support_refund_agent` and `samples/ai_generated_refund_pr`, so the shipped samples are the worked example. Also from the review: - An `override` written where the conflict is with policy-eligible evidence was silently discarded. That is the field a blocked reviewer reaches for first, and they got a byte-identical message on re-run. `conflicting_effect_evidence` now says the override does not reach it and why. - `_decision_reason` composed its sentence from three known `reason_counts` keys, so a future review-concern reason would vanish from the sentence whenever an authority or override phrase was present. The unnamed remainder is now counted explicitly. - `SemanticCoverageDecision`'s docstring -- published verbatim as the `description` in the generated report schema -- and the agent contract still described `review_concern_count` as authority-only. - The comparator block now lives entirely inside the guard that gives its locals meaning, dropping a type-checker-only conjunct that read as a real possibility. - The #409 rule, the trust boundary, the ignored override, and the deliberately accepted stale override are each pinned by a test. The `test_weaker_write_declaration_with_inferred_financial_effect_is_ie` canary now asserts the stronger property: `create_refund` declares `write` against a `financial_write` heuristic, so the discarded escalation is on the record instead of being carried only by the policy-evidence gap. Co-Authored-By: Claude Opus 5 * review: both conflicting branches name an ignored override Self-review round 2 on #411. The read-versus-side-effect conflict is the same user error as the ranked one — a reviewer reaches for `override`, the resolver discards it — so both branches now route through one message builder instead of one of them going quiet. Reworded so the sentence is accurate for both: the conflict is *in* policy-eligible source evidence, which is true whether it outranks the declaration or splits read from side-effect. `docs/passed-verdict-contract.md` enumerates what keeps human review mandatory and did not mention that an acknowledged override joins `unscoped`/`ambient` authority on that list. Swept every sample and fixture manifest for rows the stricter rule newly raises: zero, after the two overrides added in the previous commit. Co-Authored-By: Claude Opus 5 * review: the instruction names the effect it asks for Self-review round 3 on #411. The `Improve evidence:` line read "Raise action_surface.actions[].effect to the inferred effect this row names" -- and named it nowhere the user could see. Every human surface renders the short per-kind headline (`evidence_gap_headline`), not the row's `why`, so the one artifact carrying `external_communication` was report.json. A published next step that cannot be executed from what it publishes is the #362 failure mode. `expects` now names the value: "Raise action_surface.actions[].effect to 'external_communication', or acknowledge the difference with an override ...". That also settles what `accepted_values` is the vocabulary for -- the field `expects` names -- which was the other half of the complaint. Deriving the value twice is this codebase's recurring defect class, so the comparison lives in one function. `claims_above_declared_effect` is called by the resolver (which needs the claims, to name their sources) and by `strongest_effect_above_declaration`, which the release-decision projection calls for the value alone. A test asserts the instruction and the reason name the same effect. `action_surface_declaration` was spelled eight times across two modules and is now load-bearing for the projection, so it is `DECLARED_EFFECT_SOURCE`. Carrying the value on `SemanticIssue` instead would have added a published field to every semantic issue in the report schema; the CLAUDE.md surface-discipline gate says to default against that when one shared function does the job. Co-Authored-By: Claude Opus 5 * review: consume the acknowledgement, and project each exception Addresses the four findings on #411. [P1] Consume the acknowledgement in policy applicability. Policy applicability asks exactly what the override answers -- "does the higher heuristic effect apply here?" -- so leaving the acknowledged claim unresolved there traded `declaration_below_inferred_evidence` for `mixed_policy_evidence`: the reviewer followed the row's own instruction and landed on a differently-named `insufficient_evidence`. The override claim now carries `overridden_claim_ids`, and `_non_authoritative_effect_escalation_support`, the action-policy predicates, and capability-policy matching all read that one authored list rather than re-deriving the comparison. The acknowledged fixture reaches `review_required` with zero policy gaps, asserted end to end. [P1] Preserve each override in the reviewer projection. A count is not a review surface. `semantic_coverage.acknowledged_overrides[]` names the action, both readings, the hint source, any source evidence that agrees, and the human's evidence and reason. The packet's section 1 and the PR comment (`SHIP-ACTION-EFFECT-OVERRIDE-ACKNOWLEDGED`) render one row per override. Packet schema 0.12 -> 0.13 and verifier 0.9 -> 0.10, both with forward-reading paths for the frozen shapes: the field is absent there and an empty list is the honest reading. [P2] Reject visually blank override evidence. `strip()` leaves U+200B and U+2060 intact, so an override that renders as nothing to the reviewer it exists for validated, suppressed the mismatch, and restored pass-eligibility. Both fields now require visible content, using the repository's own `has_visible_content` semantics -- moved to `schemas/text.py` so the schema layer can use it without importing `core`, with `core.evidence_actions` re-exporting it. [P2] Keep the public manifest schema aligned with runtime validation. `docs/manifest-v0.1.json` is advertised for live editor validation and accepted both an `override` with no `effect` and blank `evidence`/`reason`. The dependency is published as an `if`/`then`, the visible-content rule as a `pattern` generated from the same code-point table the runtime reads (BMP as `\uXXXX` and astral as literals -- `\u{...}` is ECMA-262-with-u-flag only and Python's `re` rejects it outright), and `tests/test_manifest_schema_parity.py` runs twelve payloads through both validators and requires them to agree. `samples/support_refund_agent` and `samples/ai_generated_refund_pr` already carried the two overrides this rule asks for, so the packet golden now shows the rendered rows. Co-Authored-By: Claude Opus 5 * test: pin the property that makes consuming acknowledgements safe Policy applicability, action policies, and capability policies all drop the claims an override names. That is only safe because an override can never name a policy-eligible claim -- the resolver refuses to attach one while policy-eligible evidence outranks the declaration, and the recorded set is filtered on `not policy_eligible`. Asserted once at the producer rather than three times at the consumers. Build(deps): Bump actions/download-artifact from 7.0.0 to 8.0.1 (#377) Bumps [actions/download-artifact](https://github.com/actions/download-artifact) from 7.0.0 to 8.0.1. - [Release notes](https://github.com/actions/download-artifact/releases) - [Commits](https://github.com/actions/download-artifact/compare/37930b1c2abaa49bbe596cd826c3c89aef350131...3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c) --- updated-dependencies: - dependency-name: actions/download-artifact dependency-version: 8.0.1 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> fix(evidence): a gap subject is a label, in every kind (#408) * fix(evidence): a gap subject is a label, in every kind `LEDGER_JOINED_GAP_KINDS`, reasoning that a gap the exclusion ledger never joins may name its subject however it likes. A guard scoped to a set of kinds passes vacuously for every kind outside it, and the policy evidence gaps are exactly such a kind — so every row of `report.policy_evidence_gaps` kept a raw 64-hex canonical tool id in `subject`, in two shapes: `tool_v2_2c9ee6ae…` from the per-finding emitter and `support.search_kb [tool_v2_445a25…]` from the per-action one. `subject` is a display label. `evidence_gap_headline` prints it verbatim into the CLI's `Improve evidence:` line, `_decision_reason`, and the GitHub step summary, so a reader got a digest where a tool name belongs. `samples/support_refund_agent` carried `support.search_kb` in one gap list as both `[support_mcp_tools]` and `[tool_v2_445a25…]` at once. Both emitters now render through the shared `catalog_subject`. Identity is not dropped in the process: review 2 added `subject_id` as the field that carries it and that the ledger now joins on, but `policy_evidence_gap()` had no way to set it and left `subject_id: null` on every row — it takes one now, and both call sites pass the tool id they already hold. The rule widens to every gap kind. Since the join moved to `subject_id` it is no longer load-bearing for joinability at all; it is what keeps `subject` a label, which is the one thing the field is documented to be, and it means a kind added later arrives already conforming. `LEDGER_JOINED_GAP_KINDS` existed only to carry the carve-out and is removed with it. Field shapes are unchanged and `subject_id` already shipped in `0.35`, so `report_schema_version` stays put (STABILITY.md § every gap subject is a label, never a raw id). Regenerated the two affected sample goldens and `support_refund_agent`'s `packet.json`; no `report.md` golden moves. Co-Authored-By: Claude Opus 5 * fix(evidence): resolve every gap label from the catalog, and match ids by shape PR #408 review found three ways the previous commit's claim — that every `report.policy_evidence_gaps` row was migrated — was false. **A third emitter was missed.** `inputs/policy_packs.py` also builds policy gaps, and rendered `subject=f"{name} [{tool.id}]"` with no `subject_id`. Its existing fixture emits `create_refund [tool_v2_6dcebe…]` against a catalog label of `create_refund [api]`. I missed it because the grep I read it from was piped through `head -50`, which cut the file off. **A membership test could not see it.** The guard compared the whole subject against the run's catalog, so an id *wrapped* in a label passed. So does an id that resolves to no catalog row at all: a check plugin is validated on its declared `check_id`, not on tool membership, so it can raise a finding carrying a stale or invented `tool_v2_` that no membership test can recognise. The guard now matches a canonical id by shape, anywhere in the label, and `cli/scan/decision.py` falls through to the check id rather than printing an unresolvable id raw. **The action path invented a second label.** `ActionFact.provider` is `_normalize_token(provider or source_id or source_type)` — it folds whitespace and falls back differently from `catalog_subject` — so a source id of `my api` labelled one gap `create_refund [my_api]` and another `create_refund [my api]` for the same `subject_id`. Neither is a canonical id, so the widened invariant accepted both; only resolving the label through the catalog closes it. All three emitters now go through one index built from the tool catalog (`catalog_label_index` / `tool_label`), keyed by `tool_id`. Each finding has a regression test that fails against the unfixed code: the policy-pack fixture now asserts both fields, the whitespace-provider counterexample asserts the catalog spelling, and the id-shape guard is parametrized over bare, wrapped, and unresolvable forms. The shared ledger fixture now uses a realistically shaped `tool_v2` + sha256 id; the synthetic short id it used before did not exercise the real thing. No golden moves — no bundled sample uses a policy pack or a whitespace-bearing provider. CHANGELOG and STABILITY corrected: the migration covers three emitters, not two, and the rule is shape-based rather than catalog-membership. fix(evidence): account for every narrowing decision (#403) (#404) * fix(evidence): account for every narrowing decision (#403) Five failures across two adoption walks had one shape: a stage computed the right signal, stored it, and did not connect it to the decision. The sharpest instance is a fail-open in the reward-hacking shape this product exists to catch — github/github-mcp-server#3076 adds `delete_repository` (`destructiveHint: true`) to a published MCP server, and the run reported `unbound_tools: 1` beside `gap_count: 0` and `pass_eligible: true`. The checks that would have blocked it are correct; the tool left the analysed surface before they ran. Reports now carry `surface_exclusions` (report schema 0.34 → 0.35): one typed record per subject a stage removed, derived in one place from the facts the decision itself read. `detect --json` and `trigger --json` emit the same record for the stages they own, replacing four ad-hoc spellings of the same event. `accounting` makes each record checkable — `evidence_gap` (a gap row names this subject), `route_blocked` (the stage withheld its verdict), or `not_claimed` (nothing claims the subject as capability). A conservation invariant is enforced at emission: observed == analysed ∪ excluded, every excluded subject is in the ledger, every `evidence_gap` record is backed by a gap row with the same subject, and a subject this change newly excluded can never be `not_claimed`. The gate moves only where a diff proves it should. `binding_surface_diff` gains `added_unbound_tool_ids` — head exclusions minus base exclusions — and a tool in that set raises a `missing_binding_evidence` gap naming it. A pre-existing unbound catalog entry is unchanged: `samples/large_multi_framework_agent` has 58 by design, and gating on those would make declaring a spec self-blocking. `skip` now requires positive evidence. A non-empty change set no rule classified returns `evaluation_status: "unclassified"` with `should_run: null` and a next action routing forward to the scan; an empty change set keeps `no_match`. Trigger catalog 0.3 → 0.4 also adds `TRIGGER-MCP-TOOL-SCHEMA-CONTENT`, which recognises an MCP tool definition by its content rather than by a naming convention the repository never agreed to. Co-Authored-By: Claude Opus 5 * fix(evidence): one spelling for every gap that names a catalog tool Review of the exclusion ledger found the failure it exists to prevent, reproduced one layer up. `partial_binding_evidence` and the binding graph-issue rows spelled their subject as the raw canonical tool id, while `_unreached_tool_gaps` and `_semantic_gap` rendered `name [provider]`. The ledger joined on the second spelling, missed, and recorded a tool the decision had gapped as `not_claimed` — `binding_coverage.gap_count: 1` beside `surface_exclusions.gated: 0`, and the whole suite stayed green. Every tool-scoped gap now renders its subject through the shared `catalog_subject`, and the conservation invariant gained the two claims that would have caught it: no joinable gap may name a catalog tool by raw id, and an excluded tool the decision gapped may never be recorded `not_claimed`. The spelling rule is scoped to the gap kinds the ledger actually joins, so it does not force unrelated surfaces to change for a join that does not exist. Co-Authored-By: Claude Opus 5 * test(evidence): cover the pre-decision stages, and bound the trigger ledger Review pass 2. `build_detect_exclusions` had no coverage at all — the three paths it derives from (a capped walk, a contested scope, a rejected source candidate) were exercised only through fields it does not read. Added a test per path plus the negative case, and verified each against real `detect` output rather than a constructed result. Also bounded the trigger's own ledger at 25 entries rather than the shared 200. When no rule matches, every changed file is unclassified, so those rows enumerate a list the same payload already carries in full under `changed_files` — while the result is embedded verbatim in `verifier.json` and in the Codex boundary payload written to stdout, where a few hundred copies of one identical sentence buy nothing. `total` and `gated` stay exact, so nothing that reads the counts is affected. Co-Authored-By: Claude Opus 5 * test(evidence): state conservation as a property over the whole fixture corpus The invariant is enforced in `validate_semantic_consistency` at emission, so a scan returning at all already proves it. That made the proof depend on which samples other tests happen to scan. This sweeps every bundled manifest and states the property directly — including the half emission cannot check for itself, that a sample with a non-empty excluded set never ships an empty ledger. `benchmark/repos/` is materialized from `samples/` (eight of its nine archetypes are copies, per its README), so covering samples covers the benchmark corpus by construction rather than by a second sweep that would drift from it. Co-Authored-By: Claude Opus 5 * fix(evidence): close the four fail-open routes review found (#403) All eight review findings on PR #404, each reproduced first. [P1] A negative detector no longer discards an explicit capability match. `stop_conditions` won before `run_shipgate`, so a `.snap` file holding an MCP tool schema — invisible to detect's `suggested_sources` globs, which is why this PR added content recognition for it — matched TRIGGER-MCP-TOOL-SCHEMA-CONTENT and was skipped anyway, preserving the #403 fail-open in the pre-adoption flow that supplies a complete negative detect result. The stop is terminal only while nothing contradicts it: a matched run rule is diff evidence the whole-workspace negative never accounted for. [P1] A requested base comparison that could not be performed now fails closed. `binding_surface_diff.enabled == false` conflated "nobody asked" with "asked and could not" — a v0.30 base, a baseline, or a failed verify base scan — so a head scan concluded an unbound destructive tool was pre-existing from a comparison it never ran. `base_comparison_requested` (and `VerificationContext.base_comparison_unavailable` for the verify path) separates them; that state raises one gap naming the unusable base, and ledger rows are `unverified` rather than `not_claimed`. [P1] The tri-state verdict reaches the consumers that act on it. The hooks branched on `not should_run`, turning a withheld verdict into silence; they now read `evaluation_status` and word the two cases differently. `decide-shipgate-relevance.md` teaches the tri-state and the new precedence. `trigger_catalog_schema_version` moved in step across the contract payload, `.well-known`, the local contract render, and the docs — a drift nothing compared, so a cross-surface equality test now does. [P1] A `dry_run` match no longer hides unclassified siblings. Coverage is per changed path: a dependency bump beside an opaque capability file matched a rule whose glob leg covered only the manifest. `TRIGGER-DOCS-ONLY-NEGATIVE` is unaffected by construction — `every_file_matches` only fires when it classified everything, which is the difference between a negative rule and an absent one. [P2] `BINDING_GAP_KINDS` is derived from `AgentBindingIssue.kind` instead of restated; the copy had drifted and omitted `invalid_binding_annotation`. [P2] The `adapter_parse` rows are gone. Every `source_warning` became "part of that input never entered the catalog", which is false of most: `simple_crewai_agent`'s `FileReadTool` is in the catalog, in the inventory, reachable, and high-confidence. No adapter records a typed omission today, and every provable one already reaches the ledger through `surface_completeness`, so the decision loses nothing and the ledger loses a claim it could not support. [P2] The v0.35 schema pins the nested required lists. `surface_exclusions: {}` and a dropped `added_unbound_tool_ids` both validated, so a nominally valid report could erase this PR's evidence. [P2] The cap no longer discards gated rows. Sorting them first was not enough: `rows[:limit]` still dropped one at 201 while reporting `gated=201`. The cap applies to the rest. Co-Authored-By: Claude Opus 5 * fix(evidence): close the review-2 fail-opens and the ledger integrity gaps Ten findings on 120ccce5, each reproduced first. [P1] The GitHub Action reinstated a skip the runtime refused. `trigger_action` read the raw `stop_conditions_fired` bit before the winning verdict, so the capability-content/negative-detect case — the `.snap` shape this PR exists for — published `skip_shipgate` while the runtime said run. It also collapsed both withheld states into `none`, indistinguishable from "matched nothing". It now projects the verdict, returns `withheld`, and `action.yml` exports `trigger_evaluation_status` so a workflow can tell "run the scan" from "repair the input". The CLI no longer claims a stop overrode a published RUN. [P1] The unavailable-base gap advertised a command that removed the comparison. `scan -c shipgate.yaml --format json` is executable against the head, where it drops `--diff-from` and clears the very gap it was meant to answer — and a published command reaches `fix_task.allowed_repairs`, making it a machine-readable instruction to delete the evidence. Both this gap and the pre-existing base-regeneration one now publish no command; the two steps are in `expects`, and `path` keeps the rows addressable. [P1] First adoption was routed into the failed-comparison path. `missing_manifest` means the base was read successfully and has no gate yet — the distinction `safe_recovery` already draws one function over — so asking the adopter to regenerate a base report that cannot exist made adoption over a partially-wired catalog unfinishable without falsely binding unrelated tools. [P1] The unavailable-base state was erasable. Only `unverified => base gap` was checked, so rewriting the row to `not_claimed` and dropping `gated` to 0 passed while the base gap stood. The converse is now enforced. [P2] The ledger joined tools by `name [provider]`, which two catalog ids can share. `EvidenceGap.subject_id` carries the canonical id, and every entry now names the gap accounting for it through an explicit `accounted_by` pointer — one join for three different gap shapes, instead of three renderings assumed to agree. [P2] `gated` was unvalidated: `entries: []` beside `gated: 999` passed Pydantic, the schema, and semantic validation alike. Counts are now checked against the rows in all three. [P2] The cap's guarantee was untrue of two accountings. 201 `route_blocked` or `unverified` rows kept 200 while reporting `gated=201`. `gap_backed` is the count the cap guarantees exactly — those rows carry per-row proof and are never dropped; the other two are one whole-run fact a single row proves as well as five hundred. The published wording says so. [P2] Adapter omissions are recorded again, from `LoadedToolSource.omissions` — a typed fact the MCP loader records at both of its skip branches — rather than from warning prose. An entry that genuinely never entered the catalog reaches the ledger; warnings about tools that did load stay out. [P2] `BindingSurfaceDiff.base_report_schema_version` joins the required list, and `--diff-from` that fails to parse now counts as a requested comparison: whether the bytes parsed is not the same question as whether the caller asked. [P2] `AGENTS.md`, `llms-full.txt`, and `docs/agent-contract-current.md` teach catalog 0.4 and both withheld states; the accounting enum is documented where agents read it. The parity test covers the prose surfaces and the enum now — it named only the machine payloads, which is why they drifted. fix(control): ask for the checkout instead of dead-ending the loop (#397) (#402) * fix(control): ask for the checkout instead of dead-ending the loop (#397) `verify --preview --head ` reads project markers from the working tree, because that is the tree the `init` it recommends would write to. On a pull request whose project directory exists only on the PR branch, that established no project and routed to a human: `must_stop: true`, `command: null`, `allowed_next_commands: []`. One move, no forward path for either actor — and the remedy, derivable from the ref preview was handed, went unsaid. Nothing about the change is in doubt there. What is missing is an input, a working tree holding the commit under review, and producing it is one mechanical step the caller owns. That is what `fetch_base` already exists for, so the route is now `agent_action_required` with no command (Shipgate never writes to a caller's worktree), an `expects` naming the input and a `why` naming the ref. Checking that ref out and re-running the identical command resolves the project and emits the scoped `init`; a test follows the route through to that second envelope rather than stopping at the first. Evidence the change deleted and an unreadable inventory keep their human route: no checkout repairs either. `detect` was the second dead end. Its unresolved-scope escalation published a JSON selector inside prose and no runnable command anywhere in `next_actions[]`. It now carries one exact `init --workspace --write --json` per candidate below the unchanged rank-1 decision, with the `executable`/`args` pair — the shape `init`'s own refusal has always published. Both build the list from one helper, so the two commands an adopter runs in sequence cannot publish different recoveries for one workspace. Attaching those commands surfaced a third defect: `setup_control_envelope` kept `advance_alternatives` even when a diagnostic outranked the advance they carry out, so a workspace whose only agent evidence is two nested manifests published `stop` — "not a Shipgate target" — and then an `init --write` for each of the two. They now ride only with the decision that selected them; a no-op for `init`, whose scope refusal always selects its own advance. Co-Authored-By: Claude Opus 5 * docs(control): scope the candidate-command promise to a settled parse Review of the previous commit found five defects, all in prose and one test; the code stood. `AGENTS.md` and `docs/troubleshooting.md` promised the per-candidate `init` commands on any "unresolved scope". A truncated parse is unresolved too, and outranks that route in both commands: rank 1 is the higher-cap rerun and no candidate commands are published, because the list they would be built from is a lower bound rather than an enumeration. Both now scope the promise to `agent_scope: "ambiguous"` with a complete parse and say where the truncated case goes instead. The preview bullet explained the head-mismatch route with the reported pull request's shape — a directory that exists only on the PR branch — where the route in fact fires for every `--head` that is not the checked-out commit, whatever the worktree holds. It now states the rule and keeps the reported case as the illustration it is. The CHANGELOG cross-reference pointed "(below)" at an entry prepended above it, and the detect/init parity test compared two filtered lists without asserting either was non-empty — a regression in the shared builder drops the commands from both callers at once and would have reported green. Co-Authored-By: Claude Opus 5 * fix(control): state the real reason discovery cannot settle a deleted-evidence scope Splitting `head_mismatch` into its own route left the shared refusal prose behind unchanged, and one of its clauses had been written for the cause that moved out. `deleted_evidence` and `unreadable_inventory` are produced only after the head was confirmed to be this worktree, so telling their reader that "discovery of the current worktree would answer about a different tree" is false on exactly the runs that reach it: a pull request deleting a project's only evidence file, previewed with no `--head` at all, was told its worktree holds some other commit. The honest reason is the one the route's own docstring already gives and the prose did not: the evidence is missing from the tree everything reads, so discovery reports whatever survived as the workspace's single scope and its `init` adopts an agent the change never touched. The root-exclusion test proved its point with an `all(...)` over the emitted commands, which holds when no commands are emitted at all. It now names the two it expects before asserting the workspace root is not among them. Co-Authored-By: Claude Opus 5 * fix(docs): a fetch_base route has no command to perform Recipe 0 told an agent to "perform only the exact coding-agent action and command in `control.next_action`" for `agent_action_required`. A `fetch_base` route carries `command: null` by contract, and this change makes one reachable from `verify --preview` — the first command that same recipe prescribes — so an agent following it verbatim looks for a step that is not there. `AGENTS.md` and `docs/agents/protocol.md` already say "action" and "route"; this line was the outlier. It now names `expects` as the commandless form. `selected_is_advance` was a hand-maintained mirror of "`selected` is the `advance` object", set in two of six precedence branches. `advance` is the only route any branch selects without building a fresh `NextAction`, so identity answers the question exactly — and a precedence tier added later cannot forget to keep a mirror in step, which would silently drop the caller's alternatives. Co-Authored-By: Claude Opus 5 * docs(detect): say where its candidate commands differ from init's Both agent-facing pages called `detect`'s per-candidate commands "the same list `init --write` publishes when it refuses". They are the same list only for a flagless `init`: the builder interpolates the setup flags the invocation carried, and `init` passes the ones it was asked for while `detect` passes none. detect → init --workspace .../alpha --write --json init --write --ci → init --workspace .../alpha --write --ci --json An agent that asked for `--ci`, hit the scope refusal, and reused the command it already had from `detect` would adopt the project with no CI workflow and report success for setup the caller requested and did not get — the loss `_requested_setup_flags` exists to prevent, warned about two bullets away on the same page. Both pages now state the caveat and name the two ways out, and a test pins the divergence rather than leaving it to prose: `detect`'s command ends `--write --json` where the same workspace's `init --write --ci` refusal ends `--write --ci --json`. Co-Authored-By: Claude Opus 5 * fix(control): make every published recovery reachable and terminating (#397 review) Six review findings, four of them P1. Each is reproduced below before its fix. **A revision-expression head walked history backwards.** The step this route asks for moves `HEAD`, so a route spelled with the caller's own expression does not survive it: `--head HEAD~1` names one commit before the checkout and its parent after. Following the route three times moved `HEAD` three times and never resolved. A `HEAD`-relative `--base` re-ranges across the same checkout, which is quieter and worse — the rerun succeeds against a diff nobody asked for. Both refs are now resolved to immutable ids before either is published, and `expects` names a commit rather than a ref. **The requested checkout left the control pointer current.** Preview bound no HEAD identity, so `agent control` — the one refresh entry point — returned the same `current_control_id` and the same unmet-looking request after the caller performed exactly what `expects` named. The pointer now binds the worktree the preview actually read, so the checkout makes it stale and the caller re-runs preview instead of repeating the checkout. **`fetch_base` consumers still meant "fetch a ref".** The adoption scorer accepted only `git fetch`/`git remote update` for a command-less route, so it rejected the checkout this one requires while accepting a fetch that leaves it unmet. The union is embedded in six durable published schemas, so widening it with a new action kind is not available; `fetch_base` instead means "make this input available" in both senses, with the scorer, the Claude Code stop hook, and `docs/agents/protocol.md` updated together. **The ten-item display cap was applied to the routing.** Candidate 11 onward was selectable and unrunnable — `detect --workspace samples --json` finds 22 projects and emitted 10 commands, and the reported repository has 25. The cap belongs to the human summary alone. **Adopted candidates were routed to `init --write`.** A nested `shipgate.yaml` is itself evidence of a project, so adopted directories are candidates too: on this repository's own `samples/`, 21 of 22 are, and every command emitted for them exited 2 on a manifest `init` will not overwrite while `expects` promised a file that already existed. They now route to `doctor --config `. The exception is `--agent-instructions`, which makes `init --write` the advertised refresh and exits 0 — there the `init` route is kept, flags and all. **`detect`'s setup identity was blind to the route it published.** Every emitted command is spelled for the entry point the process came in through, so the same workspace read as `agents-shipgate` and as `/opt/custom/agents-shipgate` published different commands under one `input_id` — the documented cache boundary for the answer. The selected advance and its alternatives are folded in, as `init` and `doctor` already do. Co-Authored-By: Claude Opus 5 * fix(control): keep the printed refusal and the routes beside it in step Three defects found reviewing the previous commit; two of them it introduced. **The human refusal contradicted its own machine routes.** It listed every candidate identically and told the reader to re-run `init --workspace` on the one they were changing, while `next_actions[]` in the same payload routed an already-adopted candidate to `doctor` because that `init` refuses. Reproduced on a two-project workspace: `- p1 (p1)` and `- p2 (p2)` read the same, and picking `p1` exits 2. The list now marks such candidates and the message names the `doctor` route, from the one predicate the routing already uses. **The scorer credited `git restore` as input recovery.** It rewrites working-tree files and leaves `HEAD` where it was, so it cannot produce the commit a checkout request names — an agent doing only that would have had the obligation dropped from the outstanding list. `checkout` and `switch` are the verbs that move `HEAD`, and they are the only ones accepted now. **An unpinnable head was still published with the pinning claim.** When the evaluated head resolves to no commit, the route fell back to the caller's revision expression while the sentence beside it said such expressions re-resolve after the checkout. It now omits the rerun clause rather than printing advice that contradicts itself and reinstates the walk. Reaching that state requires a git failure between two reads, and "unreachable so it does not matter" is what left a stale rationale in place two review passes ago. Co-Authored-By: Claude Opus 5 * fix(detect): print the candidate list from the formatter init prints from Marking already-adopted candidates in `init`'s refusal reached `init` only. `detect` publishes the same routes and prints the same list, and it kept saying `- p1 (p1)` and "init --write refuses here until you name the project directory to initialize" while its own `next_actions[]` emitted `doctor --config .../p1/shipgate.yaml` for that directory. A human who named it ran an init that exits 2. The cause was a second implementation: `_echo_agent_scope` recomputed the candidate line inline instead of calling `_describe_candidate`, which is why a change to one formatter left the other behind. So the fix is consolidation rather than a second copy of the marker — `describe_candidate` now lives in `scope_routing` beside `is_adopted`, and both commands print from it. The adopted-candidate note is conditional in both, and a test pins that a workspace with nothing adopted prints neither line, so it cannot settle into boilerplate. Co-Authored-By: Claude Opus 5 * refactor(scope): require the workspace the adopted marker is read from `describe_candidate` defaulted `workspace` to `None` and dropped the "already adopted" marker when it was absent. Every caller passes one, so the default existed only to let a future surface print this list unmarked — a human summary contradicting the routes beside it, silently and with every test still green. That is the drift the shared formatter was extracted to prevent, so the parameter is required and a new caller has to decide. Co-Authored-By: Claude Opus 5 * docs(scope): stop the module summary promising an init for every candidate The leading paragraph still said both commands owe the caller "the exact `init` invocation for each candidate". The adopted-candidate route made that false, and the function docstring sixty lines down already explains why: a project that carries a manifest gets `doctor`, because `init --write` there refuses a file it will not overwrite. That belief is the one this module was changed to abandon — it is what had 21 of the 22 commands emitted for this repository's own `samples/` exiting 2 — so leaving it in the first paragraph a reader forms their model from is the worst place for it to survive. Co-Authored-By: Claude Opus 5 * test(scorer): let pytest's configured path find the harness The new test inserted the repository root into `sys.path` itself, which `pyproject.toml` already does for the session via `pythonpath = ["src", "."]`. Two mechanisms for one decision, and the hidden failure is the point: drop that config entry and every other test reaching a non-installed top-level package breaks while this one keeps passing on its private insert, so the suite reports the configuration as working when it is not. Co-Authored-By: Claude Opus 5 * fix(control): match every recovery to the input it was asked for (#397 review 2) Five residual cases, three P1. Each reproduced before its fix. **The scorer accepted either family of input recovery for every request.** A checkout obligation was satisfied by `git fetch origin main`, by `git checkout -- AGENTS.md` and `git restore` (which leave `HEAD` where it was), by `git checkout deadbeef` (the wrong commit), and by `git switch main`. The criterion then dropped an obligation the cell never met — and the test added with the previous commit codified the first of those. Which family is required is now read from `expects`, the checkout form requires the requested commit, and the six near misses are pinned in both directions. **The pointer was blind to the evidence preview actually routes on.** Binding `HEAD` catches a checkout but not an uncommitted change, and preview routinely selects a project from untracked files: deleting an untracked `app/agent.py` moved neither `HEAD` nor its tree, so `agent control` kept handing back the stale `initialize app` route. The pointer now binds the overlay of everything that differs from `HEAD`, with the path set derived from the live worktree on both sides — the reader excludes the same reports directory it was pointed at to find the pointer, so the two agree by construction. Only a plan-less worktree pointer takes this path, which no producer but preview emits. **The pins were the part that got truncated.** `why` is capped at 400 bytes and led with the diagnosis, so a 220-character branch name pushed both the checkout instruction and the pinned `--base`/`--head` out of the bounded field, leaving a consumer to re-derive the rerun from its own `HEAD`-relative request — rebuilding the walk this route exists to end. The instruction now leads, and the whole recovery is also in `expects`, which is never truncated. **`.` was selectable and unroutable.** It is a real entry in `agent_project_candidates` and rank 1 tells the caller to choose from that list, but the routing skipped it: `init` at the root is the run that just refused, and `--allow-unresolved-scope` accepts the whole workspace as one scope, a different decision that also adopts every project under it. It gets an explicit human route saying so. **Requested `--ci` did not survive an adopted candidate.** A refused `init --write --ci` writes no workflow, so routing that candidate to a bare `doctor` dropped the request. It now carries the flag on an `init` with `--write` omitted: workflow installed, manifest untouched, exit 0. `--ci` is the only flag that reaches this branch — `--claude-code` implies an `--agent-instructions` selection and takes the full-refresh route. Co-Authored-By: Claude Opus 5 * fix(scorer): read the checkout target from the checkout, not the command line `_moves_head_to` matched `git`, then `checkout|switch` anywhere after it, then the requested commit anywhere at all. In a compound command line those three pieces are free to come from three unrelated places, which is the false positive the split-by-`expects` change was written to close: git log 66f837355087 && npm run checkout-preview -> credited git fetch origin 66f837355087 && ./tools/switch-env.sh -> credited git log 66f837355087 && git checkout deadbeefcafe -> credited None of the three puts the requested commit in the worktree. Git invocations are now parsed — each `git` token, its global options consumed, its subcommand, and the arguments up to the next shell separator — so the verb has to be the subcommand and the target has to be an argument of that same invocation. Co-Authored-By: Claude Opus 5 * fix(scope): mark the root candidate in the lists people read, not only in JSON Giving `.` a route of its own fixed the JSON surface. Both printed summaries still listed it as an ordinary entry — `- . (rooty)` beside a real project — under a caption telling the reader to re-run `init --workspace` on the one they are changing, which for `.` is the run that just refused. The same pair in `detect`. That is the third time the human form of a run has contradicted its own routing, and the second caveat this list needs, so both now come from one `candidate_caveats` in `scope_routing` rather than a second conditional written into each command — writing the two surfaces separately is why they keep drifting. Each caveat is emitted only for the candidates actually present, and a workspace with neither prints neither. Co-Authored-By: Claude Opus 5 * fix(scope): compute the printed caveats from the printed candidates Both callers passed every candidate to `candidate_caveats` while printing the first ten, so a caveat could be raised by a candidate the display cap cut. The lines refer to the marking on the list they sit under — "a project marked already adopted" — and with fourteen candidates whose only adopted one is `apps/p12`, the note sits below ten unmarked entries and a "4 more" line, sending the reader hunting a mark that is not on screen. They describe the printed list, so they are computed from it. Co-Authored-By: Claude Opus 5 * docs(scope): enumerate the routes this module actually decides between The leading paragraph promised "the one command that advances that project" for every candidate and listed two cases. Two more have been added since: an adopted candidate with requested setup gets an `init` with `--write` omitted, and the workspace root gets a human route and no command at all. This is the second time the summary has outlived the routing under it. It now enumerates the branches instead of describing the two that existed when it was written, so a reader does not build a consumer that pairs one command with one candidate and finds `.` off by one. fix(scaffold): carry the vocabulary, and reach the scans that need it (#361, #388) (#401) * fix(scaffold): carry the vocabulary, and reach the scans that need it (#361, #388) `suggested-declarations.yaml` was written only once the binding layer was already closed, so during the two scans where an adopter is most stuck it did not exist — and where it did exist, it handed out blanks whose accepted values lived in `report.json`. Two templates close the two missing stages (#361). A source whose agent lists tool symbols static analysis cannot resolve now raises one `incomplete_surface` row per source carrying the exact `tool_inventories` entry, `source_id` bound to the source it completes, and the inventory skeleton is written with those symbol names pre-filled — recovered through the canonical decoder in `core.source_warnings`, which owns both halves of that prose. Per source, not per symbol: six warnings are one mechanism restated six times, and a repair on each row would put raw loader prose back in the headline grouping removed. Once the inventory is declared, the unbound catalog scaffolds the closed-world `agent_bindings.declarations` row — agent, every catalog tool's exact selector, and observed handoffs pre-filled; `complete` and `reason`, the two judgements, stay ``. Merging it verbatim closes `binding_coverage.gap_count` in one iteration. Past 50 tools it is withheld, never truncated: `complete: true` claims the list is everything the agent reaches. Every blank now carries what it accepts (#388) — the field's `accepted_values` rendered from the gap's own list, so the two artifacts cannot drift, or the shape and co-requirements where the answer is not a closed set. The `agent_bindings.root` block lists the agent objects the scan observed, which also fixes guessing the wrong one: `object` matches the declared name, not the Python variable. Nothing is filled in. The renderer emits comments PyYAML cannot carry, so a test pins its output to `safe_dump` byte for byte, and repository-controlled names are escaped injectively — an agent name holding a newline would otherwise close the comment and write a filled-in root selector for someone to paste. And a pasted unfinished scaffold now says so whatever field it lands in: `complete` accepts only `true`, so its type used to answer first with "Input should be True". Co-Authored-By: Claude Opus 5 * fix(scaffold): make the prescribed cold-start route terminate (PR #401 review) Six findings from review, led by one that mattered most: the route this PR advertises did not terminate. Following it to the end — inventory, binding declaration, effect and authority for every tool — still returned `insufficient_evidence`, because the six unresolved-import warnings stayed on the report and `evidence_below_ie_threshold` gates on their raw count. A repository that did exactly what it was told had nothing left to act on and still could not be gated. A reviewed `tool_inventories` entry naming a source in `source_id` is the answer shipgate itself prescribes for that source's warnings, so those warnings are now withdrawn once the manifest declares the source. Per source, keyed on the reviewed completion relationship — never on tool names. The name subtraction this replaces was wrong both ways: an unrelated source exposing a same-named tool cleared a repair nobody had made, and an inventory that had correctly split a toolset symbol into the tools it exposes never matched the symbol, so its source was prescribed the same inventory forever. Only the warning is withdrawn; the loader's `surface_gaps` entry stays, so extraction confidence is untouched, and an empty inventory still cannot get past `SHIP-INVENTORY-NOT-ENUMERABLE`. Also from review: - `complete:` closes the world over handoffs as well as tools, and the hint said only "tools" — a reviewer could ratify the tool set while silently asserting a downstream agent surface they never looked at. - `handoffs:` is a bare list of names with nowhere to put a source qualifier, so an ambiguous target resolves to neither agent and the block reports an unresolved binding instead of closing its own gap. Withheld entirely, since dropping the handoff would understate a closed world. - The raw-input sentinel check recursed forever on a recursive YAML alias, replacing a structured config error and its agent-mode recovery payload with a stack overflow. - `display_literal` passed Unicode noncharacters through, and PyYAML rejects U+FFFE outright, so an agent name carrying one made the generated scaffold unloadable. Fixed in the escape predicate, so every sink benefits; the encoding stays injective. The integration test now walks every declaration layer and asserts the final verdict, so an unreachable remedy fails the suite rather than a review. Co-Authored-By: Claude Opus 5 * fix(scaffold): bound the sentinel walk, and complete per source correctly (PR #401 re-review) Four of the five findings on d4415abb. The fifth is analysed and deliberately left open — see below and docs/engineering/insufficient-evidence-cold-start.md. - The raw-input sentinel walk guarded only the current branch, which stops a true cycle but leaves an acyclic alias DAG visited once per path: 20 levels of `{left: *prev, right: *prev}` is 1.5 KB of YAML, doubles at every level, and materializes 2**n path strings for one placeholder. The visited set is now traversal-wide, so the walk is bounded by the size of the document; an aliased placeholder is reported at one representative path, which is all rejecting it needs. - A display name two sources share is now withdrawn against once *every* source publishing it is complete. Dropping the name outright was conservative only while some candidate was still incomplete; past that it left an answered warning standing forever. - Both sides of the source-id comparison are stripped. The manifest permits surrounding whitespace, so ' adk ' and 'adk' compared unequal and left an answered warning in place. - The cold-start walk test now consumes the scaffold's own action blocks, reviews send_quote_email as what it is, and asserts the substantive result: all four evidence-gap counts zero, evidence_gaps empty, and `blocked` on a finding about the declared surface. Declaring every effect `read` left a mixed_policy_evidence gap, so the old assertions proved warning withdrawal rather than the claim in their own comment. - The closed-world declarations row is also scaffolded for a partly-read tool list, listing the agent's existing edges alongside the unbound catalog — a row omitting a tool the repository wires to the agent would be false. Not fixed: the mixed local/imported surface. Three defects sit behind it — `incomplete_surface` prescribing an inventory for an unproven tool *set* (and printing a self-referential source_id after the merge), an agent whose list is not exhaustive still presenting a "complete" structural set so the prescribed declaration is rejected as conflicting, and framework partials never being superseded by a reviewed declaration. Each is reproduced. The routing half alone points the reader at a target the next defect rejects, and the obvious fix for that one drops resolved tools out of the analyzed surface. It needs the resolver to separate edge uncertainty from list exhaustiveness, and to attribute partials to an agent — a separate change with its own review. Build(deps): Update claude-agent-sdk requirement (#380) Updates the requirements on [claude-agent-sdk](https://github.com/anthropics/claude-agent-sdk-python) to permit the latest version. - [Release notes](https://github.com/anthropics/claude-agent-sdk-python/releases) - [Changelog](https://github.com/anthropics/claude-agent-sdk-python/blob/main/CHANGELOG.md) - [Commits](https://github.com/anthropics/claude-agent-sdk-python/compare/v0.2.130...v0.2.136) --- updated-dependencies: - dependency-name: claude-agent-sdk dependency-version: 0.2.136 dependency-type: direct:production ... Signed-off-by: dependabot[bot] Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> fix(google-adk): measure extraction confidence instead of assuming it (#393) (#400) * fix(google-adk): measure extraction confidence instead of assuming it (#393) The Python AST path hardcoded `extraction_confidence="medium"` on every tool it produced, and the only code that ever set `"high"` applied to tools loaded from a `tool_inventory` artifact. Every gate tests `!= "high"`, so no ADK repository could reach a pass from source however statically analysable it was: `insufficient_evidence` was not a property of a repository, it was the framework's default first-run verdict. A condition that holds for every input carries no information. It could not tell a toolkit factory from twelve annotated module-level functions, and the remedy it prescribed was transcription — copy the tools Shipgate had just extracted correctly into `suggested-inventory.json`, adding no fact to the system. An ADK Python entrypoint now reports `high` when the adapter can show it read the whole surface, and `medium` with a named reason when it cannot. The proof is scoped to the file: one unresolved construct anywhere holds every tool the file produced, because a fully-resolved agent in a half-resolved module can still reach tools nobody enumerated. `_surface_is_complete` changed with it. An AST source type used to be disqualified outright, which was the same constant one layer down and would have kept `incomplete_surface` open on a proven surface. Membership now poses the question and the adapter's attestation answers it; saying nothing still reads as incomplete, so adapters not yet taught to answer keep their previous verdict. Promoting anything to `high` is where the risk is, so the guards are the bulk of the change: anything reaching `agent.tools` after construction (including through an alias or `setattr`), `Agent(**config)`, and four name-resolution shapes where the flat scope-blind name map answers `tools=[helper]` with a definition the running agent will not use. All eight were fail-open in the first draft. An unaccounted-warning backstop demotes a module whose ambiguity nobody classified, so a future `warnings.append` fails closed by default. Net effect on the reported subject: a fully static ADK project with its actions declared reaches `passed`; one without them is asked for the effect and authority declarations a human genuinely owes. Co-Authored-By: Claude Opus 5 * fix(google-adk): close five fail-open paths in the surface proof (#400 review) Each case below reached `release_decision="passed"` with no evidence gaps while a reachable tool was omitted from the catalog or its interface was represented by a guessed schema. All five are reproduced as regressions. Unresolvable inline wrapper. `_extract_tool_expr` returned unconditionally from the `FunctionTool`/`LongRunningFunctionTool` branch, so a `func=` this module does not define — an import, an attribute, a lambda, or none at all — produced no tool, no warning, and no gap. The agent could call it; the report described a strictly smaller surface and called it proven. Incomplete name-binding model. `def` and `=` are not the only binding forms: parameters are `ast.arg`, classes bind through `ClassDef.name`, and `except ... as`, `case ... as`, and `global` each have their own shape. Collecting only `Name` stores let a parameter named after a module function resolve as that function. Proof now requires exactly one binding occurrence of any kind, at module scope, through a direct child of the module body — and the wrapper and toolset *variable* names are checked too, since those maps are last-write-wins. Module gaps that reached nothing. The finalizer walked only `canonical_function_tools`, so a module whose tools all came from a resolved OpenAPI/MCP toolset recorded `dynamic_agent_kwargs` or `mutable_tool_binding` into an empty loop. Module-scoped reasons now reach every tool the file contributed. Those tools are only ever lowered, never raised: this step cannot promote a tool the adapter did not extract. Reflective access. `getattr(root_agent, "tools").append(...)` carries the attribute name as data and contains no `Attribute` node named `tools`, so it walked past the structural check. `getattr`/`setattr`/`delattr` with a constant `"tools"` are matched on the trailing callee name, so `builtins.setattr` counts. The module-path exemption is also binding-aware now: `from x import agents` followed by `agents = LlmAgent(...)` left the name in the alias map while it referred to an agent. Annotation presence taken as schema fidelity. `_json_schema_type` reads the unparsed string and falls back to `"string"`, so `set[str]`, `int | None`, `tuple[...]`, a Pydantic model, and even `typing.List[str]` all shipped as `{"type": "string"}` on a tool marked high and enumerated. Faithfulness is now established by asking the emitter what it would produce and comparing it to what the annotation denotes, which keeps the two from drifting and covers any spelling added later. The return annotation is checked on the same footing; an absent one stays an honest omission. Also updates docs/examples.md, which still said an AST-only binding graph cannot qualify for an evidence-backed pass. Co-Authored-By: Claude Opus 5 * fix(google-adk): prove framework provenance, not just spelling (#400 review 2) Six more paths where an otherwise-complete scan returned `release_decision="passed"` with high/enumerated evidence and no gaps while the runtime capability surface or schema was incomplete or wrong. All six are reproduced as regressions. Identity merging discarded member incompleteness. `_merge_bound_observations` starts from the primary and copies nothing about extraction across, which is deliberate: promoting the primary's fidelity is what naming a reviewed inventory is for (#386). That reasoning holds only for claims about one tool's own interface. An identity assertion proves two observations describe the same operation; it says nothing about whether the module one of them came from exposes further tools. Reasons are now split — the interface ones still resolve in the primary's favour, the set ones travel with the observation as `extraction["tool_set_proven"]` and cap the canonical tool. Framework symbols were trusted by spelling. `_qualified_name` maps a call back to `google.adk...` through an alias table, so `FunctionTool = replacement` after the import left a foreign factory being read with Google's semantics. Rebinding `LlmAgent` did the same one level up. Agent, wrapper, and toolset callees now have to root in a name bound exactly once, by that import. `from x import *` can rebind anything and is recorded only under `"*"`, so a local `def known(...)` looked singly-bound and proven. Nothing in such a module is provable, and it now says so once. Injected context was identified by name. ADK decides this by the parameter's *type*, with `tool_context` as the name fallback, so dropping everything spelled `ctx` or `context` deleted ordinary model-visible inputs: `def known(context: str, record_id: str)` shipped a one-property schema and called it proven. Only a statically verifiable injection is omitted now. Annotation spellings were trusted without provenance. `from domain import Account as str` made ADK see `Account` where the emitter wrote `{"type": "string"}`. A builtin counts only while the module binds nothing of that name; `List`/`Dict` only when their single binding imports them from `typing`. Reflective mutation had three more spellings. `from builtins import getattr as read_attr` calls the real builtin under a local name, and `vars(agent)["tools"]` and `agent.__dict__["tools"]` reach the attribute through a mapping. The builtin is matched through the alias table; the mapping forms are matched on `vars`/`__dict__` specifically, so an ordinary config dict with a "tools" key is not a mutation. Two fixes fall out of the review's non-blocking notes. An unproven tool *set* no longer routes its remediation by source type — it asked for an inventory or spec the repository had already supplied — and names the construct to fix instead. `_json_schema_type` now recognises `List[...]`/`Dict[...]`, which the new faithfulness check correctly refused to certify as `string`; without it `from typing import List` held a tool at medium for an emitter gap rather than anything in the user's code. fix(discovery): route preview to the changed project and stop reporting a capped walk as complete (#394, #395) (#399) * fix(discovery): route preview to the changed project and stop reporting a capped walk as complete (#394, #395) `verify --preview` on a monorepo emitted `init --workspace --write`, which `init` then refused deterministically. The change-scope resolver had everything it needed and returned nothing: `evidence_dirs` was omitted at the preview call site, and that parameter is what unlocks the weak project markers a `requirements.txt` beside `agent.py` depends on. `detect` built and passed it, so the two commands an adopter runs in sequence answered the same question differently. Preview now builds the evidence itself, with `weak_marker_evidence_dirs`: the directories at or above the changed paths, filtered to the ones whose answer can change anything — no strong marker, one weak marker — and answered by `_score_python_signals`, the same per-file rule `detect` uses, over the files each directory directly holds. `detect`'s own evidence directories are the immediate parents of its evidence files, so the two ask the same question of the same files. 60ms on a 1271-file change set, against 9s for a whole workspace walk. Separately, `_agent_scope` short-circuited on `ambiguous`, so `"unknown"` — the state whose purpose is to say the parse was cut short — was reachable only when one or fewer candidates were found, and the cap warning went unprinted on exactly the repositories the cap had cut. Truncation is now computed first and reported beside the verdict as `agent_scope_truncated`, with `workspace_signals.project_root_count` (uncapped, filename-only) to bound the claim. The human summaries, the `init` refusal, and its ranked recovery actions all say the list is a lower bound and name the `--max-python-files` remedy; a capped walk that reached no agent at all no longer reports "workspace does not appear to be an agent project" either. `tools/shipgate-detect.py` carries both new fields (script 0.4.0), pinned by the zero-install parity test. Co-Authored-By: Claude Opus 5 * fix(discovery): settle what the preview probe cannot establish, and stop every terminal negative from a capped walk (#399 review) Four P1 gaps and five follow-ups from the review of 45b57ae. **P1 · the implicit workspace scope was uncounted.** `_project_root_count` counted marker directories, but `_agent_project_candidates` also attributes unmarked agent evidence to the workspace as `.`. A repository with one marked sub-project and an unmarked root agent past the cap censused a single root, kept `single` with no truncation warning, and let `init --write` write a root manifest carrying the sub-project's agent name. The workspace is now always counted, in both detectors. **P1 · the truncation guard only reached the human summary.** Every negative-control diagnostic publishes a `stop`, which routing turns into `setup_not_applicable`; `bootstrap` read the same negative as "nothing to do"; the trigger stop block could bury a rule that had already fired on the diff. None of the three diagnostics fires while `agent_scope_truncated` is true, so `detect` falls through to the higher-cap recovery `_detect_advance` already produces; `bootstrap` falls through to `init`'s own refusal, whose structured error it forwards verbatim; the catalog's stop block gained `agent_scope_truncated: false`. A payload missing any key the block reads is now `stop_conditions_evaluated: false` — the evaluator's existing "could not evaluate" answer, applied per key rather than per payload, because absent is not false. **P1 · artifact evidence.** `services/api/{requirements.txt,openapi.yaml}` beside an unrelated strong project: the probe found nothing, preview emitted a root `init`, and that exact command exited `refused_unresolved_scope`. The probe now reads the same three evidence families `_agent_project_candidates` does — framework-attributed Python, `_suggested_sources`, and `_codex_plugin_candidates` — by calling them, not by copying them. **P1 · a deleted agent is not an absent one.** The probe reads the head tree, so a PR deleting the sole `agent.py` beside a `requirements.txt` left the directory looking empty and routed to a root `init` that could adopt an unrelated project. `weak_marker_evidence_dirs` now returns a tri-state: directories it confirmed, and directories it could not settle. Unsettled routes to discovery. Follow-ups: the probe reads the git-aware `_candidate_files` inventory, so an ignored file cannot make preview narrow to a directory `detect` never saw; it spends the same `max_python_files` budget, and exhausting it is undetermined rather than negative; the zero-install `unknown` recovery names `--max-python-files` instead of a rerun that reproduces `unknown`, with a parity fixture that actually truncates; `init` serializes the `workspace_signals` block its refusal quotes; the docs qualify every "`is_agent_project: false` means stop" on truncation, and say the list *may* be incomplete rather than that a project *is* missing. `_candidate_files_matching` gained an optional inventory, which also stops `_suggested_sources` from re-running the git walk once per pattern. Co-Authored-By: Claude Opus 5 * fix(discovery): gate whole-workspace negatives on raw parse completeness, and route each unsettled cause to a recovery that can settle it (#399 review 2) Seven P1s and one P2 from the second review of 7ffce75. **The truncation guard was the wrong field.** `agent_scope_truncated` also requires more than one candidate scope — right for a claim about the candidate list, wrong for a claim about the workspace. A root-`pyproject.toml` repository whose only agent sorts past the cap leaves it false while hiding an agent, so every consumer gating on it still published a terminal negative. `DetectResult` now carries `python_parse_truncated`, the raw fact, and the negative-control diagnostics, `bootstrap`, the `detect` summary, first-look, and the trigger stop block all gate on that instead. **A capped walk now emits an executable retry.** Raising `--max-python-files` needs no decision, so publishing it as prose inside a `human_review_required` route with `command: null` left the only actionable step in a string. `workspace_signals.python_file_total` — the uncapped `.py` count — gives the retry a bound that cannot hit the cap again, and on an `unknown` scope that command leads the ranked recovery: nothing was chosen there because nothing was seen. **Two routes that could not succeed.** The artifact-only and Codex-plugin diagnostics name a root `init --write` and outrank the advance, so on an unresolved scope they published a command that refuses deterministically; both now require `agent_scope == "single"`. And `bootstrap`'s no-surface stop never read `codex_plugin_candidates`, so a plugin-only repository — deliberately `is_agent_project: false` — stopped at detect and never ran init. **The preview probe read a subset of the evidence.** `_agent_project_candidates` counts every framework's `candidate_files`, including the artifact-glob detectors; `requirements.txt` beside an `openai-config.json` was a project detect reported and preview did not. `_collect_glob_hits` now takes a scoped inventory and the probe calls it. Deletion uncertainty covers the same families, by name, since a deleted file cannot be parsed. **One generic unresolved route could not recover anything.** A `detect` at the same cap hits the same cap; a head-only `detect` cannot see evidence the change deleted — it reports the surviving project as the single scope and its `init` adopts an agent the PR never touched. `WeakMarkerEvidence` carries the cause, and preview routes budget exhaustion to a concrete higher-cap command and deleted evidence to a human route with no command. **P2:** `_holds_agent_python` returns whether it stopped with a file unread instead of inferring truncation from a zero remainder — a directory holding exactly `max_python_files` modules was read completely. Co-Authored-By: Claude Opus 5 * fix(discovery): make every capped or removed-evidence route settle the question it was published for (#399 review 3) Six P1s and two P2s from the third review of a14e32b. The pattern in all of them: the route was published but the command behind it could not reach the answer. **`init` ran its own discovery at the default cap.** Following detect's `--max-python-files 1002` found the agent and recommended `init --write`, which re-parsed at 1,000, missed it, and wrote a `CHANGE_ME`/no-tools manifest at exit 0. `init` now takes `--max-python-files`, refuses to write while `python_parse_truncated`, serializes the field, and leads its recovery with the same setup at a bound that covers every Python file — settling the scan and completing the setup in one step. That ordering now applies whenever the parse was cut short, not only on `unknown`: asking a human to choose from a list the refusal itself calls incomplete was the thing to avoid. **A settled scope is not a complete parse.** The artifact-only and Codex-plugin nudges were gated on `agent_scope == "single"`, which a one-project workspace is however early the parse stopped; the nudge then outranked the full-count retry and adopted a truncated surface. Both now require a complete parse as well. **`DetectResult.next_action` still ended at the capped negative.** The CLI overwrites it with the routed action, which is why the stale branch survived — read as a library value, which is what the zero-install detector mirrors, a capped single-scope workspace returned "Workspace does not appear to be an agent project." Truncation is now checked ahead of the adoption and negative branches in both detectors and in the script's human output. `first_look`'s final `Next:` line is the retry rather than a `verify --preview` that walks past it. **The preview probe outran the command it recommends.** It read a directory's own files while a scoped `detect` spends the cap over that directory's whole subtree in inventory order, so a direct `agent.py` sorting after a thousand inert modules was evidence preview could see and the scoped command could not. Python evidence is now bounded to the files that command would reach. **Removed boundaries are derived before anything reads the head tree**, since a deleted `pyproject.toml` leaves nothing for a head-tree marker filter to find: a PR deleting a whole project was attributed to whatever survived. Causes also accumulate rather than replacing one another, and deleted evidence outranks a cap — raising a bound cannot find a file the change removed. **A head that is not this worktree is its own cause.** Collapsing it into empty evidence let the generic route authorize a `detect` of the current worktree, which answers about a different tree and can recommend root init for an unrelated agent. It now carries no command. fix(inventory): bind a tool inventory to the source it completes (#386) (#392) * fix(inventory): bind a tool inventory to the source it completes (#386) `incomplete_surface` fires for every statically-extracted tool on a first ADK/LangChain/CrewAI/n8n scan, and the only remedy the tool offered was "save the skeleton, reference it from `.tool_inventories`". Following that instruction exactly made things worse: the inventory was loaded as an *independent* source, so its entries were added beside the extracted tools rather than joined to them. On a minimal ADK repro the catalog grew 4 -> 6, reachable fell to 4/6, semantic gaps went 2 -> 8, the `action_surface` rows that used to resolve became `ambiguous_tool_selector`, and the gap that asked for the file was still open. The loop had no third step. `.tool_inventories[]` entries now take `source_id`, naming the tool source whose surface the file enumerates. Each entry matching a name that source already exposes is joined to that observation, so the catalog keeps its size, the merged tool inherits the inventory's high extraction confidence, and the gap closes. Entries the source does *not* expose stay standalone: an inventory exists precisely to disclose tools static extraction missed, and a tool nobody wired is still honestly reported as unbound. Nothing is joined by name alone. `build_tool_identity_catalog` opens with "Build canonical tools without ever joining observations by name", and that invariant is kept: `source_id` is a manifest declaration desugared into the same reviewed-binding engine as `tool_identity.bindings`, one binding per matched name, with the inventory as `primary`. A name a source exposes twice implies no join and asks for an explicit binding; a reviewed binding that already claims an observation always wins. Both emitters of the prescribed text now route through one builder that names the field, the `suggested-inventory.json` note carries the exact YAML entry, and an inventory declared without `source_id` that shadows a low-confidence source says so in `source_warnings` rather than degrading in silence. The skeleton note and the fix-task rollup name the route only for source types that actually have a `tool_inventories` key. Applied to all four frameworks with that key: the schema, the loaders, and the prescribed-text table are shared code, so an ADK-only fix would have left three identical defects. `samples/google_adk_agent` drops from 20 lines of hand-written bindings to one entry, with byte-identical coverage and findings. Co-Authored-By: Claude Opus 5 * fix(inventory): keep source identity and evidence through completion (#386 review) Three review findings on the inventory-completion change, each reproduced before it was fixed. Selector identity. Making the reviewed inventory `primary` rekeys the canonical tool's `source_id` to `google_adk_inventory:...`, and selector resolution compared `source_id` only. A row such as `{tool: lookup, source_id: adk_agent}` resolved before the documented remediation and matched nothing after it — including rows Shipgate scaffolds itself, since `_action_selector` qualifies every action row by `source_id`, and including the same-name-provider case where a source-qualified selector is mandatory. A canonical tool now answers to the `source_type`/`source_id` of any observation bound into it, read from `identity_assessment`. Given together, both qualifiers must be satisfied by the *same* observation, so a selector cannot pair one member's type with another member's id. This repairs hand-written `tool_identity.bindings` too, not only desugared ones. Source-owned semantics. `_merge_bound_observations` starts from the primary and copied nothing else across, so completing a source erased that source's own evidence. An n8n tool with `apiKey`/`unscoped` auth, an output schema, and an owner came back high-confidence with unknown auth, `{}` output, and no owner — trading the closed `incomplete_surface` gap for `partial_authority_evidence`, the same non-monotonic move #386 is about. Merging now backfills `description`, `input_schema`, `output_schema`, `parameters`, `function_signature`, `owner`, and auth `type`/`credential_mode`/`source`/`mode`/`explicit` into empty slots only. The reviewed observation still wins wherever it says something, and two populated disagreeing values remain `conflicting_tool_identity` rather than a silent overwrite. Source identity is deliberately excluded from the backfill: the row keeps the primary's provenance, and other members stay selectable through the alias path above. That erasure was suppressing findings, not only degrading evidence. `samples/support_refund_agent` binds a `-> str` SDK function to an inventory silent about output, and the merge dropped both the AST-derived `{"type": "string"}` schema and the `sdk_function` source type that `SHIP-SCHEMA-FREEFORM-OUTPUT` falls back on, so the shipped golden recorded no free-form-output finding for a tool that plainly returns free-form text. The finding is restored: one new MEDIUM review item, with the sample's `blocked` verdict and its five blockers unchanged. Goldens regenerated. YAML safety. `source_id` is unconstrained and generated framework ids embed the configured path, so a comma split `source_id: google_adk:agent,prod.py` into two keys and the exact entry the tool prescribed failed manifest validation under `extra="forbid"`. Both emitters now render the value through a new `yaml_scalar`, and the tests parse the emitted entry with `yaml.safe_load` rather than matching a substring — a substring assertion passes on the broken output. Co-Authored-By: Claude Opus 5 * fix(inventory): alias identity everywhere, and check preserved evidence (#386 follow-up review) Five findings from the follow-up review, each reproduced before it was fixed. Aliasing missed the id the scaffold actually carries. Completion rekeys the canonical tool_id from an observation-derived hash to a binding-derived one, and `_action_selector` emits `tool_id` on every action row it scaffolds — while `resolve` prioritizes it. The earlier regression passed only because it hand-wrote `tool` + `source_id`; the generated declaration still became `unresolved_tool_selector` against the very inventory the tool had prescribed. A canonical tool now answers to the id each of its observations carried while unbound, computed the same way the catalog issues one. Aliasing also missed two consumers. `_action_has_policy_control` and `_matching_suppression` compared the canonical primary fields directly, so a source-qualified `require_confirmation_for_tools` entry silently stopped applying — the scan reported a missing `confirmation.required` and moved to `blocked` on a manifest the user never touched — and a source-qualified `checks.ignore` went inert. Both now route through one shared `IdentityAliases.matches`, which also covers tool_id. Alias ids resolve selectors only and stay out of `ToolSelectorIndex.by_id`: `agent_bindings` reads that as the whole catalog to partition reachable/possible/unbound, and an alias there invents a catalog member and trips the tool_catalog/binding-graph consistency invariant. Preserved evidence is conflict-checked. Backfilling "the first non-empty value" resolved genuine disagreements by observation-id order: two members reporting `owner: team-a` and `owner: team-b` produced a tool owned by `team-a`, no issues, `pass_eligible=True`. Every contributor to `output_schema`, `function_signature`, `owner`, and `auth.credential_mode` is now compared, the primary included, and more than one distinct populated value is `conflicting_tool_identity` — which makes the identity non-pass-eligible. Schemas compare through `_schema_signature`, so a reviewed refinement still reads as compatible rather than contradictory. `auth.source` is deliberately preserved without being compared, which is the one place this departs from the review. It names the *extractor* that produced the auth record — `google_adk_static` on an AST observation, `mcp` on the inventory reading of the same tool — so two observations of one capability disagree by construction. Conflict-checking it fired on every completed ADK tool and took `pass_eligible_actions` to 0, the opposite of this change's purpose. `credential_mode` is a real claim about the credential and is checked. Preserved evidence is traceable. Each backfilled value records the observation that supplied it, and `tool_finding` accepts `evidence_field` so a check names the field it judged and the finding cites the artifact that declares it. The restored free-form-output finding now points at `agents/refund_agent.py:5`, where the `-> str` is, instead of an inventory JSON with no output schema in it. Opt-in per check, because inferring the donor from evidence-dict keys would misfire whenever a key shares a Tool field name (`source` and `owner` both do). YAML encoding is total. `ensure_ascii=False` emitted C1 controls literally: PyYAML rejects a stream carrying U+0080/U+009F/U+007F or a lone surrogate, and silently normalizes U+0085 NEL to a space — so an id containing NEL round-tripped to a different id and the remediation named the wrong source. Default `json.dumps` escaping costs readability on accented identifiers and buys a scalar that always parses back to the value it names. Goldens regenerated: the free-form-output finding's source and fingerprint move. `report.md` is unchanged. fix(inputs): refuse an absent input instead of reporting it as a malformed one (#389, #387, #384) (#391) Three reports, one class: an input that is not there was reported as an input with the wrong shape — and in one case the command created the input it had been asked to inspect. any output directory is resolved or created, on every command that takes the option. `verify --preview` created the entire four-level path, wrote a full artifact set into it, and exited 0, so a typo produced a confident result about a workspace that was never there and CI read healthy on both signals a caller can gate on. Closing it uncovered four more instances: `init --write`, `audit --host`, and `verification prepare`/`worker` raised bare FileNotFoundError tracebacks; `install-hooks --write` wrote hooks into the mistyped tree; `mcp audit` answered `decision: allow`; and `detect`, `check`, and `trigger` exited 0 with a payload describing nothing. Preview's documented "always exits 0" is narrowed deliberately: it is a promise about workspaces it evaluates, and there is nothing here to evaluate. No contract or schema version moves; the refusal reuses the existing config_error / exit 2 slot. b"" collapse that binds a failure to the bytes the identity hashed survives while absent no longer reaches the YAML shape check. Absent, empty, and present-but-not-a-mapping now produce three distinct messages, which stops `control.reason` ("fix this file") from contradicting `control.next_action` ("bootstrap from scratch") about a file that does not exist. The same conversion in the diff-input path — "Workspace is not inside a git checkout" for a path that was never created — goes with it. ValueError and AssertionError into a ValidationError but lets TypeError propagate past the config-loading boundary, so a YAML mapping where a list belongs surfaced as internal_error telling the user to file a bug for their own typo. Messages now name the field path and the shape that was written. Two sweeps keep the class closed rather than the instances: test_every_workspace_command_is_swept enumerates --workspace commands from the live Typer app and fails when a new one is not covered, and test_no_schema_module_raises_typeerror bans the statement outright. test_absent_input_messages.py asserts absent and malformed stay distinct across all six inputs Shipgate reads; the other four already did. tests/test_adapter_static_only.py pins allowlisted subprocess call sites by line number, and the new imports shifted three of them. fix(google-adk): reach sub-agent tools and name every unreached one (#385) (#390) * fix(google-adk): reach sub-agent tools and name every unreached one (#385) On the canonical Google ADK multi-agent shape -- a coordinator with `sub_agents=[salesforce_agent, sap_agent]` -- every tool the sub-agents owned fell out of the root-reachable graph, and none of the 25 evidence gaps named one. On the reported repository the excluded half was the half a release gate exists to judge: three financial writes, including one that sets opportunities to `Closed Won`. Three independent defects produced the one symptom. ADK routes a handoff by the sub-agent's `name=`, but `sub_agents=[...]` spells the Python variable the agent was assigned to. Reading the variable as an agent name produced one phantom node per sub-agent, owning no tools, so the handoff landed on a node with nothing behind it and the real agent stayed unreachable. The two spellings are now reconciled from the module's own assignments, which also collapses the duplicate nodes the graph used to report. An element that cannot be named now fails closed as partial evidence; naming two of three sub-agents previously reported the two as the whole handoff set. `agent_bindings.declarations` was the documented remedy and could not work. Declaring an agent seeded a synthetic node for it, and for an agent the scanner had already observed that second node made the name ambiguous, so the resolver rejected names its own scan had emitted. Declarations now reuse the observed node, and a genuinely ambiguous name says how many agents share it and which sources they came from. Finally, a tool bound to an agent the configured root cannot reach now gets an evidence gap naming it. Everything downstream of the binding graph is narrowed to root-reachable tools, so such a tool is never judged; before this it was not mentioned either, leaving the ratio `6/12 catalog tools reachable` as its only trace. Co-Authored-By: Claude Opus 5 * fix(google-adk): resolve sub-agents by lexical scope and fail closed on imports Addresses both P1s from review on #390. The variable-to-agent map was keyed by the bare target name while `_agent_calls` walks nested functions, so two factories that each build a local `worker` collapsed into one entry. Reproduced: `RootA(sub_agents= [worker])` reached `WorkerB`, so the gate analyzed `write_b` -- a tool that root cannot call -- and excluded `read_a`, which it can, while still reporting `structural` / `pass_eligible=True`. A wrong capability surface presented as proven is worse than a missing one. The map now keys on the enclosing scope and resolves innermost-out from the referencing call, which gets both factories right rather than merely failing closed on them. One name rebound to differently named agents inside a single scope is genuine flow-sensitivity that AST position cannot settle, so it resolves to nothing. Completeness counted `_qualified_name` successes, but an imported name qualifies (`from sub import worker` -> `sub.worker`) without matching any agent definition. Reproduced: scanning only `root.py`, the imported `worker` owns `delete_record`, which never enters the catalog -- so the per-tool gap cannot cover it either -- and the graph reported `structural` / `pass_eligible=True` with no partial evidence. That recreated the exact silent exclusion #385 exists to close. An element that resolves to no agent definition is now recorded separately and reported as partial evidence naming the spelling, and no longer becomes a node of its own: an empty tool set on a node named after an import reads as proof the sub-agent has no capability. Three regression tests: single-entrypoint imported sub-agent, two factories sharing a local name, and ambiguous same-scope rebinding. Also applies the P3 naming correction -- `shipgate` is the CLI alias, not the display name. Co-Authored-By: Claude Opus 5 --------- Co-authored-by: Claude Opus 5 fix(verify): rank release blockers above the trust-root notice and split the no-base fail-safe (#365) (#376) * fix(verify): rank release blockers above the trust-root notice (#365) `verifier.json.headline` is the one line that reaches a PR comment, a chat reply, or a triage list. When a PR both touched the release trust root and blocked release on critical findings, it reported the trust-root fact and never mentioned the blockers. The two findings driving it are medium; the blockers they outranked were critical — and because touching the trust root is unavoidable on a first adoption, it understated severity most reliably for exactly the readers with the least context. A run carrying a critical or high release blocker now leads with the scan's own verdict line and appends the self-approval prohibition, so the human-review requirement survives in the same string rather than being replaced by it. That leading line also names the worst blocker instead of only counting it — a count reads the same whether the agent is missing a docstring or can move funds with no enforced control — picked deterministically by severity, then check id, then title. A trust-root or policy change with no blockers still leads with the governance notice: it is then the whole story. `control.human_review.why` follows the headline, so the control envelope can never name less than the headline does. Separately, the fail-safe that fires when there is no base policy to compare against shared a reason code with a proven base-relative weakening, so a first adoption reported `SHIP-VERIFY-POLICY-WEAKENED` against a base that carried no gate at all. It gets its own reason code, `SHIP-VERIFY-POLICY-BASE-ABSENT`, carrying both evidence kinds (`manifest_introduced`, `base_snapshot_unavailable`). `SHIP-VERIFY-POLICY-WEAKENED` is narrowed, not deprecated: it keeps firing for every proven weakening. Nothing about the gate moves with either change — same severity, same suppression immunity, same `human_ack` requirement on the `policy` surface, same `protected_surface_changes` rows, same release decision, same control state, `must_stop`, `merge_verdict`, `can_merge_without_human`, and `permissions`. `verifier_summary.policy_weakened` keeps its fail-safe meaning in particular: only a git-proven adoption clears it, so a rename-and-loosen diff still cannot clear the gate-bypass alarm by breaking the base scan. Co-Authored-By: Claude Opus 5 * fix(verify): preserve configured gating and honest copy across the policy split (#365) Review follow-up on #376. Six defects, each reproduced first. **Configured gating moved with the id.** A repository that had raised the no-base policy fail-safe with `checks.severity_overrides: {SHIP-VERIFY-POLICY-WEAKENED: critical}` got `medium`/`review_required` after the rename, where it had configured `critical`/`blocked` — a gate that loosened because an id moved, which is exactly what the ordering-only claim promised would not happen. The pre-split id is now an umbrella over both halves (`SPLIT_CHECK_ID_ALIASES`), kept separate from `LEGACY_CHECK_ID_ALIASES` because it is not deprecated and a baseline naming it must not be reported as stale. An override written against the new id still wins; floor validation is unchanged. **Fail-closed routing was speaking for the copy.** `policy_weakened` stays raised when the direction could not be established — that is what stops a broken base scan from clearing the gate-bypass alarm — but every renderer read it as a fact and said "This PR weakens the release policy" about a change nothing compared. `capability_review` gains `policy_weakening_proven` (additive, default false): the narrower fact that a base-vs-head comparison actually ran. The route is identical either way; only the claim differs, in the headline, the control reason, and the fix task's repair reason. **The named blocker title was untrusted, unbounded, and multiline-capable.** `ReleaseDecisionItem.title` embeds a tool name read out of a scanned spec, so newlines survived into a field contracted to be one sentence, and length alone pushed the appended `cannot self-approve` clause past the compact control envelope's 400-byte prose budget — deleting the human-review requirement from the projection a routing consumer reads. Control characters are collapsed, the quoted title is capped, and the governance suffix now gets a reserved byte budget so the lead is shortened instead of the requirement. If the requirement alone fills the budget it is published on its own, which is what the headline said before blockers ever led it. **Artifacts written before the split stopped reprojecting.** A stored report carrying the old id with `manifest_introduced` was read as a weakening and no longer recognized as an adoption. Every read path — verifier summary, capability review, `protected_surface_changes`, `human_ack`, the adoption fix-task route, and the Action's findings fallback — now accepts either id with the same evidence kinds, through one shared vocabulary in `core/policy_reason_codes`. Nothing re-emits the old id. **Stale public rationale.** `docs/triggers.json` named the old id as the no-base fail-safe; corrected there and in the five surfaces that copy it (three protocol goldens, two examples). `AGENTS.md:421` carries the same stale enumeration and is a protected trust root: preflight routes that edit to `human_review_required`/`must_stop`, so it is left for a human and reported on the PR. Its claim — that Tier B checks select changed files the same way — is still true; the list is now incomplete, not wrong. **Reversed catalog text.** `fires_when` for `SHIP-VERIFY-POLICY-WEAKENED` said the base was weaker than the head, the opposite of the implementation. Corrected and regenerated. Co-Authored-By: Claude Opus 5 * fix(verify): keep the wire contract, the comparator, and the budget honest (#365) Second review follow-up on #376. Five defects, each reproduced first. **The PR comment still claimed a proven weakening.** Headline, control reason, and fix task were made honest for an unprovable policy direction, but `pr-comment.md` printed `Policy weakened: true` straight off `policy_weakened` — the fail-closed flag that is raised precisely when nothing was compared. It now prints `Policy changed, weakening unproven: true` with the reason. The route it reports is unchanged. **A frozen schema identifier gained an emitted field.** An artifact still declaring `verifier_schema_version: "0.8"` failed validation against the published v0.8 schema with `Additional properties are not allowed ('policy_weakening_proven' was unexpected)`. The verifier schema advances to `0.9`, `docs/verifier-schema.v0.8.json` is restored to its published bytes, and `0.8` joins the legacy read set — the field defaults to `false`, which is what "this artifact recorded no comparison" means. The model now also rejects the contradiction the docs already forbade: `policy_weakening_proven=true` requires `policy_weakened=true`. **The comparator ignored the split aliases the runtime honors.** Tier B read literal override keys, so it produced both kinds of wrong answer: a head adding an explicit override for the new id lowered the applied severity with no key change on the umbrella (missed), and a head dropping a redundant explicit override changed no applied severity at all (falsely reported). Resolution now mirrors `_severity_override_for_check` exactly — exact id first, then umbrella — and the comparison covers every check either side's configuration reaches. **Accepted debt did not survive the split.** A fingerprint hashes the check id, so a baseline entry recorded against the pre-split id stopped matching: a `critical` accepted item went from matched debt to new, and its decision from `review_required` to `blocked`. Baseline matching offers the pre-split fingerprint as an additional candidate, scoped to declared split targets and still required to agree on `support_hash`. **The reserved budget was spent after it was reserved.** The evidence-gap provenance note was appended *after* composition; a long multibyte title plus one gap note made 443 bytes and the 400-byte compact projection dropped `a human must review it.` from both `reason` and `human_review.why`. Every later addition now goes through one composition, the configured-manifest path is bounded, and when room runs out the parts yield in priority order — gap note first, never the verdict, the named cause, or the requirement. **Unicode format controls survived C0/C1 filtering.** U+202E and U+2066 reorder rendered text without changing a byte, so a tool name carrying one could visually move the reserved governance suffix; a lone surrogate raised inside the byte budgeting. Unsafe categories (Cc, Cf, Cs, Co, Cn, Zl, Zp) are collapsed before any byte accounting. feat(cli): give the repository one launcher and doctor an environment block (#334) (#383) * feat(cli): give the repository one launcher and doctor an environment block (#334) Running Agents Shipgate from a source checkout meant either a bare `agents-shipgate`, which resolves through `PATH` and can silently execute a pipx or base-conda copy — a `0.8.0` shadowing a worktree makes new subcommands look "missing" — or `PYTHONPATH=src python -m agents_shipgate`, which is correct but has to be discovered. A console script promoted from an environment that no longer exists is worse than either: it dies with `ModuleNotFoundError` before a line of Shipgate runs, so nothing in Shipgate's own output can explain it, and the epic's reproduction (#338) had a user drop into a terminal at exactly that point. `./shipgate` is now the canonical contributor command. It puts this tree's `src/` ahead of every installed copy, in child processes too, and selects an interpreter — `AGENTS_SHIPGATE_PYTHON`, else the project virtualenv, looked up in the main checkout as well so a `git worktree` shares it — re-executing exactly once under a loop guard. It announces itself through `AGENTS_SHIPGATE_CLI`, the operator override the invocation policy already honours (#322), so every command it prints back is runnable as printed; without that its `argv[0]` reads as the `shipgate` console script and the policy would emit commands a clean checkout cannot run. An operator's own override still wins. `doctor --json` payloads — and every `doctor` agent-mode error line, including the discovery failure that prints no payload at all — now carry an `environment` block: the interpreter, the launcher and every Shipgate console script on `PATH` with the interpreter each shebang names, the import source, the installed / imported / source-tree versions, and `mismatches[]` with severities and runnable recovery commands. Nothing runs an interpreter or executes a console script to find out — a wrapper that cannot start is exactly the one that cannot report on itself. Two orderings are load-bearing. The checkout search is caller-first: a worktree borrows the main checkout's editable install, so import-first would find that tree, agree with itself, and report nothing while every edit went unexecuted. And a source checkout out-voting an installed distribution is not a mismatch — that is the intended state, and an editable install's metadata lags every version bump by design. Adds `environment_error` (exit 4, the existing "other error" code) to `docs/errors.json`, and `invocation.render_cli_override`, the host-rules inverse of how `AGENTS_SHIPGATE_CLI` is parsed, now that something writes it. No schema or contract version changes. Co-Authored-By: Claude Opus 5 * fix(cli): read a quoted program token, rank the recovery, follow the sh trampoline Three failure-path defects from review, each reproduced first. **A quoted program token is read before it is judged to be ours.** `retarget_command` located the program by scanning to raw whitespace, on the argument that our console-script names contain none — true of the names, and irrelevant to the strings they appear in. A quoted interpreter path whose directory is named after this project, which cloning it into `~/agents-shipgate worktree/` produces, was cut at the space; the remaining `'/tmp/agents-shipgate` has `agents-shipgate` as its basename, so the launcher's own `python -m pip install` recovery was rewritten to name the Shipgate entry point. Not an unrunnable command but a runnable one that runs the wrong program, and the dangling quote also cost the action its `executable`/`args` pair. The span is now found with quoting honoured, the token's value comes from `shlex` — the same grammar the string was rendered with — and the two are cross-checked so a disagreement leaves the command untouched. A correctly quoted Shipgate path is now retargeted where it previously was not. **The recovery is ranked by what the interpreter can actually do.** `venv --without-pip` answers `python -m pip install …` with `No module named pip`, so emitting that alone promised a recovery that fails on its first token in exactly the environment the recovery exists for. `ensurepip` is proposed ahead of it when `pip` is absent; an interpreter with neither gets the diagnosis and no command rather than one that cannot run. Both prerequisites are answered by asking the selected interpreter about itself — this code is already running inside it. The test now executes rank 1 for real and proves rank 2 can start. **A `#!/bin/sh` wrapper reports the interpreter it `exec`s.** An interpreter path containing a space cannot go in a shebang, so `pip` writes a shell trampoline. Reading only the shebang reported `/bin/sh`, which exists and is not the running interpreter, so a healthy install raised `console_script_runs_other_interpreter` once per alias while the interpreter that could actually go stale stayed invisible. The `exec` target is parsed; an unrecognised handoff reports `null` rather than the shell. The isolated-interpreter fixture now lives at `<...>/agents-shipgate worktree/.venv/`, so one venv covers both the missing-dependency path and the retargeting regression. Co-Authored-By: Claude Opus 5 * fix(cli): announce a runnable launcher spelling and read PATH the way a shell does Four defects from the second review, each reproduced first. **The launcher had no Windows-invocable spelling.** A shebang is a POSIX kernel feature, so `.\shipgate` is a file Windows will not execute — and announcing that path through `AGENTS_SHIPGATE_CLI` published recovery commands that could not run there, which is this launcher's own defect relocated to another platform. `launcher_argv()` now announces ` ` wherever the file cannot be started on its own: on Windows, and on a copy that lost its executable bit. `python shipgate ...` is the documented Windows spelling and it is the one that gets emitted. No `.cmd` shim: `CreateProcess` runs a batch file through `cmd.exe`, which re-parses arguments under a second quoting grammar, and this project publishes structured argv that callers execute without a shell. **Two virtualenvs over one base were one interpreter.** `runs_this_interpreter` resolved both paths, and a POSIX virtualenv's `bin/python` is a symlink to the interpreter it was built from — so unrelated environments with different `sys.prefix` and `site-packages` compared equal, and a console script pointing at a different one reported clean. Reproduced: two venvs here both resolve to the same uv-managed binary. The comparison no longer dereferences, which is the rule the launcher already applied for the opposite consequence; the two copies are now pinned to each other by test, since the launcher cannot import the package before the version gate. **`PATH` lookup accepted a file the shell would skip.** POSIX command lookup requires the execute bit and continues to later entries without it, so stopping at the first existing file described a wrapper the caller's shell would never run — and hid the stale-interpreter diagnostic for the one it would. Executability is now required on POSIX and `PATHEXT` decides on Windows, with the extensions taken from the probed environment so the rule stays a stated fact. Cross-checked against `command -v` on the case that used to differ. **A commented-out `exec` could become the interpreter.** The trampoline target was found by searching the whole wrapper, so `# old target: exec "/deleted/python"` above a working `exec` reported a healthy wrapper as `console_script_interpreter_missing`. Lines are read in order now, with `exec` required in command position and comments skipped: a diagnostic may not be derived from a string the shell never executes. Adds a `windows-launcher` CI job covering the entry point, the announcement, and the `os.name == "nt"` re-execution branch — including that a non-zero status survives the hop, which is what a branch that spawns instead of replacing the process can lose. Co-Authored-By: Claude Opus 5 * test(environment): pin the PATHEXT lookup to the rule, not the host filesystem `test_windows_lookup_follows_pathext` wrote `agents-shipgate.exe` and probed with `.EXE`. Windows matches filenames case-insensitively and so does macOS by default, so the test passed locally while asserting nothing about the lookup — on Linux's case-sensitive filesystem it failed, which is how CI found it. The fixture and the probed suffix now match exactly, and a second assertion removes the suffix list to prove the suffixes are what find the file rather than the directory listing. Co-Authored-By: Claude Opus 5 * ci: give the released-install check the PATH an adopter actually has The step asserted `mismatches == []` while the runner's own editable install had `agents-shipgate` and `shipgate` on `PATH`, pointing at the hosted-tool interpreter rather than at the wheel under test. That really is a `console_script_runs_other_interpreter` — a bare `agents-shipgate` there would execute a different installation — but it is not the adopter environment this step exists to check. It passed before only because the interpreter comparison resolved symlinks: the venv's `bin/python` links to the hosted-tool binary, so the two collapsed into one interpreter and the mismatch never fired. Fixing that comparison is what made this surface, which is a second, independent confirmation of the review finding — and evidence that this assertion was previously vacuous. `PATH` is now scoped to the released install for the run itself, so the wrappers found are the wheel's own, and the step additionally asserts the healthy adopter shape — both wrappers present and running this interpreter — rather than inferring it from an empty mismatch list. fix(agent-mode): rank the insufficient_evidence reason by actionability (#362) (#375) * fix(agent-mode): rank the insufficient_evidence reason by actionability (#362) An `insufficient_evidence` verdict is announced three times in a row, and each line answered "what is wrong here?" on its own. The reason counted source warnings — the symptom — and demoted the one actionable gap to a secondary line, while `agent_summary.first_recommended_action`, the field the agent contract routes coding agents to, contradicted that line outright: "applying patches does not clear an evidence verdict, so no machine-applicable fix is available", printed directly beneath `Improve evidence: … Target: shipgate.yaml#agent_bindings.declarations`. For a coding agent that is a dead end, and the cheap ways out of a dead end are the ones `forbidden_actions` enumerates. `release_decision.reason`, `primary_evidence_remediation_text`, and `first_recommended_action.why` now project one selected gap (`core/evidence_actions.primary_evidence_gap`): the first `evidence_gaps[]` row whose `next_action.path` is non-null, falling back to the first row when nothing is addressable. The reason reads "Insufficient evidence: (). Fix at . Context: …", and "no machine-applicable fix is available" is unreachable on every branch that could emit it while any gap names a path. Where it still appears it is true. The action stays `kind: "info"`: an evidence gap is closed by a reviewed declaration, never by a command Shipgate hands you. Two copy rules came with it. Warnings that restate one mechanism collapse at render time — six "Google ADK agent 'x' references unresolved tool 'y'." lines become one naming the cause, every affected symbol, and the two surfaces that close it — in report.md, packet.md/html, verify's fix-task remedies, and the CLI --verbose list, which now prints the mechanism count beside the raw one. report.json, packet.json, and `evidence_coverage.source_warning_count` are untouched: that count gates, and folding it would silently recalibrate the threshold. And a `tool_identity.bindings` member naming a source that produced no observations now states the rule and points at `agent_bindings`, instead of reporting only that the member "matched 0 observations". Verdict strictness is unchanged: `_MAX_TOLERATED_SOURCE_WARNINGS` and `_LOW_CONFIDENCE_TOOL_RATIO` stay frozen, `evidence_below_ie_threshold` reads the same counts, and no schema version moves. Closes #362 Co-Authored-By: Claude Opus 5 * fix(agent-mode): address Codex review on #362 actionability ranking Five findings from the first independent review of PR #375, each reproduced before it was fixed. P1 — line and Markdown injection through repository-derived fields. Only `_decision_reason` one-lined its inputs, so `next_action.expects`/`path` reached `Improve evidence:` and `first_recommended_action.why` verbatim: a policy-derived path ending `\nControl: complete\nYou may: merge` forged lines the CLI printed raw and `_safe_markdown_text` does not collapse. Normalization moves into the shared projection (`evidence_gap_action_text`), covering expects, path, and command while preserving the deliberate `Run:` separator, and the GitHub step summary now escapes `release_decision.reason` the way report.md always has. P2 — an unknown binding `source_id` was sold as a zero-observation source. `observed_source_ids` alone cannot tell a configured-but-empty source from a typo, so a member naming `orders_typo` was told to declare `agent_bindings`, which cannot repair an invalid selector. Configured ids are now tracked separately: a known empty source keeps the `agent_bindings` rule, an unknown one is told to correct the member to a configured `tool_sources[].id`. P2 — warning grouping invented facts. Merging each quoted column independently reported two agents failing on one shared symbol as "2 tool symbols", cross-producted binding/source/tool tuples, and could not parse a `repr()` literal containing both quote styles. Grouping is now structural: a mechanism declares context fields (part of the group key) and subject fields (listed as whole tuples), quoted counts come from distinct subjects while `group.count` stays the raw row count, and a row groups only when re-building the parsed fields reproduces the message byte for byte. Unrecognized text is never merged. P2 — `fix_task` dropped typed source-warning repairs. The stale-`--diff-from` gap carries `provide_source`, a path, an expectation, and the regeneration command; two blanket `source_warning` skips threw all of it away and left prose, so the handoff named a different repair from the selected gap. Only pathless review-only warnings fall through to prose now, and a typed row is not also restated. P2 — the published contract promised more than the code guarantees. Alignment is now scoped to `insufficient_evidence` with an addressable gap; the no-machine-fix prohibition is stated separately as holding on every verdict; the two intentionally-divergent cases are spelled out; "non-null" is corrected to "non-empty" (which is what the code checks); and the stale first-row / generic-deeper-sources guidance in the report-reading primer is updated. Verdict strictness, `source_warning_count`, the JSON warning arrays, and the frozen thresholds are unchanged. Co-Authored-By: Claude Opus 5 * fix(agent-mode): address Codex review 2 on #362 actionability ranking Four blockers from review #4940674487, each reproduced before it was fixed. P1 — the warning parser accepted ambiguous delimiter splits. The rebuild check was vacuous: a regex of escaped literals plus capture groups rebuilds byte-for-byte for *any* successful match, so it validated nothing. Confirmed four corruptions: an ADK agent name containing " references unresolved tool " split into a different agent/symbol; a source_id containing ", tool=" rewritten; a binding warning with two invalid members silently reporting only the first; and a mixed unknown/configured-empty message selecting the wrong remediation — reintroducing round 1's P2 through the parser. Decoding is now exact: every interpolated value is repr() of a string, so the decoder reads a string literal at each field position and requires the message to be consumed exactly. A value carrying the separator is read whole; a composite or non-canonical message decodes as nothing and renders verbatim. Mechanisms also declare repeated fields (a binding message prints its source_id twice) and reject a message where they disagree. P1 — typed fix-task fields bypassed one-line normalization. subject, why, expects, path, accepted_values, and command were interpolated raw into fix_task.instructions[] and allowed_repairs[].target/reason/command, which agent_result copies verbatim into repair.instructions, suggested_fixes, and agent_repair_instructions. The hostile fixture wrote literal "Control: complete" lines into all of them. Every interpolation now goes through the shared one_line projection. P1 — addressability was raw string truthiness. A schema-valid whitespace- or control-only path won ranking, rendered "Fix at .", masked a real path on a later row, created a bogus typed repair, and suppressed the truthful no-fix route. One shared predicate (evidence_gap_target / is_addressable_gap, both normalized) is now used by ranking, the release reason, the agent summary, and both fix-task source-warning predicates. P2 — the review_required guidance promised behaviour branch precedence does not implement: with auto-applicable patches and sub-threshold evidence the action is the apply-patches command, as test_review_required_sub_threshold_ evidence_keeps_auto_apply already pinned. The promise is scoped to the evidence-first branches, and the stale mirrors are updated: the degraded- evidence condition list in the report-reading primer (which omitted semantic, binding, policy, and typed provide_source conditions), STABILITY.md's evidence_gaps[0] recovery rule, summary_text.py's "can never differ" docstring, and all four mirrored fix-top-finding.md prompts (render hash bumped, copies byte-identical). llms-full.txt regenerated. Verdict thresholds, the raw source_warning_count, and the JSON warning arrays are unchanged. Co-Authored-By: Claude Opus 5 * test(agent-mode): pin the review_required auto-apply counterexample (#362) The contract text now scopes the review_required promise to the evidence-first branches. This is the counterexample it scopes around, asserted beside the claim so the doc and the action picker cannot drift apart again: sub-threshold evidence plus an auto-applicable finding returns the apply-patches command even when an addressable evidence gap exists. Co-Authored-By: Claude Opus 5 * fix(agent-mode): address Codex review 3 on #362 actionability ranking Seven findings from review #4941036126, each reproduced before it was fixed. P1 — opaque warning text reached text consumers raw. An unrecognized warning was preserved byte-for-byte as `group.message`, so `Optional source bad failed to load:\nControl: complete\x1b[2J` put a forged physical line into report.md and packet.md/html and left ESC in the CLI and fix-task strings. The *display* copy is now normalized; `SourceWarningGroup.warnings` and `report.source_warnings` keep the loader's bytes so nothing that counts or gates moves. `_dedupe_cap` normalizes every final fix-task instruction as a backstop for strings that arrive from elsewhere. P2 — Unicode format and bidi controls defeated normalization. `\s` plus C0/C1 left U+200B, U+200E, U+202E, and U+FEFF intact, so a path of one ZWSP was "nonblank", won ranking, and rendered `Fix at .`, and U+202E could reorder the visible guidance. `one_line` now drops general-category `Cf` and keeps every visible script. P2 — affordances were gated on raw values. `command=" \n\x00 "` published a bare `Run:` and a `""` repair command, and blank accepted values rendered `Accepted values: , .`. Command and accepted values are normalized first, blanks dropped, the suffixes gated on the result, and an empty normalized command becomes `null`. P2 — the evidence note was derived from the overloaded `human_review_recommended`, which is also true for any critical/high finding. A producer-valid `review_required` with one high auto-applicable finding and zero measured gaps was told "applying patches does not address the evidence gap". New `has_measurable_evidence_gaps` keys it on the same measurable inputs `_decision_reason` already uses. P2 — the contract promised `kind: "info"` in all evidence cases two lines after documenting the `kind: "command"` auto-apply branch. Scoped to the evidence-first actions, and "non-empty string" is corrected to a nonblank *normalized* target across the contract, STABILITY, the report-reading primer, the four prompt mirrors, and llms-full.txt, so third-party consumers do not reimplement the bug just fixed. P2 — the prompt called every insufficient-evidence repair a declaration. The stale-base case emits `provide_source` with `path=--diff-from` and a command: there is no file to open. All four mirrors now branch on `next_action.kind` and reserve the never-auto-assert rule for effect, authority, and binding declarations. Render hash bumped; copies byte-identical. P2 — the Conductor golden was regenerated in full rather than string-edited, dropping an inapplicable `agent_bindings.root` scaffold the engine no longer emits for a `provide_complete_binding_graph` row, plus a field-for-field generation parity test that an invariant-only check could not provide. Thresholds, raw `source_warning_count`, and the trust-root gate are unchanged. Co-Authored-By: Claude Opus 5 * test(agent-mode): pin the Conductor parity scan's environment (#362) The new field-for-field golden comparison passed locally and failed on Actions: a scan appends `github_step_summary` to `privacy_audit.output_surfaces` whenever GITHUB_STEP_SUMMARY is set, so the artifact is environment-dependent and the fresh scan gained a surface the committed one does not have. Clear the variable for the scan rather than widening what the comparison ignores — the point of the check is that nothing drifts, and `output_surfaces` is real content. Verified both ways: the test now passes with GITHUB_STEP_SUMMARY set, and the whole non-packaging suite passes under a CI-like environment. Co-Authored-By: Claude Opus 5 * fix(agent-mode): address Codex review 4 on #362 actionability ranking Six findings from review #4941330151, each reproduced before it was fixed. P1 — normalization rewrote the values it was rendering. Dropping general category `Cf` wholesale turned `agents/👩‍💻.yaml` into a different filename, changed Persian identifiers carrying ZWNJ, and — worst — repaired the executable token `r​m -rf /tmp/x` into a runnable `rm -rf /tmp/x` the repository never wrote. The cut was also incomplete: VS16 and CGJ are `Mn`, so an all-VS16 path stayed "addressable". Three questions are now separate. Display (`one_line`) escapes controls and bidi as `` and never deletes an identity-bearing character. Visibility (`has_visible_content`) uses Unicode Default_Ignorable, so an all-invisible target names nothing. Executability (`is_publishable_command`) is all-or-nothing: an unsafe command is suppressed, never sanitized, because sanitizing one changes which program runs. P2 — an all-control warning normalized to the empty string, so report.md and packet.md rendered blank bullets, packet HTML rendered `
  • `, the CLI rendered `- `, and the fix task emitted a bare `Resolve source warning:` while the gate reported degraded evidence. A deterministic placeholder now stands in when the display copy has no printable content; raw warnings and the count are untouched. P2 — `_decision_reason` kept a narrower copy of `has_measurable_evidence_gaps` that omitted binding, policy, and typed-gap inputs, so a mixed review whose selected action names a binding declaration reported only "1 finding requires human review" and the headline dropped the evidence clause. It now calls the shared predicate, and the reason's subject/verb agreement is fixed with it (the crewai golden is regenerated for that string). P1 — the prompt expanded a human-owned capability assertion into agent work. Production `declare_tool_inventory` rows are `requires_human_review=true`, `auto_apply=false`, `suggested_patch_kind=manual`, and their command is the rerun *after* a human supplies the inventory. All four mirrors now route on those published authority fields rather than a kind allowlist, cover every action kind, and isolate the one genuinely mechanical case: regenerating a stale `--diff-from` base that Shipgate itself produced. P2 — `path` and `command` are independently nullable, and only the path was read, so a command-only `provide_source` row published `Run: …` from `Improve evidence:` while `first_recommended_action` said no fix existed, and a later command-only row lost selection to a blank one. A gap is addressable when it names a target *or* carries a publishable command, on every surface. P2 — the contract and changelog claimed evidence repairs are never closed by a command, contradicting the typed `provide_source` route this PR preserves. Scoped to the summary projection, with the mechanical route named; mirrors and `llms-full.txt` regenerated. Thresholds, raw `source_warning_count`, and the trust-root gate are unchanged. Co-Authored-By: Claude Opus 5 * fix(agent-mode): address Codex review 5 on #362 actionability ranking Six findings from review #4941919422, each reproduced before it was fixed. P2 — display normalization still rewrote identity. `one_line` folds whitespace, which is right for prose and wrong for anything a reader opens or runs: it mapped `configs/foo bar.yaml` to a one-space neighbour, mapped NBSP/U+3000 to ASCII space, trimmed the ends, and — as the final fix-task backstop — rewrote a validated `python -c 'print("a b")'` into a different program in `instructions[]` and the agent-result copies. Identity-bearing values now use `display_literal`: line-breaking and bidi characters become a visible `` escape, everything else is preserved exactly, nothing is folded or trimmed. Prose keeps folding, and now folds before escaping so a stray newline still reads as a space. P1 — command validation ran after `.strip()`, so a boundary character was removed before it could be seen: `" agents-shipgate scan"` has `shlex.split(...)[0] == " agents-shipgate"` yet published a clean `agents-shipgate` invocation. Validation now runs on the authored value, and rejects every whitespace character except U+0020 anywhere in the string; only U+0020 is trimmed, which cannot change argv[0]. P2 — `is_addressable_gap` means target *or* command, so it could not guard target prose: a command-only row wrote `Regenerate the base report at .` into a durable instruction. The clause is gated on the target string. P2 — the instruction cap ran before de-duplication, so three warnings rendering as one placeholder consumed the budget and a fourth, visible mechanism disappeared. Placeholders now carry a digest of the raw bytes so distinct warnings stay distinct, rendered groups are de-duplicated before the cap, and readable diagnostics are ranked ahead of unprintable stand-ins. The cap is still three; raw warnings and the count are unchanged. P1 — the prompt's stale-base exception overrode the authority fields the line above it tells the reader to trust. That row is `requires_human_review: true`, `auto_apply: false`, `suggested_patch_kind: manual`, and `build_fix_task` emits `actor=human` / `safe_to_attempt=false`. The "Run it" instruction is removed: a command tells the human what to run, and agent authority comes only from `fix_task.actor == "coding_agent"` with `safe_to_attempt`. P2 — the canonical contract carried two competing definitions. The original block is rewritten rather than appended to, and STABILITY.md, the report-reading primer, the changelog, and the module docstrings now describe the same non-mutating, command-capable behaviour. Mirrors and `llms-full.txt` regenerated. Thresholds, raw `source_warning_count`, and the trust-root gate are unchanged. Co-Authored-By: Claude Opus 5 * fix(agent-mode): address Codex review 6 on #362 actionability ranking Five findings from review #4942947250, each reproduced before it was fixed. P1 — trimming a validated command could still break it. `printf foo\ ` is a two-token command whose second argument ends in a space; stripping U+0020 left `printf foo\`, which `shlex.split` refuses to parse, and the broken value propagated through Reason, Improve, the summary, fix-task instructions, `allowed_repairs[].command`, and the agent-result copies. A publishable command is now published byte for byte — no trimming at all. P1 — the display encoding was neither injective nor spoof-proof. `a\nb.yaml` and the literal filename `ab.yaml` rendered identically, and an embedded ZWSP passed through invisibly so `shipgate​.yaml` impersonated `shipgate.yaml`. Identity-bearing values now escape the introducer `<` as well, and escape Default_Ignorable code points rather than passing them through, so the rendering is reversible (`undisplay_literal` proves it) and no two repository objects can render the same way. Prose keeps `<` as ordinary punctuation — escaping it there mangled `