Skip to content

Name what a bound tool reaches in diff --application rows (#872) - #906

Merged
pengfei-threemoonslab merged 3 commits into
mainfrom
claude/issue-872-reach
Sep 30, 2026
Merged

pengfei-threemoonslab merged 3 commits into
mainfrom
claude/issue-872-reach

Conversation

@pengfei-threemoonslab

@pengfei-threemoonslab pengfei-threemoonslab commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Closes #872 (part of #868). Each side of a diff --application row now names what the bound tool reaches: the endpoint and method, the request fields, the credential it sends and which arguments the model chooses. It also carries effect evidence from the engine's own semantic assessment.

On tensorflow/tensorflow#128063, the signature alone gave all three tools write with unknown evidence. Now submit_pr_code_review is shown posting a review whose event the model may set to APPROVE, using GITHUB_TOKEN, and get_pull_request_details only reads.

What changed

Reach (inputs/tool_reach.py, new). The tool's own code and its same-scope helpers, up to three calls deep, are read statically and never run. Each outbound requests, httpx, aiohttp or urllib call records:

  • the method and URL template ({pr_number}, {env OWNER}, {…});
  • literal request fields, and the literals a field is chosen from, with the parameters that decide;
  • model_supplied: which parameters flow where;
  • credential_sources: environment variable names only, never a value;
  • for GraphQL, query or mutation, only at a GraphQL endpoint.

Every call the read cannot follow is a named limit with its location.

Effect evidence. effect_evidence is assess_tool_semantics over the tool, with the reach as one more structural source (source_http_call). It is not a second classifier. A write or delete call supports that effect. read is claimed only when:

  • there is at least one outbound call and every one reads;
  • no limit was hit;
  • nothing was truncated.

Structural rules bound what read can rest on (see docs/application-comparison.md):

  • A client or request built outside the calling function is a limit.
  • A value read out of a module-level dict, or a copy of one, is never taken as written.
  • A module constant or function changed elsewhere in the scope is unknown. That covers a change by name, and one under names the read cannot see: setattr, vars(), __dict__, globals(), exec, sys.modules. The changed object is followed through names, parameters at any depth, returns, displays and loops. Anything that is not a module withholds nothing, so an ORM row's setattr(user, field, value) does not.
  • A patch to an HTTP library anywhere in the scope is a limit on every sending tool.

Surface. reach, effect_evidence and construction_sites are evidence outside the compared meaning, so they move no row or status. application_comparison_schema_version is 0.2. For a Google ADK name constructed twice, each row's binding_location names a construction that lists the tool, and construction_sites lists every one. Credentials and token-shaped values are never printed.

Evidence (pinned 134-PR corpus)

Rows, statuses and exit codes are identical to main, and run time is flat. Of 96 row sides in 19 PRs:

  • 9 name outbound calls (3 PRs) and 5 name a credential;
  • 7 carry structural effect evidence (3 read, 4 write);
  • 65 name at least one limit, mostly database helpers, tracing and SDK clients.

Review

An independent reviewer agent ran 25 adversarial rounds, ending in a confirmation round with no P0/P1 left unfixed or undocumented, against frozen exports of this branch. Each round tried to make read appear where the code in fact changes what a request sends. Every P0/P1 was fixed with a regression test that fails on the previous round, and was confirmed against runtime behaviour (the wire method actually sent). The rounds moved from direct patches (requests.get = …) to patches through aliases, re-exports, holders, introspection and copies. The last rounds found only contrived layouts.

Two whole-repository censuses guarded against noise after every widening:

  • a stack-patch census over the 99 corpus scopes: 4 flagged, 3 of them real socket/ssl patches;
  • a namespace-wildcard census: 12 of 99 scopes, 65 of 183 clones.

Both were held at the same numbers across rounds. A third census over 183 whole clones flags 5. Four are real patches: ddtrace instrumenting http.client, and a patch.object(requests.Session, "request", …) in three PRs. The fifth, code_puppy, is a fail-closed limit.

Known limits (documented, not claimed). Each of these either reads as the scan's boundary or fails closed:

  • builtins and stdlib callables aliased or used as values (sa = setattr, partial(setattr, m), from operator import attrgetter as ag);
  • a module changed by a function outside the scope;
  • a namespace copied or iterated (dict(vars(m)), vars(m).items(), inspect.getmembers, gc.get_objects());
  • getattr with a computed name, and pydoc.locate;
  • with/except … as targets, and globals()["X"] = … rebinding;
  • name-based callee matching;
  • a patch outside the compared scope, including one reached through a relative import that climbs above the scope root, runpy.run_path or a plugin imported by computed name;
  • code run from a string (exec, eval, compile). Failing closed there would limit most of the 13 corpus scopes (of 99) that run an interpreter or a calculator this way;
  • json._default_encoder and __defaults__ mutation;
  • telemetry and log-shipping wrappers;
  • function-level relative imports and PEP 562 __getattr__ re-exports;
  • relative imports in a class body;
  • import a.b as c where a/__init__ rebinds b;
  • an unbound-method call on the method's own class;
  • a factory named like a fresh constructor.

Noise is fail-closed. When a stack module or class is kept and an unrelated store ends in a stack name or goes through a Request/Session/Client attribute, the tool gets a limit, never a false read. A synthetic web of 500 modules with cyclic from .x import * takes about 39 s to scan; no real repository in the corpus comes close.

CI

The first CI run failed for reasons outside the change:

  • suite (3) timed out. All 4,362 tests passed, but the job hit its 15-minute cap and was cancelled. On Derive the application comparison scope from the change when --scope is omitted (#875) #904, two of the three shards were already at 13 of 15 minutes. As ci.yml prescribes, the shard seconds are re-measured and the suite now runs in four shards. tests/test_release_pipeline.py now ties the matrix to SHIPGATE_TEST_SHARDS instead of pinning three.
  • A reach perf guard exceeded its bound. It took 45 s in CI (coverage on a shared runner) against 4 s locally. Its bound is now 120 s: it catches a blow-up, which on its 40,000-entry table would be minutes.

Needs a maintainer: the Protect main ruleset requires suite (1)–suite (3). Add suite (4) there, and to the table in docs/release-runbook.md, so the fourth shard is required too.

Verification

  • pytest (full suite, -n 8, 14,759 tests): all pass except three local, environmental failures:
    • two test_check_unmodelled_host_config_keys cases (local_settings_enabled_plugins) fail on a clean checkout of main as well. This machine's global git ignore holds **/.claude/settings.local.json, so the fixture's git commit has nothing to commit;
    • test_setup_control.py::test_an_absent_manifest_is_not_reported_as_a_malformed_one passes on its own. Under -n 8 the hint it looks for is truncated because this worktree's absolute path is very long.
  • Corpus (134 PRs, main vs this branch): rows, statuses and exit codes are identical, and run time is flat.
  • scope_mutations time on the heaviest scopes is within noise of round 23: QwenPaw 18.9 s, dd-trace 8.4 s, ms-agent 4.4 s, Upsonic 6.4 s.

🤖 Generated with Claude Code

Each side of a `diff --application` row now carries `reach`: the outbound
`requests`, `httpx`, `aiohttp` and `urllib` calls the tool's own code and its
same-scope helpers make, read statically up to three calls deep. For each
call it records:
- the method and URL template;
- literal request fields, and the literals a field is chosen from;
- which model-supplied parameters flow where;
- the environment variables sent as credentials, by name only.

Every call the read cannot follow is a named limit. `effect_evidence` is the
engine's own `assess_tool_semantics` over the tool, with the reach as one
more structural source (`source_http_call`). `read` is claimed only when
every call was followed, every outbound call reads, and no limit was hit.

What `read` can rest on is bounded by structural rules:
- a client built elsewhere is a limit;
- module state changed anywhere in the scope, under any name, is not taken
  as written;
- a patch to the HTTP stack anywhere in the scope is a limit on every
  sending tool, however the stack was reached: aliases, re-exports,
  holders, introspection, copies.

An adversarial reviewer ran 25 rounds against this. Each round's P0/P1 was
fixed with a regression test and confirmed against the wire method actually
sent. The remaining known limits are documented in
docs/application-comparison.md.

For a Google ADK name constructed twice, `binding_location` names a
construction that lists the tool, and `construction_sites` lists every one.
`application_comparison_schema_version` is 0.2.

On the pinned 134-PR corpus, rows, statuses and exit codes are unchanged, and
run time is flat.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rd seconds

The guard took 45 s in CI's suite (coverage on a shared runner) against 4 s
locally, over its 20 s bound. It catches a blow-up, which on a 40,000-entry
table is minutes, so 120 s keeps it meaningful.

tests/shard_seconds.json is re-measured with the two new test files
(test_tool_reach 10 s, test_application_diff_tool_reach 64 s), so they are
balanced by measured time rather than by item count.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On #904 two of the three shards took 13 of their 15 minutes, and this PR's
tests took one past the cap (suite (3), cancelled at 15 minutes after
4362 tests passed). As the workflow says: re-measure, then add a shard,
rather than raise the timeout. The release-pipeline test now ties the matrix
to SHIPGATE_TEST_SHARDS instead of pinning three.

The `Protect main` ruleset requires suite (1)-(3); suite (4) needs adding
there to be required too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@pengfei-threemoonslab
pengfei-threemoonslab merged commit 6ced6f7 into main Sep 30, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Application diff rows should name what a bound tool reaches: endpoint, action, credential, model-supplied arguments

1 participant