Name what a bound tool reaches in diff --application rows (#872) - #906
Merged
Merged
Conversation
Each side of a `diff --application` row now carries `reach`: the outbound `requests`, `httpx`, `aiohttp` and `urllib` calls the tool's own code and its same-scope helpers make, read statically up to three calls deep. For each call it records: - the method and URL template; - literal request fields, and the literals a field is chosen from; - which model-supplied parameters flow where; - the environment variables sent as credentials, by name only. Every call the read cannot follow is a named limit. `effect_evidence` is the engine's own `assess_tool_semantics` over the tool, with the reach as one more structural source (`source_http_call`). `read` is claimed only when every call was followed, every outbound call reads, and no limit was hit. What `read` can rest on is bounded by structural rules: - a client built elsewhere is a limit; - module state changed anywhere in the scope, under any name, is not taken as written; - a patch to the HTTP stack anywhere in the scope is a limit on every sending tool, however the stack was reached: aliases, re-exports, holders, introspection, copies. An adversarial reviewer ran 25 rounds against this. Each round's P0/P1 was fixed with a regression test and confirmed against the wire method actually sent. The remaining known limits are documented in docs/application-comparison.md. For a Google ADK name constructed twice, `binding_location` names a construction that lists the tool, and `construction_sites` lists every one. `application_comparison_schema_version` is 0.2. On the pinned 134-PR corpus, rows, statuses and exit codes are unchanged, and run time is flat. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rd seconds The guard took 45 s in CI's suite (coverage on a shared runner) against 4 s locally, over its 20 s bound. It catches a blow-up, which on a 40,000-entry table is minutes, so 120 s keeps it meaningful. tests/shard_seconds.json is re-measured with the two new test files (test_tool_reach 10 s, test_application_diff_tool_reach 64 s), so they are balanced by measured time rather than by item count. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On #904 two of the three shards took 13 of their 15 minutes, and this PR's tests took one past the cap (suite (3), cancelled at 15 minutes after 4362 tests passed). As the workflow says: re-measure, then add a shard, rather than raise the timeout. The release-pipeline test now ties the matrix to SHIPGATE_TEST_SHARDS instead of pinning three. The `Protect main` ruleset requires suite (1)-(3); suite (4) needs adding there to be required too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes #872 (part of #868). Each side of a
diff --applicationrow now names what the bound tool reaches: the endpoint and method, the request fields, the credential it sends and which arguments the model chooses. It also carries effect evidence from the engine's own semantic assessment.On tensorflow/tensorflow#128063, the signature alone gave all three tools
writewith unknown evidence. Nowsubmit_pr_code_reviewis shown posting a review whoseeventthe model may set toAPPROVE, usingGITHUB_TOKEN, andget_pull_request_detailsonly reads.What changed
Reach (
inputs/tool_reach.py, new). The tool's own code and its same-scope helpers, up to three calls deep, are read statically and never run. Each outboundrequests,httpx,aiohttporurllibcall records:{pr_number},{env OWNER},{…});model_supplied: which parameters flow where;credential_sources: environment variable names only, never a value;Every call the read cannot follow is a named limit with its location.
Effect evidence.
effect_evidenceisassess_tool_semanticsover the tool, with the reach as one more structural source (source_http_call). It is not a second classifier. A write or delete call supports that effect.readis claimed only when:Structural rules bound what
readcan rest on (seedocs/application-comparison.md):setattr,vars(),__dict__,globals(),exec,sys.modules. The changed object is followed through names, parameters at any depth, returns, displays and loops. Anything that is not a module withholds nothing, so an ORM row'ssetattr(user, field, value)does not.Surface.
reach,effect_evidenceandconstruction_sitesare evidence outside the compared meaning, so they move no row or status.application_comparison_schema_versionis0.2. For a Google ADK name constructed twice, each row'sbinding_locationnames a construction that lists the tool, andconstruction_siteslists every one. Credentials and token-shaped values are never printed.Evidence (pinned 134-PR corpus)
Rows, statuses and exit codes are identical to main, and run time is flat. Of 96 row sides in 19 PRs:
Review
An independent reviewer agent ran 25 adversarial rounds, ending in a confirmation round with no P0/P1 left unfixed or undocumented, against frozen exports of this branch. Each round tried to make
readappear where the code in fact changes what a request sends. Every P0/P1 was fixed with a regression test that fails on the previous round, and was confirmed against runtime behaviour (the wire method actually sent). The rounds moved from direct patches (requests.get = …) to patches through aliases, re-exports, holders, introspection and copies. The last rounds found only contrived layouts.Two whole-repository censuses guarded against noise after every widening:
socket/sslpatches;Both were held at the same numbers across rounds. A third census over 183 whole clones flags 5. Four are real patches: ddtrace instrumenting
http.client, and apatch.object(requests.Session, "request", …)in three PRs. The fifth, code_puppy, is a fail-closed limit.Known limits (documented, not claimed). Each of these either reads as the scan's boundary or fails closed:
sa = setattr,partial(setattr, m),from operator import attrgetter as ag);dict(vars(m)),vars(m).items(),inspect.getmembers,gc.get_objects());getattrwith a computed name, andpydoc.locate;with/except … astargets, andglobals()["X"] = …rebinding;runpy.run_pathor a plugin imported by computed name;exec,eval,compile). Failing closed there would limit most of the 13 corpus scopes (of 99) that run an interpreter or a calculator this way;json._default_encoderand__defaults__mutation;__getattr__re-exports;import a.b as cwherea/__init__rebindsb;Noise is fail-closed. When a stack module or class is kept and an unrelated store ends in a stack name or goes through a
Request/Session/Clientattribute, the tool gets a limit, never a falseread. A synthetic web of 500 modules with cyclicfrom .x import *takes about 39 s to scan; no real repository in the corpus comes close.CI
The first CI run failed for reasons outside the change:
suite (3)timed out. All 4,362 tests passed, but the job hit its 15-minute cap and was cancelled. On Derive the application comparison scope from the change when --scope is omitted (#875) #904, two of the three shards were already at 13 of 15 minutes. Asci.ymlprescribes, the shard seconds are re-measured and the suite now runs in four shards.tests/test_release_pipeline.pynow ties the matrix toSHIPGATE_TEST_SHARDSinstead of pinning three.Needs a maintainer: the
Protect mainruleset requiressuite (1)–suite (3). Addsuite (4)there, and to the table indocs/release-runbook.md, so the fourth shard is required too.Verification
pytest(full suite,-n 8, 14,759 tests): all pass except three local, environmental failures:test_check_unmodelled_host_config_keyscases (local_settings_enabled_plugins) fail on a clean checkout of main as well. This machine's global git ignore holds**/.claude/settings.local.json, so the fixture'sgit commithas nothing to commit;test_setup_control.py::test_an_absent_manifest_is_not_reported_as_a_malformed_onepasses on its own. Under-n 8the hint it looks for is truncated because this worktree's absolute path is very long.scope_mutationstime on the heaviest scopes is within noise of round 23: QwenPaw 18.9 s, dd-trace 8.4 s, ms-agent 4.4 s, Upsonic 6.4 s.🤖 Generated with Claude Code