Derive the application comparison scope from the change when --scope is omitted (#875) - #904
Merged
Merged
Conversation
pengfei-threemoonslab
added a commit
that referenced
this pull request
Sep 28, 2026
The suite's three CI shards were balanced by collected item count. Item count is a poor proxy for time: a file of forty git-fixture tests costs more than a file of four hundred pure ones. - On `main`, shard 3 took 13 of its 15 minutes while shards 1 and 2 took 8. - Any new test file reshuffled most files. #904's one new file moved 298 of 363, which put shard 1 at 14 minutes and cancelled shard 3 at the cap. Shards are now balanced by measured seconds per file, in `tests/shard_seconds.json`. `scripts/measure_shard_seconds.py` writes that file from a `--junitxml` run. A file without a measurement costs its item count at the measured seconds per item. A stale measurement only unbalances, never drops a file. The union property, determinism and fail-loud rules are unchanged. On one full-suite measurement: - balancing by count gives main 25/30/45% of the work, matching CI's 6.8/6.8/11.6-minute shards; - balancing by time gives 33/33/33%, for main and for #904. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
The 🤖 Generated with Claude Code |
… when --scope is omitted (#875) Without --scope, each changed Python file is related to the OpenAI Agents SDK and Google ADK agent files that import it, or that it imports, within six import hops. It is read from the two commits' objects on each side. A package's __init__.py and a literal importlib.import_module count. An agent file does one of these: - constructs or subclasses an agent class; - copies an agent with capabilities of its own; - changes an agent's capabilities after construction; - builds its agent through another file's factory. Each related group is compared in the outermost package that holds the change, its agents and what they import. Vesta#58 derives backend/app. - Independent applications become separate comparisons under `comparisons`, never the repository root. - A relocated application is one comparison. - A change that touches no supported agent is an explicit not_established answer naming the files considered. - These are named in scope_selection.limits and make the result partial: - a changed file unrelated to an agent-building module, with why; - a changed link or submodule; - an agent outside the scopes that reaches the change; - a consumer: a module that imports the change and an on-request builder and builds, copies or rewires an agent, calls repository code with its own arguments, or sets its module state. It is exempt only when one compared scope holds all three. - any bound reached. `scope_selection` records the mode, the scopes and why. An explicit --scope always wins, `--scope .` keeps the root, and --base-scope needs --scope. Partial clones are refused with the established hydration message before anything reads the tree. On the pinned 48-PR corpus, derived scopes establish the same 296 rows as the root. The changed-wiring directory establishes 18. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ew 11) The derived scope decided whether a module rewires an agent with its own predicate, which missed spellings the SDK reader already counts: a slice store into an agent's tools list, and a change made through a handle to that list. It now asks #880's `_capability_changes`, so the reader and the scope derivation agree on what changes an agent. On the 48-case corpus, one extra changed file in alliance-genome #842 and #860 is named as unrelated (`prompt_builder.py` reads `agent.tools` into a local); nothing else moves. Of 156 fixtures, the four this round added move to partial. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The suite's three CI shards were balanced by collected item count. Item count is a poor proxy for time: a file of forty git-fixture tests costs more than a file of four hundred pure ones. - On `main`, shard 3 took 13 of its 15 minutes while shards 1 and 2 took 8. - Any new test file reshuffled most files. #904's one new file moved 298 of 363, which put shard 1 at 14 minutes and cancelled shard 3 at the cap. Shards are now balanced by measured seconds per file, in `tests/shard_seconds.json`. `scripts/measure_shard_seconds.py` writes that file from a `--junitxml` run. A file without a measurement costs its item count at the measured seconds per item. A stale measurement only unbalances, never drops a file. The union property, determinism and fail-loud rules are unchanged. On one full-suite measurement: - balancing by count gives main 25/30/45% of the work, matching CI's 6.8/6.8/11.6-minute shards; - balancing by time gives 33/33/33%, for main and for #904. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Three #865 cases change only `README.md` and assert what a full-tree comparison says about an untouched agent. With the scope derived from the change (#875), a README-only change touches no agent and is `not_established`, so these cases now pass `--scope .`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pengfei-threemoonslab
force-pushed
the
claude/issue-875-derived-scope
branch
from
September 29, 2026 01:06
5122984 to
1a0a40b
Compare
pengfei-threemoonslab
added a commit
that referenced
this pull request
Sep 29, 2026
On #904 two of the three shards took 13 of their 15 minutes, and this PR's tests took one past the cap (suite (3), cancelled at 15 minutes after 4362 tests passed). As the workflow says: re-measure, then add a shard, rather than raise the timeout. The release-pipeline test now ties the matrix to SHIPGATE_TEST_SHARDS instead of pinning three. The `Protect main` ruleset requires suite (1)-(3); suite (4) needs adding there to be required too. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pengfei-threemoonslab
added a commit
that referenced
this pull request
Sep 30, 2026
* Name what a bound tool reaches in diff --application rows (#872) Each side of a `diff --application` row now carries `reach`: the outbound `requests`, `httpx`, `aiohttp` and `urllib` calls the tool's own code and its same-scope helpers make, read statically up to three calls deep. For each call it records: - the method and URL template; - literal request fields, and the literals a field is chosen from; - which model-supplied parameters flow where; - the environment variables sent as credentials, by name only. Every call the read cannot follow is a named limit. `effect_evidence` is the engine's own `assess_tool_semantics` over the tool, with the reach as one more structural source (`source_http_call`). `read` is claimed only when every call was followed, every outbound call reads, and no limit was hit. What `read` can rest on is bounded by structural rules: - a client built elsewhere is a limit; - module state changed anywhere in the scope, under any name, is not taken as written; - a patch to the HTTP stack anywhere in the scope is a limit on every sending tool, however the stack was reached: aliases, re-exports, holders, introspection, copies. An adversarial reviewer ran 25 rounds against this. Each round's P0/P1 was fixed with a regression test and confirmed against the wire method actually sent. The remaining known limits are documented in docs/application-comparison.md. For a Google ADK name constructed twice, `binding_location` names a construction that lists the tool, and `construction_sites` lists every one. `application_comparison_schema_version` is 0.2. On the pinned 134-PR corpus, rows, statuses and exit codes are unchanged, and run time is flat. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * test(#872): bound the large-constant guard for CI, and re-measure shard seconds The guard took 45 s in CI's suite (coverage on a shared runner) against 4 s locally, over its 20 s bound. It catches a blow-up, which on a 40,000-entry table is minutes, so 120 s keeps it meaningful. tests/shard_seconds.json is re-measured with the two new test files (test_tool_reach 10 s, test_application_diff_tool_reach 64 s), so they are balanced by measured time rather than by item count. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * ci: split the suite into four shards On #904 two of the three shards took 13 of their 15 minutes, and this PR's tests took one past the cap (suite (3), cancelled at 15 minutes after 4362 tests passed). As the workflow says: re-measure, then add a shard, rather than raise the timeout. The release-pipeline test now ties the matrix to SHIPGATE_TEST_SHARDS instead of pinning three. The `Protect main` ruleset requires suite (1)-(3); suite (4) needs adding there to be required too.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Closes #875. Without
--scope,diff --applicationnow derives the comparison scope from the change instead of comparing the repository root.Neither obvious scope works for someone reviewing a PR in an unfamiliar repository:
partialwhenever any file anywhere uses another framework.backend/app/services/and the agents live inbackend/app/agents/, so--scope backend/app/servicesanswersnot_established.What changed
Relating the change to agents. Each changed Python file is related to the OpenAI Agents SDK and Google ADK agent files that import it, or that it imports, within six import hops, on each side of the change.
__init__.pyon the way counts, and so does a literalimportlib.import_module(...)._capability_changespredicate, so the scope and the reader agree;Choosing the scope. Each related group is compared in the outermost package that holds the change, its agents and what they import. Vesta#58 derives
backend/app.comparisons, never the root.not_establishedanswer naming the files considered.Never a silent
compared. Each of these is named inscope_selection.limitsand makes the resultpartial:Surface.
--jsonrecordsscope_selection: the mode (derivedorexplicit), the scopes, why, and each scope's relations. An explicit--scopealways wins,--scope .keeps the root, and--base-scopeneeds--scope. Everything is read from the two commits' objects and never run. Partial clones are refused with the established hydration message.Evidence (pinned 48-PR application corpus)
Review
Twelve adversarial review rounds, each with fixtures and a prior-engine comparison. Every fix has a test in
tests/test_application_scope.pythat fails on the commit before it.Rounds 6–7 settled the design: a reverse import closure that walks through agent files, and builder qualification. Two rules were tried and reverted after regressing true consumers:
Round 11, after rebasing onto #880, replaced this module's own rewire predicate with #880's
_capability_changes. A slice store into an agent's tools, or a change made through a handle to that list, now counts for the scope as it does for the reader.partial. Corpus time went from 291 s to 278 s.comparedtopartial.Round 12 found no P0/P1. It checked:
Known, non-blocking:
t = agent.tools; t = [...]) as a change. That adds the alliance-genome entry above.bot.tools, bot.name = ...),for/withattribute targets, or walrus/globalhandles. The derivation and the reader agree on these cases, since they now share one predicate. That belongs in a follow-up on fix(diff --application): an unobserved agent is not "no change" (#876) #880.probar_conversacion.py), and the unrelated-file limit on 21 of 48 corpus repos.Verification
pytest -n 8): 14,389 passed, 7 skipped. The only failures are the two environmentaltest_check_unmodelled_host_config_keys.py[local_settings_enabled_plugins-*]cases, which fail the same way on untouched main. Three Google ADK: trace bounded FunctionTool factories into agent tool bindings #865 cases that change only a README now pin--scope .(a README-only change derives no scope).ruff check .passes.🤖 Generated with Claude Code