Skip to content

Derive the application comparison scope from the change when --scope is omitted (#875) - #904

Merged
pengfei-threemoonslab merged 4 commits into
mainfrom
claude/issue-875-derived-scope
Sep 29, 2026
Merged

pengfei-threemoonslab merged 4 commits into
mainfrom
claude/issue-875-derived-scope

Conversation

@pengfei-threemoonslab

@pengfei-threemoonslab pengfei-threemoonslab commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Closes #875. Without --scope, diff --application now derives the comparison scope from the change instead of comparing the repository root.

Neither obvious scope works for someone reviewing a PR in an unfamiliar repository:

  • The repository root is partial whenever any file anywhere uses another framework.
  • The changed directory often holds no agent. On feat: Gmail API integration mezgoodle/Vesta#58, the PR changes backend/app/services/ and the agents live in backend/app/agents/, so --scope backend/app/services answers not_established.

What changed

Relating the change to agents. Each changed Python file is related to the OpenAI Agents SDK and Google ADK agent files that import it, or that it imports, within six import hops, on each side of the change.

  • A package's __init__.py on the way counts, and so does a literal importlib.import_module(...).
  • An agent file is one that does any of these:
    • constructs an agent class under any name it is imported as;
    • subclasses an agent class;
    • copies an agent with capabilities of its own;
    • changes an agent's capabilities after construction, counted with fix(diff --application): an unobserved agent is not "no change" (#876) #880's own _capability_changes predicate, so the scope and the reader agree;
    • builds its agent through another file's factory.
  • A test file is related only when an agent imports it, and it never widens a scope.

Choosing the scope. Each related group is compared in the outermost package that holds the change, its agents and what they import. Vesta#58 derives backend/app.

  • Independent applications in a monorepo are separate comparisons under comparisons, never the root.
  • A relocated application is one comparison.
  • A change that touches no supported agent is an explicit not_established answer naming the files considered.

Never a silent compared. Each of these is named in scope_selection.limits and makes the result partial:

  • a changed file not related to a module that builds an agent, with why;
  • a changed link or submodule;
  • an agent outside the compared scopes that reaches the change;
  • a consumer: a module that imports the change and a builder that builds on request, and builds, copies or rewires an agent, calls the repository's code with its own arguments, or sets its module state. It is exempt only when one compared scope holds all three.
  • any bound reached: 6 hops, 2000 files to relate, 5000 files or 64 MB to find agent files.

Surface. --json records scope_selection: the mode (derived or explicit), the scopes, why, and each scope's relations. An explicit --scope always wins, --scope . keeps the root, and --base-scope needs --scope. Everything is read from the two commits' objects and never run. Partial clones are refused with the established hydration message.

Evidence (pinned 48-PR application corpus)

scope rows established
derived (this PR) 296
repository root 296
common parent of the changed wiring files 18
  • The derived scope is narrower than the root for 20 PRs.
  • For 7 PRs it answers that the change touches no supported agent; the root reported unrelated agents there.
  • Median run time is 13 s, the same as the root.

Review

Twelve adversarial review rounds, each with fixtures and a prior-engine comparison. Every fix has a test in tests/test_application_scope.py that fails on the commit before it.

Rounds 6–7 settled the design: a reverse import closure that walks through agent files, and builder qualification. Two rules were tried and reverted after regressing true consumers:

  • skipping entry scripts whose imports are all in scope, which dropped FEDOT.MAS's examples;
  • stopping the closure at agent files.

Round 11, after rebasing onto #880, replaced this module's own rewire predicate with #880's _capability_changes. A slice store into an agent's tools, or a change made through a handle to that list, now counts for the scope as it does for the reader.

Round 12 found no P0/P1. It checked:

  • 7,570 corpus files through the new predicate, with no exception and none over 1 s;
  • 27 adversarial syntax snippets;
  • every earlier round's fixtures, with no regression.

Known, non-blocking:

Verification

🤖 Generated with Claude Code

pengfei-threemoonslab added a commit that referenced this pull request Sep 28, 2026
The suite's three CI shards were balanced by collected item count. Item count
is a poor proxy for time: a file of forty git-fixture tests costs more than a
file of four hundred pure ones.

- On `main`, shard 3 took 13 of its 15 minutes while shards 1 and 2 took 8.
- Any new test file reshuffled most files. #904's one new file moved 298 of
  363, which put shard 1 at 14 minutes and cancelled shard 3 at the cap.

Shards are now balanced by measured seconds per file, in
`tests/shard_seconds.json`. `scripts/measure_shard_seconds.py` writes that file
from a `--junitxml` run. A file without a measurement costs its item count at
the measured seconds per item. A stale measurement only unbalances, never drops
a file. The union property, determinism and fail-loud rules are unchanged.

On one full-suite measurement:
- balancing by count gives main 25/30/45% of the work, matching CI's
  6.8/6.8/11.6-minute shards;
- balancing by time gives 33/33/33%, for main and for #904.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@pengfei-threemoonslab

Copy link
Copy Markdown
Contributor Author

The suite (3) cancellation was a shard-balance timeout, not a test failure: the three shards were balanced by test count, and this PR's new test file reshuffled 298 of 363 files, putting shards 1 and 3 at the 15-minute cap. The fix is #905 (balance shards by measured time per file), cherry-picked here as 5122984 so this PR's CI can pass now; it drops out on rebase once #905 merges.

🤖 Generated with Claude Code

… when --scope is omitted (#875)

Without --scope, each changed Python file is related to the OpenAI Agents SDK
and Google ADK agent files that import it, or that it imports, within six
import hops. It is read from the two commits' objects on each side. A
package's __init__.py and a literal importlib.import_module count. An agent
file does one of these:
- constructs or subclasses an agent class;
- copies an agent with capabilities of its own;
- changes an agent's capabilities after construction;
- builds its agent through another file's factory.

Each related group is compared in the outermost package that holds the
change, its agents and what they import. Vesta#58 derives backend/app.

- Independent applications become separate comparisons under `comparisons`,
  never the repository root.
- A relocated application is one comparison.
- A change that touches no supported agent is an explicit not_established
  answer naming the files considered.
- These are named in scope_selection.limits and make the result partial:
  - a changed file unrelated to an agent-building module, with why;
  - a changed link or submodule;
  - an agent outside the scopes that reaches the change;
  - a consumer: a module that imports the change and an on-request builder
    and builds, copies or rewires an agent, calls repository code with its
    own arguments, or sets its module state. It is exempt only when one
    compared scope holds all three.
  - any bound reached.

`scope_selection` records the mode, the scopes and why. An explicit --scope
always wins, `--scope .` keeps the root, and --base-scope needs --scope.
Partial clones are refused with the established hydration message before
anything reads the tree.

On the pinned 48-PR corpus, derived scopes establish the same 296 rows as
the root. The changed-wiring directory establishes 18.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ew 11)

The derived scope decided whether a module rewires an agent with its own
predicate, which missed spellings the SDK reader already counts: a slice store
into an agent's tools list, and a change made through a handle to that list.
It now asks #880's `_capability_changes`, so the reader and the scope derivation
agree on what changes an agent.

On the 48-case corpus, one extra changed file in alliance-genome #842 and #860
is named as unrelated (`prompt_builder.py` reads `agent.tools` into a local);
nothing else moves. Of 156 fixtures, the four this round added move to partial.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The suite's three CI shards were balanced by collected item count. Item count
is a poor proxy for time: a file of forty git-fixture tests costs more than a
file of four hundred pure ones.

- On `main`, shard 3 took 13 of its 15 minutes while shards 1 and 2 took 8.
- Any new test file reshuffled most files. #904's one new file moved 298 of
  363, which put shard 1 at 14 minutes and cancelled shard 3 at the cap.

Shards are now balanced by measured seconds per file, in
`tests/shard_seconds.json`. `scripts/measure_shard_seconds.py` writes that file
from a `--junitxml` run. A file without a measurement costs its item count at
the measured seconds per item. A stale measurement only unbalances, never drops
a file. The union property, determinism and fail-loud rules are unchanged.

On one full-suite measurement:
- balancing by count gives main 25/30/45% of the work, matching CI's
  6.8/6.8/11.6-minute shards;
- balancing by time gives 33/33/33%, for main and for #904.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Three #865 cases change only `README.md` and assert what a full-tree
comparison says about an untouched agent. With the scope derived from the
change (#875), a README-only change touches no agent and is `not_established`,
so these cases now pass `--scope .`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@pengfei-threemoonslab
pengfei-threemoonslab force-pushed the claude/issue-875-derived-scope branch from 5122984 to 1a0a40b Compare September 29, 2026 01:06
@pengfei-threemoonslab
pengfei-threemoonslab merged commit 524f0a7 into main Sep 29, 2026
12 checks passed
pengfei-threemoonslab added a commit that referenced this pull request Sep 29, 2026
On #904 two of the three shards took 13 of their 15 minutes, and this PR's
tests took one past the cap (suite (3), cancelled at 15 minutes after
4362 tests passed). As the workflow says: re-measure, then add a shard,
rather than raise the timeout. The release-pipeline test now ties the matrix
to SHIPGATE_TEST_SHARDS instead of pinning three.

The `Protect main` ruleset requires suite (1)-(3); suite (4) needs adding
there to be required too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pengfei-threemoonslab added a commit that referenced this pull request Sep 30, 2026
* Name what a bound tool reaches in diff --application rows (#872)

Each side of a `diff --application` row now carries `reach`: the outbound
`requests`, `httpx`, `aiohttp` and `urllib` calls the tool's own code and its
same-scope helpers make, read statically up to three calls deep. For each
call it records:
- the method and URL template;
- literal request fields, and the literals a field is chosen from;
- which model-supplied parameters flow where;
- the environment variables sent as credentials, by name only.

Every call the read cannot follow is a named limit. `effect_evidence` is the
engine's own `assess_tool_semantics` over the tool, with the reach as one
more structural source (`source_http_call`). `read` is claimed only when
every call was followed, every outbound call reads, and no limit was hit.

What `read` can rest on is bounded by structural rules:
- a client built elsewhere is a limit;
- module state changed anywhere in the scope, under any name, is not taken
  as written;
- a patch to the HTTP stack anywhere in the scope is a limit on every
  sending tool, however the stack was reached: aliases, re-exports,
  holders, introspection, copies.

An adversarial reviewer ran 25 rounds against this. Each round's P0/P1 was
fixed with a regression test and confirmed against the wire method actually
sent. The remaining known limits are documented in
docs/application-comparison.md.

For a Google ADK name constructed twice, `binding_location` names a
construction that lists the tool, and `construction_sites` lists every one.
`application_comparison_schema_version` is 0.2.

On the pinned 134-PR corpus, rows, statuses and exit codes are unchanged, and
run time is flat.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(#872): bound the large-constant guard for CI, and re-measure shard seconds

The guard took 45 s in CI's suite (coverage on a shared runner) against 4 s
locally, over its 20 s bound. It catches a blow-up, which on a 40,000-entry
table is minutes, so 120 s keeps it meaningful.

tests/shard_seconds.json is re-measured with the two new test files
(test_tool_reach 10 s, test_application_diff_tool_reach 64 s), so they are
balanced by measured time rather than by item count.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: split the suite into four shards

On #904 two of the three shards took 13 of their 15 minutes, and this PR's
tests took one past the cap (suite (3), cancelled at 15 minutes after
4362 tests passed). As the workflow says: re-measure, then add a shard,
rather than raise the timeout. The release-pipeline test now ties the matrix
to SHIPGATE_TEST_SHARDS instead of pinning three.

The `Protect main` ruleset requires suite (1)-(3); suite (4) needs adding
there to be required too.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Application diff: derive the comparison scope from the change when --scope is omitted

1 participant