Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 6 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
name: CI

on:
Expand All @@ -22,7 +22,12 @@
# (add a shard).
#
# `conftest.py` assigns whole test *files* to shards, deterministically,
# from the collection alone — nothing is exchanged between these jobs.
# from the collection and the measured seconds per file in
# `tests/shard_seconds.json` — nothing is exchanged between these jobs.
# Balancing item counts instead left one shard at 13 of its 15 minutes on
# `main` while another took 8, and any new test file reshuffled the rest
# (#904). When a shard nears the cap, re-measure with
# `scripts/measure_shard_seconds.py` before adding a shard.
# `tests/test_shard_partition.py` asserts the union of the shards is the
# suite, because a partition that silently drops a file leaves every job
# green while a test stops running.
Expand Down
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@
### Application comparison without prior setup

- Add `diff --application` for OpenAI Agents SDK and Google ADK source-observed per-agent wiring, with exact base/head refs and independently selected scopes. No manifest, saved baseline or authored declarations are needed. The advisory `application_comparison_schema_version: "0.1"` result records before/after evidence, scoped coverage gaps and explicit uncertain candidates; it supplies no release verdict or merge permission. Includes scoped Git materialization, partial-clone recovery, and definition lookup using reader-resolved Python symbols and locations. See [application comparison](docs/application-comparison.md). (#871)
- `diff --application` without `--scope` now derives the scope from the change instead of comparing the repository root. Each changed Python file is related to the OpenAI Agents SDK and Google ADK files that import it, or that it imports, within six import hops (a package's `__init__.py` included); a changed test relates only when an agent imports it. Agent files are those that construct an agent class under any name it is imported as, subclass one, copy one with capabilities of its own, or change an agent's capabilities after construction. A changed file not related to a module that builds an agent is named, with why (no import path, a path longer than six hops, only a rewiring module, a file too large to read), and makes the result `partial` even inside a compared scope; so does a changed link to a module or directory, a changed submodule, an agent outside the compared scopes whose imports go past the bound, or a module that imports the change (directly, through a package or agent file's re-export, or through other modules), imports a builder whose agent's capabilities are not a fixed list of its own names, and builds, copies or rewires an agent itself, calls the repository's code with its own arguments, or sets its module state, unless one compared scope holds all three. A module that builds its agent through a builder's factory is an agent file. A relocated application is one comparison, whatever it holds. Each related group is compared in the outermost package that holds it. On mezgoodle/Vesta#58 a change to `backend/app/services/` is compared in `backend/app`, where its agents are, rather than in the changed directory, which holds none. Independent applications in a monorepo are separate comparisons under `comparisons`, never the root. A change that touches no supported agent is an explicit `not_established` answer naming the files considered. A bound reached is named and makes the result `partial`. `--json` records `scope_selection` (`derived` or `explicit`, the scopes and why). `--scope .` keeps the root, an explicit `--scope` always wins, and `--base-scope` now needs `--scope`. On the pinned corpus of 48 application PRs, derived scopes establish the same 296 rows the repository root does, where the changed-wiring directory establishes 18. The derived scope is narrower than the root for 20 PRs. For 7 PRs it answers that the change touches no supported agent; the root reported unrelated agents there. The median run takes 13 s, the same as the root. (#875)
- `diff --application` no longer reads another library's `Agent` or `function_tool` as the OpenAI Agents SDK's, and no longer reports an agent it stopped observing as removed. On speechmatics/speechmatics-academy#142, a LiveKit voice agent (`from livekit.agents import Agent, function_tool`) moved its tools into an `Agent` subclass that passes them through `super().__init__`. The comparison printed `compared` with four `REMOVED agent → …` rows although the same four tools were still bound. It now prints `not_established`: no supported application agent on either side. (#871 follow-up; related #580, #864, #865)
- **Framework identity.** Discovery and the SDK reader count `function_tool` and `Agent` as the SDK's only when the name is imported from the absolute `agents`/`openai_agents` package, or is not imported at all (the existing reading, unless a wildcard import from another module could supply it). The name is resolved in the scope that uses it, as Python resolves it, so an import in a sibling function never decides it, and a parameter or local assignment of that name is not the SDK's. A name imported from anywhere else — `livekit.agents`, a relative `.agents` package — is another library's. A LiveKit file is no longer an SDK candidate, and a file using both reads only the SDK's symbols. A relative import is no longer any framework's import signal. `scan` of a declared SDK source reads the same way.
- **Unobserved is not removed.** One side may observe an agent while the other side's same file still assigns or imports its name (`from factory import agent`), or passes it as `name=`, through a construction the reader does not support: an `Agent` subclass, a factory, `Agent[Context](...)` or `.clone()`. That side then records a coverage gap for the agent. The result is `partial`, and its rows are `not_established` with `candidate_change` `removed` or `added`. An agent referenced only in another agent's `handoffs` is not an observed construction. An agent whose name is gone from the file is still an established removal, and rows for other agents are unaffected. No schema changes.
Expand Down
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,8 @@ agents-shipgate diff --application --workspace /path/to/repo --base BASE_SHA --h
```

It shows source-observed changes per agent, including before/after signatures
and source locations. See [application comparison](docs/application-comparison.md)
and source locations, in the package that holds the changed code's agents unless
`--scope` names one. See [application comparison](docs/application-comparison.md)
for scoped applications, moves, exact refs and coverage limits. Version availability
is recorded in the [CHANGELOG entry](CHANGELOG.md#application-comparison-without-prior-setup);
while it is under Unreleased, use a source build containing the feature.
Expand Down
70 changes: 60 additions & 10 deletions ci_sharding.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,17 +8,50 @@
from anywhere and has one definition.

Pure. It reads nothing, writes nothing, and takes no environment: the caller
supplies the collection and gets back an assignment. That is what lets
``tests/test_shard_partition.py`` assert the properties directly.
supplies the collection (and the measured seconds, when it has them) and gets
back an assignment. That is what lets ``tests/test_shard_partition.py`` assert
the properties directly.
"""

from __future__ import annotations

import json
import math
from collections.abc import Mapping
from pathlib import Path

#: Measured seconds per test file, written by
#: ``scripts/measure_shard_seconds.py`` from a full ``--junitxml`` run.
SECONDS_FILE = Path(__file__).resolve().parent / "tests" / "shard_seconds.json"

def shard_assignment(paths: Mapping[str, int], shards: int) -> dict[str, int]:
"""Assign whole test *files* to shards, balancing collected item counts.

def load_seconds(path: Path = SECONDS_FILE) -> dict[str, float]:
"""The measured seconds per file, or nothing when there is no measurement.

A missing or unreadable file balances on item counts alone, as before:
the measurement only ever improves the balance, never the coverage.
"""

try:
payload = json.loads(path.read_text(encoding="utf-8"))
except (OSError, ValueError):
return {}
files = payload.get("files") if isinstance(payload, dict) else None
if not isinstance(files, dict):
return {}
return {
str(name): float(value)
for name, value in files.items()
if isinstance(value, int | float) and not isinstance(value, bool) and value >= 0
}


def shard_assignment(
paths: Mapping[str, int],
shards: int,
seconds: Mapping[str, float] | None = None,
) -> dict[str, int]:
"""Assign whole test *files* to shards, balancing their expected time.

Two properties, and both are load-bearing.

Expand All @@ -32,15 +65,32 @@ def shard_assignment(paths: Mapping[str, int], shards: int) -> dict[str, int]:
is exchanged between jobs, so the union of the shards is exactly the suite
and no test can fall between two of them — which a test asserts.

Item count is a proxy for time, and an imperfect one; it is used because it
is free and needs no stored measurements to go stale. Balance is checked in
``tests/test_shard_partition.py`` rather than assumed.
**Measured time where there is one.** Item count alone was a poor proxy:
a file of forty git-fixture tests costs more than a file of four hundred
pure ones. Balancing counts left one shard at 13 of its 15 minutes on
``main``, while another took 8. Adding any test file reshuffled most files
between shards, so an unrelated PR could tip a shard past its cap (#904).
``seconds`` holds each file's measured time. A file without a measurement
(new, or renamed since) costs its item count times the measured seconds
per item. A stale measurement only unbalances; it never drops a file.
Balance is checked in ``tests/test_shard_partition.py`` rather than
assumed.
"""

load = [0] * shards
cost = _costs(paths, seconds or {})
load = [0.0] * shards
owner: dict[str, int] = {}
for path, count in sorted(paths.items(), key=lambda item: (-item[1], item[0])):
for path in sorted(paths, key=lambda item: (-cost[item], item)):
target = min(range(shards), key=lambda index: (load[index], index))
owner[path] = target
load[target] += count
load[target] += cost[path]
return owner


def _costs(paths: Mapping[str, int], seconds: Mapping[str, float]) -> dict[str, float]:
known = {path: seconds[path] for path in sorted(paths) if path in seconds}
known_items = sum(paths[path] for path in known)
# ``fsum`` over a sorted order: every shard computes the same rate, bit
# for bit, whatever order its collection listed the files in.
rate = math.fsum(known.values()) / known_items if known_items else 1.0
return {path: known.get(path, paths[path] * rate) for path in paths}
4 changes: 2 additions & 2 deletions conftest.py
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@

import pytest # noqa: E402

from ci_sharding import shard_assignment # noqa: E402
from ci_sharding import load_seconds, shard_assignment # noqa: E402


@pytest.fixture(autouse=True)
Expand Down Expand Up @@ -90,7 +90,7 @@ def pytest_collection_modifyitems(config, items) -> None: # noqa: ANN001
counts: dict[str, int] = {}
for item in items:
counts[item.location[0]] = counts.get(item.location[0], 0) + 1
owner = shard_assignment(counts, shards)
owner = shard_assignment(counts, shards, load_seconds())
keep = [item for item in items if owner[item.location[0]] == index - 1]
dropped = [item for item in items if owner[item.location[0]] != index - 1]
if not keep:
Expand Down
94 changes: 91 additions & 3 deletions docs/application-comparison.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,8 +13,9 @@ agents-shipgate diff --application --workspace /path/to/repository \

No `init`, manifest, saved baseline or authored declaration is required. The
command writes no files in the subject repository and runs no application code.
Fetch the refs first; the command never fetches. `--scope` defaults to the
repository root; `--head` defaults to committed `HEAD`, excluding dirty edits.
Fetch the refs first; the command never fetches. Without `--scope` the scope
is derived from the change (below); `--scope .` selects the repository root.
`--head` defaults to committed `HEAD`, excluding dirty edits.
The base is the requested base ref's merge base with the selected head.

The result identifies each observed agent object and its added, removed or
Expand All @@ -33,6 +34,92 @@ A reviewer can inspect the new callable and decide whether that agent should
receive it. A body change is a request to review the implementation; it does not
establish widening, narrowing, business impact or runtime behavior.

## A scope derived from the change

Without `--scope`, the comparison does not read the whole repository, nor only
the directory the change touched. Each changed Python file is related to the
agent files — files discovery scores as OpenAI Agents SDK or Google ADK sources
that construct an agent class (under any name they import it as), subclass
one or copy one with capabilities of its own, and a module that changes an
agent's capabilities after construction (`support.tools.append(x)`) and
imports one of those, or that builds its agent through one's factory
(`root_agent = make("bot", [tool])` with `make` returning an `Agent`); a
module that only defines tools is not one — that
import it, or that it imports, up to six import hops away, on each side of the
change. A module under a directory discovery skips (`build/`, `fixtures/`) is
related like any other, never an agent file. Importing `pkg.impl` runs `pkg/__init__.py`
first, so a package's imports count, and so does a literal
`importlib.import_module("app.tools")`. A file named like a test is related
when an agent imports it; it is never an agent file, and a changed test that
imports an agent does not widen a scope. A changed Python file not related to
a module that builds an agent — no import path, one longer than six hops, only
a module that rewires an agent, a file too large to read — is named in
`scope_selection.limits` with why, and makes the result `partial`, even when a
compared scope holds it: its consumer may be elsewhere. So is a changed link
to a module or a directory, or a changed submodule, whose content is not read,
and so is an agent outside the compared scopes whose imports go past the hop
bound. A module that imports the change — directly, through a package's
re-export, an agent file's, or other modules, up to six levels — and imports
within six hops an agent builder that builds on request — one that
constructs an agent whose tools, handoffs, MCP servers or sub-agents are not a
fixed list of the module's own names (`tools=tools`, `list(REGISTRY.items)`,
`self.tools`, `**config`) or gives it such capabilities afterwards (a
rewire, or a copy such as `BASE.clone(tools=tools)`), or subclasses an agent
class — is named too,
however it builds its agent from them, when it builds, copies
(`clone(tools=...)`, `clone(update={"tools": ...})`) or rewires an agent
itself, calls the repository's code with arguments of its own, or sets the
repository's module state (`factory.TOOLS = [...]`,
`setattr(factory, ...)`), unless
one compared scope holds it, the change and that builder. A module that only
imports another module's agent, or only calls an entry point (`main()`), is
not one. A module that imports the change and imports onward past six hops
without reaching a builder is named as not established, and so is the search
itself when modules outside the compared scopes still import onward after six
levels. A rename's two paths are one file.

```text
scope: backend/app (derived: backend/app: backend/app/services/gemini_tools.py is
related to agent file backend/app/agents/support.py through …)
```

Each changed file is compared in the outermost package holding it and every
agent that reaches it — with what those agents import from the repository and
the path entry a namespace package is imported through (`src` for
`myapp.tools` in `src/myapp/`) — so the reader can follow each chain. A
change to `backend/app/services/gemini_tools.py` that
`backend/app/agents/support.py` imports is compared in `backend/app`. When only
the repository root holds them all — a library module its own agents and an
example app elsewhere both import — the change is compared where its nearest
agents are, the others are named in `outside` and in `scope_selection.limits`,
and the answer is `partial`, never a silent `compared`.

Independent applications of a monorepo are separate comparisons, never the
repository root; `--json` then holds each one under `comparisons`, with every
path spelled from the repository root, and `comparison_status` is `compared`
only when each one is. A Python module moved inside one package is one
comparison of that package; an application directory moved whole — its old
path gone from the head, its new one absent from the base — is one relocation,
compared old path to new path as `--base-scope`/`--scope` would, with every
scope inside it; a renamed
non-Python file joins nothing. A change that touches no supported agent — a
README, or a script no agent imports — is `not_established` with the reason
naming the files considered; nothing is compared, and that is an answer, not a
failure. An absolute import is found where the importer's own path would find
it, or at one path entry holding a package of that name (`libs/shared`); a
standard-library name is the standard library unless a module on the importer's
path shadows it.

Everything is read from the two commits' objects, never run. The reading is
bounded per side: six import hops, 2000 files read to relate the change, and
5000 files or 64 MB read to find the agent files. A bound reached is named in
`scope_selection.limits` and makes the result `partial` — so does a
no-agent answer whose import following stopped at the hop bound — and it never
falls back to the root. `--json` records `scope_selection`: `mode` (`derived`
or `explicit`), `scopes`, the `reason`, the changed files considered and each
scope's relations (`base_scope` when it moved, `outside` agents). An explicit
`--scope` always wins, and `--base-scope` needs it.

## Scopes, moves and incomplete inputs

For an application directory moved by the PR, select its old location separately:
Expand All @@ -54,7 +141,8 @@ exit 2. A removal describes the selected source path, not the entire repository.

`--json` emits `application_comparison_schema_version: "0.1"`, engine identity,
requested and compared refs/tree IDs, per-side scope/coverage, rows, source
correspondence, and a deterministic `comparison_id`. This is a separate advisory
correspondence, `scope_selection`, `comparisons` when a derived change spans
more than one application, and a deterministic `comparison_id`. This is a separate advisory
artifact from the existing host diff JSON and verifier receipt.

- `compared`: the selected supported source observations were compared. An
Expand Down
2 changes: 1 addition & 1 deletion docs/distribution-surfaces.md

Large diffs are not rendered by default.

Loading
Loading