Skip to content

Application diff reports 'compared, no changes' when an SDK agent built by 'return Agent(...)' gains a tool #876

Description

@pengfei-threemoonslab

Summary

diff --application can report compared / "No established binding/interface/implementation changes in the observed surface" when a real agent gains a tool. Two defects combine:

  1. The OpenAI Agents SDK reader does not observe an agent built as return Agent(...) inside a function, and it records no limit for it. The ADK reader does observe return LlmAgent(...); see Vesta and TensorFlow in Application diff: resolve tool lists passed into agent-building functions and constructors #874 and Application diff rows should name what a bound tool reaches: endpoint, action, credential, model-supplied arguments #872.
  2. Test files take part in agent discovery. A test double with resolvable tools is enough for the scope to count as established. Separately, a duplicated tool name in a test file refuses the whole comparison (exit 2).

This is a false complete answer, and a release blocker for the application route. The principle is the same as #873's "an unobserved agent is not a removal", applied the other way round: an unobserved agent is not "no change".

Real case: kkmiecik-coder/CRM#5 (merged 2026-08-31)

Merge base dae503cc671c…, head 0935d0931c58…, repository root, main a430e81a.

  • integrations/chat_bridge/bots_pro/narzedzia.py: NARZEDZIA_WYCENY gains wyslij_obraz, a new @function_tool that sends an image to the customer.
  • integrations/chat_bridge/bots_pro/agenci.py:65-92: zbuduj_agenta_wyceny() does return Agent(name="Wycena", …, tools=NARZEDZIA_WYCENY). Three more builders in the same file (:95, :112, :121) do the same.
  • Observed: compared, 0 rows, head.limits == []. The only head agent is agent at integrations/chat_bridge/tests/test_pro_tura.py:195, a test double. agenci.py is listed in sources, but no agent is established from it and nothing names the gap.

Minimal reproduction

# app/tools.py
from agents import function_tool
@function_tool
def quote(item: str) -> str: ...
@function_tool
def send_image(name: str) -> dict: ...
QUOTE_TOOLS = [quote]          # head: [quote, send_image]

# app/agents_def.py
from agents import Agent
from app.tools import QUOTE_TOOLS, quote
def build_quote_agent():
    return Agent(name="Quote", instructions="q", tools=QUOTE_TOOLS)

# app/tests/test_turn.py
from agents import Agent, function_tool
@function_tool
def fake_lookup(q: str) -> str: ...
agent = Agent(name="TestAgent", instructions="t", tools=[fake_lookup])

./shipgate diff --application --base HEAD~1 --head HEAD --scope app prints:

Application comparison: compared (dcd7e7becd37 → 0913ee4d15f6)
No established binding/interface/implementation changes in the observed surface.

JSON: head.agents == [agent (tests/test_turn.py)], rows == [], head.limits == [].

Without the test file, the same change gives not_established ("No supported application agents were established"). That is incomplete, but it is not a false no-change. With a module-level x = Agent(...) beside the builders, only x is observed, and Quote is still silently absent.

Refusal from test files

Each of these exited 2 on a whole-repository comparison because one file defines a tool name twice:

The message tells the user to "remove the duplicate definition" from their test file. One reader-level duplicate should become a named limit on that file, not a refusal of every other agent.

Acceptance

  • The SDK reader observes return Agent(...) in functions and methods, with the same name, location and binding evidence as an assigned agent. The CRM fixture above yields ADDED Quote → send_image (once Google ADK: resolve repository-local imported functions and module-qualified tool bindings #864 resolves the imported list, or else a named unresolved-tool limit on Quote), never compared with no rows.
  • Any agent construction the reader sees syntactically but does not establish produces a named limit, attributed to its file and line, so the comparison cannot be compared. Add a guard test: in a scope where every construction site is established or limited, compared requires zero unaccounted construction sites.
  • Test files do not establish the application. Files matched by the existing test-path conventions (tests/, test_*.py, *_test.py, conftest.py) are excluded from agent establishment by default, or reported separately and never counted toward compared. State the rule in docs/application-comparison.md.
  • A duplicate tool definition is a named limit on that file. It does not refuse the comparison, and it is ignored entirely when the file is excluded as a test.
  • Rerun on the pinned corpora (134 open, 50 modify-existing): no case is compared while a construction site in scope is unaccounted. Publish how many CRM-shaped silent cases existed.
  • CHANGELOG entry. Paired tests that fail before the fix.

Blocks the release step of the #868 plan. Related: #873 (unobserved ≠ removed), #864, #874, #875 (derived scope must also exclude tests).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions