Skip to content

Read tools lists built by an expression in diff --application (#909) - #939

Merged
pengfei-threemoonslab merged 6 commits into
mainfrom
claude/issue-909-list-expressions
Oct 3, 2026
Merged

pengfei-threemoonslab merged 6 commits into
mainfrom
claude/issue-909-list-expressions

Conversation

@pengfei-threemoonslab

Copy link
Copy Markdown
Contributor

Why

On the 49-PR development corpus (#908), the most common reason diff --application produced no row was "uses a dynamic tools expression". An agent whose tools= was [*FINANCE_TOOLS, *([handoff] if wanted else [])], base + (extra or []), a filter, or another module's list had none of its bindings compared, even when every member was a plain tool the reader already understood. A reviewer got partial and nothing to review.

What changes

One resolver both readers share (src/agents_shipgate/inputs/list_expressions.py). The OpenAI Agents SDK reader (tools=, handoffs=) and the Google ADK reader (tools=, sub_agents=) read these expressions member by member:

  • list and tuple literals with * spreads;
  • a + b;
  • a if c else b and a or b;
  • a comprehension that keeps its elements, and filter(...);
  • list(...), tuple(...) and sorted(...);
  • a name bound once to any of these, in the builder, at module level, or in another module reached through a repository-local import.

Nothing is imported or run.

  • A part the reader cannot read is named where it is, on its agent. Examples: a call, a builder's parameter, self.tools, a name changed in place, a name bound twice. The agent stays incomplete, so the answer is never compared, and the members that were read are still compared. The Google ADK dynamic-tools limit now names its file, line and agent instead of covering the whole file.
  • Conditions are shown, never evaluated. A member held only under a condition carries bound_when on its row's side: one entry per way in, each the condition's whole source text. A change to the condition alone is a changed row whose why names both conditions and says the direction is not established. It is not_established when an unread part of the list could hold the tool another way.
  • application_comparison_schema_version 0.2 → 0.3. The text output prints bound only when ….
  • Fixed on the way:
    • an imported handoffs= list no longer reads as one handoff named after the list;
    • an ADK starred element the reader cannot read no longer leaves the agent's list marked complete.
  • OpenAI Agents SDK: resolve bounded literal tool-list concatenation in static binding evidence #584 (literal-list concatenation) is subsumed. Its recovery reason sdk_literal_tool_list_concatenation_unsupported is retired because that form is now read.

scan

  • Tool proof is unchanged. Google ADK's high extraction confidence is a proven-surface claim, and a list read through a name is checked for changes only in its own module, the modules its import passes through and the packages above it. So an ADK tools= or sub_agents= list that came through a name, or from another module, keeps the module at medium, and no scan decision moves. A literal spread into a literal ([a] + [b]) stays proven.
  • The binding graph now trusts an expression-built sub-agent list. It already trusted the SDK's handoff lists the same way since Resolve repository-local imported tools for Google ADK and OpenAI Agents SDK (#864) #879. Please check this choice.

Measured

Review

Five reviewers, each verifying with fixtures:

  • line-by-line review of the resolver;
  • removed behavior and cross-file callers;
  • fail-open hunting through diff --application;
  • tests, docs and conventions;
  • reuse and efficiency.

They found about 40 confirmed problems, 15 after merging overlaps; all are fixed and each has a regression test in tests/test_list_expressions.py. The ones that mattered:

  • Changes the change index missed:
    • an import inside a function changed in place;
    • mod.LIST.append through the module's own spelling;
    • a package __init__ that re-exports the list and changes it;
    • a change through an agent built with the list (helper.tools.append), in its own file or another;
    • defaults, decorators and class bodies scoped as Python evaluates them;
    • walrus and match-capture rebinding;
    • a wildcard import after the list;
    • a list passed after a * spread to a helper;
    • decorated helpers;
    • a user-defined replace or a non-agent .clone;
    • another module's own Agent class.
  • Condition text was cut at 160 characters before comparison. Conditions are now compared whole.
  • No bound on work: a list spread into itself many times grew exponentially (44 s, 1,000,000 members). Members are now deduplicated and memoized, with caps on members and visits.
  • A quadratic slowdown on large files (5.6 s → 349 s): callee verdicts are cached, and one cache is shared per load.
  • ADK:
    • sub_agents= lists read from expressions had let scan pass;
    • twins differing only in what they hold unconditionally were merged;
    • a partly read list blocked other agents' rows in the file.
  • SDK:
    • hosted tools lost their recovery evidence;
    • another module's unresolved member was bound by name to a local tool (a false row);
    • a handoff spelled like a local agent was taken for it.
  • A with block was treated as conditional, which regressed a list the old reader read.

Left as a follow-up: ADK's #865 reader for builder-local tools.append(...) uses its own rules beside this resolver. Folding it in changes documented behavior, so it is proposed separately.

Surface discipline

Closes #909. Closes #584.

🤖 Generated with Claude Code

pengfei-threemoonslab and others added 6 commits October 2, 2026 15:23
An agent whose tools=, handoffs= or sub_agents= is built by a spread,
concatenation, conditional, `or`, filter or another module's list is read
member by member through one resolver both readers share, instead of being
reported as a dynamic expression. A part that cannot be read is named where
it is, on its agent, which stays incomplete; a member held only under a
condition carries bound_when (application comparison schema 0.3). Subsumes
#584 and retires its recovery reason.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@pengfei-threemoonslab
pengfei-threemoonslab merged commit 16dc027 into main Oct 3, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant