Skip to content

fix(graph): prefer node tools before synthetic handoff routing - #59

Merged
andrewklatzke merged 1 commit into
mainfrom
aklatzke/AIC-3365/tighten-graph-handoffs
Sep 16, 2026
Merged

andrewklatzke merged 1 commit into
mainfrom
aklatzke/AIC-3365/tighten-graph-handoffs

Conversation

@andrewklatzke

Copy link
Copy Markdown
Contributor

Summary

When a graph node has multiple outgoing edges, route() augments the node's config with synthetic __handoff_* tools. In practice this crowded out the node's own tools — a node that reliably called its real tool standalone would call only a handoff tool inside a graph, so the response carried routing information instead of the tool's data.

Three things in the routing augmentation pushed the model that way, all fixed here:

  • The appended routing directive led with "Select exactly one transfer tool to route to the next agent", making the transfer read as the task and the node's real work optional. It now asks the model to complete its task with its available tools first, and only then call exactly one transfer tool.
  • Handoff tool descriptions fell back to the target node's instructions verbatim, so __handoff_<target> advertised itself as the tool that does the target's work. Descriptions now always lead with Transfer control to <key>., with any handoff.description or target-instructions text appended as trailing detail.
  • The handoff handler returned Transferring to <key>, which reads as though control had already left, so the model wrapped up instead of continuing. Selecting an edge only records the choice — execution continues until the model produces its final text — so the result now says the handoff is recorded and asks the model to finish its own work.

Behavior is unchanged for nodes with zero or one outgoing edge, which never enter this path. Framework-native runners build their own transfer_to_* tools and are untouched.

Test plan

  • Existing client suite passes (30 tests, includes src/__tests__/graph.test.ts)
  • Verified against a real multi-node graph whose root node was skipping its tool: the node's own tool now fires and the handoff still routes correctly
  • Reviewer sanity check on the routing copy — this is model-steering, so wording matters

Note: the matching change for the Python SDK is launchdarkly/python-ai-sdk#90

Made with Cursor

Handoff tools were crowding out real node tools because routing copy led with "select a transfer tool", descriptions reused the target agent's instructions, and the tool result implied the turn was already over.

Co-authored-by: Cursor <cursoragent@cursor.com>
andrewklatzke added a commit to launchdarkly/python-ai-sdk that referenced this pull request Sep 16, 2026
## Summary

When a graph node has multiple outgoing edges, `route()` augments the
node's config with synthetic `__handoff_*` tools. In practice this
crowded out the node's own tools — a node that reliably called its real
tool standalone would call only a handoff tool inside a graph, so the
response carried routing information instead of the tool's data.

Three things in the routing augmentation pushed the model that way, all
fixed here:

- The appended routing directive led with "Select exactly one transfer
tool to route to the next agent", making the transfer read as the task
and the node's real work optional. It now asks the model to complete its
task with its available tools first, and only then call exactly one
transfer tool.
- Handoff tool descriptions fell back to the *target* node's
instructions verbatim, so `__handoff_<target>` advertised itself as the
tool that does the target's work. Descriptions now always lead with
`Transfer control to <key>.`, with any `handoff.description` or
target-instructions text appended as trailing detail.
- The handoff handler returned `Transferring to <key>`, which reads as
though control had already left, so the model wrapped up instead of
continuing. Selecting an edge only records the choice — execution
continues until the model produces its final text — so the result now
says the handoff is recorded and asks the model to finish its own work.

Behavior is unchanged for nodes with zero or one outgoing edge, which
never enter this path. Framework-native runners build their own
`transfer_to_*` tools and are untouched.

## Test plan

- [x] Existing graph suite passes (27 tests,
`packages/client/tests/test_graph.py`)
- [x] Verified against a real multi-node graph whose root node was
skipping its tool: the node's own tool now fires and the handoff still
routes correctly
- [ ] Reviewer sanity check on the routing copy — this is
model-steering, so wording matters

Note: the matching change for the JS SDK is
launchdarkly/js-ai-sdk#59

Co-authored-by: Cursor <cursoragent@cursor.com>
@andrewklatzke
andrewklatzke merged commit 0578819 into main Sep 16, 2026
9 checks passed
@andrewklatzke
andrewklatzke deleted the aklatzke/AIC-3365/tighten-graph-handoffs branch September 16, 2026 17:45
@github-actions github-actions Bot mentioned this pull request Sep 16, 2026
andrewklatzke pushed a commit that referenced this pull request Sep 22, 2026
🤖 I have created a release *beep* *boop*
---


<details><summary>@launchdarkly/ai-server: 0.3.0</summary>

##
[0.3.0](https://github.com/launchdarkly/js-ai-sdk/compare/@launchdarkly/ai-server-0.2.0...@launchdarkly/ai-server-0.3.0)
(2026-09-18)


### Features

* **AIC-3106:** add multimodal history support to graph().invoke()
([#18](#18))
([9737530](9737530))
* **client:** stamp modelKey and modelVersion from _ldMeta on tracking
events
([758fe7c](758fe7c))
* **client:** stamp modelKey and modelVersion from _ldMeta on tracking
events ([#66](#66))
([89e4f0e](89e4f0e))


### Bug Fixes

* **client:** harden model stamps, share node trackData builder, keep
judge results from inheriting parent model identity
([d374962](d374962))
* **client:** reject blank-string modelVersion and non-string modelKey
in model stamps
([f80c016](f80c016))
* extract LangChain content-block text and apply model parameters after
eval ([#54](#54))
([e34e779](e34e779))
* **graph:** prefer node tools before synthetic handoff routing
([#59](#59))
([0578819](0578819))
* **telemetry:** a tool that returned nothing did not return null
([#21](#21))
([92cfcee](92cfcee))
</details>

<details><summary>@launchdarkly/ai-claude-agents: 0.3.0</summary>

##
[0.3.0](https://github.com/launchdarkly/js-ai-sdk/compare/@launchdarkly/ai-claude-agents-0.2.0...@launchdarkly/ai-claude-agents-0.3.0)
(2026-09-18)


### Features

* **AIC-3106:** add multimodal history support to graph().invoke()
([#18](#18))
([9737530](9737530))
* **client:** stamp modelKey and modelVersion from _ldMeta on tracking
events
([758fe7c](758fe7c))
* **client:** stamp modelKey and modelVersion from _ldMeta on tracking
events ([#66](#66))
([89e4f0e](89e4f0e))


### Bug Fixes

* **client:** harden model stamps, share node trackData builder, keep
judge results from inheriting parent model identity
([d374962](d374962))
</details>

<details><summary>@launchdarkly/ai-claude-messages: 0.3.0</summary>

##
[0.3.0](https://github.com/launchdarkly/js-ai-sdk/compare/@launchdarkly/ai-claude-messages-0.2.0...@launchdarkly/ai-claude-messages-0.3.0)
(2026-09-18)


### Features

* **AIC-3106:** add multimodal history support to graph().invoke()
([#18](#18))
([9737530](9737530))
</details>

<details><summary>@launchdarkly/ai-openai-agents: 0.3.0</summary>

##
[0.3.0](https://github.com/launchdarkly/js-ai-sdk/compare/@launchdarkly/ai-openai-agents-0.2.0...@launchdarkly/ai-openai-agents-0.3.0)
(2026-09-18)


### Features

* **AIC-3106:** add multimodal history support to graph().invoke()
([#18](#18))
([9737530](9737530))
* **client:** stamp modelKey and modelVersion from _ldMeta on tracking
events
([758fe7c](758fe7c))
* **client:** stamp modelKey and modelVersion from _ldMeta on tracking
events ([#66](#66))
([89e4f0e](89e4f0e))


### Bug Fixes

* **client:** harden model stamps, share node trackData builder, keep
judge results from inheriting parent model identity
([d374962](d374962))
</details>

<details><summary>@launchdarkly/ai-openai-messages: 0.3.0</summary>

##
[0.3.0](https://github.com/launchdarkly/js-ai-sdk/compare/@launchdarkly/ai-openai-messages-0.2.0...@launchdarkly/ai-openai-messages-0.3.0)
(2026-09-18)


### Features

* **AIC-3106:** add multimodal history support to graph().invoke()
([#18](#18))
([9737530](9737530))
</details>

<details><summary>@launchdarkly/ai-langchain-agents: 0.3.0</summary>

##
[0.3.0](https://github.com/launchdarkly/js-ai-sdk/compare/@launchdarkly/ai-langchain-agents-0.2.0...@launchdarkly/ai-langchain-agents-0.3.0)
(2026-09-18)


### Features

* **AIC-3106:** add multimodal history support to graph().invoke()
([#18](#18))
([9737530](9737530))
* **client:** stamp modelKey and modelVersion from _ldMeta on tracking
events
([758fe7c](758fe7c))
* **client:** stamp modelKey and modelVersion from _ldMeta on tracking
events ([#66](#66))
([89e4f0e](89e4f0e))


### Bug Fixes

* **AIC-3382:** support Bedrock configs in LangChain handlers
([#72](#72))
([a1487cf](a1487cf))
* **client:** harden model stamps, share node trackData builder, keep
judge results from inheriting parent model identity
([d374962](d374962))
* extract LangChain content-block text and apply model parameters after
eval ([#54](#54))
([e34e779](e34e779))
</details>

<details><summary>@launchdarkly/ai-langchain-messages: 0.3.0</summary>

##
[0.3.0](https://github.com/launchdarkly/js-ai-sdk/compare/@launchdarkly/ai-langchain-messages-0.2.0...@launchdarkly/ai-langchain-messages-0.3.0)
(2026-09-18)


### Features

* **AIC-3106:** add multimodal history support to graph().invoke()
([#18](#18))
([9737530](9737530))


### Bug Fixes

* **AIC-3382:** support Bedrock configs in LangChain handlers
([#72](#72))
([a1487cf](a1487cf))
* extract LangChain content-block text and apply model parameters after
eval ([#54](#54))
([e34e779](e34e779))
</details>

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).

<!-- CURSOR_SUMMARY -->
---

> [!NOTE]
> **Overview**
> **Release Please cut** that bumps `@launchdarkly/ai-server` and the
Claude, OpenAI, and LangChain agent/message packages from **0.2.0 →
0.3.0**, updates `.release-please-manifest.json`, embedded
`LD_AI_PACKAGE_VERSION` constants, and adds **0.3.0** changelog sections
(no runtime code in this diff).
> 
> The **0.3.0** notes capture already-merged work: **multimodal
conversation history** on `graph().invoke()`, **modelKey/modelVersion**
on tracking events from `_ldMeta` (with validation and judge events not
inheriting parent model identity), plus fixes for **graph tool
routing**, **LangChain** text extraction / post-eval model params and
**Bedrock** configs, and **telemetry** for empty tool results.
> 
> <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit
6130d2c. Bugbot is set up for automated
code reviews on this repo. Configure
[here](https://www.cursor.com/dashboard/bugbot).</sup>
<!-- /CURSOR_SUMMARY -->

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
jeffdupont added a commit that referenced this pull request Sep 28, 2026
## Summary

Adds `graph().stream()` — an async-generator counterpart to
`graph().invoke()` that yields node boundaries while the router keeps
ownership of handoffs and graph-level telemetry.

- New public `GraphStreamEvent` union: `node_start`, `chunk` (tagged
with `nodeKey`), `node_done`, `handoff`, and a final `done`. Its `chunk`
and `done` variants stay structurally assignable to `StreamEvent`, so
renderers written against `config().stream()` type-check against graph
events unchanged. Handoff events use `sourceKey` / `targetKey`, the same
names as `$ld:ai:graph:handoff_*`.
- `bindSpanContext` (a sibling of `bindConversationId`, next to it in
`conversation.ts`) re-enters the graph span's context on **every**
`next()`. A generator body suspends at each `yield`, so wrapping it once
is not enough — without this, a streamed two-node graph emitted node
spans with no parent across three separate trace ids, while the same
graph through `invoke()` produced one correctly nested trace.
- Both the OTel parent and the conversation id are captured at
`stream()` call time rather than on first `next()`, so a caller can hand
the generator to a renderer and have it iterated later without
`launchdarkly.graph` detaching into its own trace.
- The graph span is `launchdarkly.graph` with `launchdarkly.graph.key`,
matching the prefix rename on main (#16). It is opened with `startSpan`
(parent captured at call time) rather than `startActiveSpan`, because
the generator suspends. `invoke()` drains this same span; it does not
open a second one.
- The graph-level judge runs with the graph span re-entered, so its
spans share the graph trace the way `invoke()`'s already did.
- Consumer abandonment (`break` mid-stream) ends the graph span with
`launchdarkly.stream.abandoned` and tracks neither success nor failure —
abandonment is neither, per the convention documented above
`endSpanOnce`.
- `graph().invoke()`, `route()`, and `runNode()` drain this walk.
Handoffs, judges, and graph telemetry live in one place; the blocking
callers read it to completion.
- Multi-edge routing is built once, in `buildHandoffRouting`, and only
`streamRoute()` calls it. `route()` drains that generator, so the two
entry points cannot drift. An earlier streaming copy had been written
against a pre-#59 `route()` and silently reverted that PR's three prompt
fixes.
- Handlers without `.stream` fall back to a single chunk per node,
matching `config().stream()`.
- The OpenAI Agents handler yields on `output_text_delta`.
`@openai/agents` 0.11.6 rewrites the Responses API
`response.output_text.delta` to that name before the handler sees it;
matching the wire name produced a successful run with no model text.

Also corrects `packages/client/README.md`'s `ProviderGraphResponse` row,
which listed a `trackData` field the type does not have.

`examples/graph-streaming.ts` is wired through `main.ts`, and
`integration-config.json` lists `graph-streaming` with the same flag
keys as `graph`.

## Test plan

- [x] `packages/client/src/__tests__/graph.test.ts`: 59 tests — event
ordering, per-node `nodeKey` tagging, aggregate usage, handoff placement
between `node_done` and the next `node_start`, disabled-graph and
missing-handler errors, graph telemetry, per-node
`$ld:ai:generation:success` carrying `graphKey`, blocking-handler
fallback, and judge results on `done`
- [x] OTel parenting pinned from both sides: handler spans nest under
`launchdarkly.graph` on the stream path **and** the invoke path, every
finished span sharing one `traceId`
- [x] Deferred iteration keeps `launchdarkly.graph` under the caller's
span; graph-judge spans nest under `launchdarkly.graph`
- [x] Abandonment: graph span carries `launchdarkly.stream.abandoned`,
and no `$ld:ai:graph:invocation_success` is tracked
- [x] Multi-edge branching fixture: `handoff_success` from the route
branch, `handoff_failure` after a choice was captured, judges receiving
the un-augmented config, and handoff descriptions / tool reply / routing
suffix matching the blocking path
- [x] Each span-parenting and routing fix was confirmed **red before**
its fix, then confirmed load-bearing after by reverting each in turn —
exactly one test fails each time, and it is the intended one
- [x] OpenAI agents streaming matches the event `@openai/agents` 0.11.6
yields (`output_text_delta`). `yarn test` re-run on the merge commit
`11bd299` exited 0: every workspace package, including 95 tests in
`@launchdarkly/ai-openai-agents` and 421 in `@launchdarkly/ai-server`
- [x] Live `graph-streaming` success key `travel-agent-flow` ("I was
double charged for my flight"), re-run on the merge commit `11bd299`:
exit 0, model text on stdout, one `[conversation]` line,
`[node_start]`/`[node_done]` for `travel-agent-orchestrator`, Usage
input 366 / output 49 / total 415 matching that node. No handoff, which
is valid for a single-node run. No JSON file. stderr had no
`RuntimeError` or `aclose`
- [x] Live `graph-streaming` failure key `travel-agent-flow-wrong-key`:
exit 1, one `Error: Agent graph "travel-agent-flow-wrong-key" is
disabled`, no unhandled rejection, no JSON file
- [x] After merging main, re-run on `11bd299`: `yarn test` includes
`packages/client` — 14 files, 421 passed (59 of them in `graph.test.ts`,
covering event order, `launchdarkly.graph` parenting on both paths,
deferred iteration, abandonment, and multi-edge routing). The three
tests added by the merge are main's judge-config skip coverage (#17)

Note: `packages/client/tsconfig.json` uses `"include": ["src/*.ts"]`, so
`tsc --noEmit` does not typecheck anything under `src/__tests__/`. Left
as-is — out of scope here.

## Known behaviour carried over, not introduced

`handoff_success` double-emits on a multi-edge hop (once from the route
branch, once from the next node's `opts.from`). That was already true of
`route` / `runNode`. The single walk keeps it, and the test pins the
count at 2.

`judgeResults` is the same on both entry points now. `invoke()` reads
the stream's `done` event, which omits the field when judges return
`{}`. The earlier split — stream normalizing `{}` to `undefined` while
`invoke()` passed `{}` through — is gone.

`graph().invoke()` drains `handler.stream` when the handler has one, so
graph nodes on **both** entry points follow TESTING.md §1.9: streaming
ignores `outputFormat`. `config().invoke()` still uses the blocking
handler. Python `graph().invoke()` still does too (see
[python-ai-sdk#104](launchdarkly/python-ai-sdk#104));
only its `stream()` path drops the schema. The next node receives the
previous response as text (`[Previous agent response]\n...`) on either
path, not as a parsed object. What `outputFormat` changes is whether the
provider was constrained to that schema while producing the text.

## Relationship to #40 and #39

#40 implements this same ticket and predates this PR by four weeks — it
was open, unreviewed, and nobody caught the overlap before this was
built. It is being closed as overly complex, and two specific things
drove that:

- Its `done` event mirrors `invoke()`'s return **including `path` and
`nodes`**, which is why it needed #39 (AIC-3211) stacked underneath it.
`TESTING.md` §3.11 has required since the spec's initial commit that the
graph return "contain only `response`, `usage`, and optionally
`judgeResults` — no `path` or `nodes` fields". This PR's `done`
conforms; #40's required changing that rule and shipping a second ticket
first.
- It put `streamRoute` on the public `GraphDefinition`. Here it stays
internal to `buildGraph`, so `resolveGraph()`'s contract is unchanged
and no constructor signature moves when Python mirrors this.

Credit where it is due: #40 factored the handoff-tool setup into a
shared `prepareRoute()` from the start, and this PR originally forked it
— which silently reverted #59's three routing-prompt fixes on the
streaming path. `buildHandoffRouting` is the same idea, arrived at the
hard way.

## On `streamNode` / `streamRoute`

`runNode` and `route` drain `streamNode` and `streamRoute`.
`graph().invoke()` drains `graph().stream()`. There is no second copy of
the router.

The two generators are still split: one outgoing edge delegates to the
node walk, and multiple edges build the handoff tools. That is the same
split `route` had before it became a drain. Collapsing them would mix
the single-edge path with the synthetic-tool path, which is the opposite
of what §3.15a's routing-parity rule is pinning.

## Integration coverage — read before approving

`examples/graph-streaming.ts` is wired through `main.ts`, and
`integration-config.json` lists `graph-streaming` with the same flag
keys as `graph`. Both keys were run live.

The first success run exited 0 and reported tokens, but stdout had no
model text. The OpenAI agents handler was still matching
`response.output_text.delta`. `@openai/agents` 0.11.6 yields
`output_text_delta` instead. After that check was updated, the same
success key printed the model reply. The failure key was a clean
disabled-graph error. Details are in the test plan above.

Spec: `TESTING.md` §3.15a and Appendix A.13. The companion spec PR is
merged: launchdarkly/ai-sdks-monorepo#20. Python
implementation: launchdarkly/python-ai-sdk#104.

Jira: AIC-3210
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants