Skip to content

feat: send a traceparent so a gateway joins the trace - #42

Open
kevinb361 wants to merge 2 commits into
langfuse:mainfrom
kevinb361:feat/traceparent-propagation
Open

kevinb361 wants to merge 2 commits into
langfuse:mainfrom
kevinb361:feat/traceparent-propagation

Conversation

@kevinb361

Copy link
Copy Markdown

The problem

When the provider endpoint is a tracing-aware proxy — LiteLLM, or anything that exports its own OpenTelemetry spans — it traces the same request separately. You end up with two unrelated traces of one call: this plugin's LLM Call and the proxy's generation, with nothing linking them.

On my self-hosted setup that was not a corner case. Measured over a 9.7h window: 1,156 traces litellm-only, 156 pi-only, 0 shared. The agent's view and the proxy's view of the same request never met, so following a turn down into the model call was impossible and any rollup across both sources double-counted tokens.

The cause is simply that pi sends no traceparent.

The change

Every provider request now carries a W3C traceparent naming the current turn root:

traceparent: 00-<turn trace id>-<turn root span id>-01

A gateway that honours inbound trace context then produces one trace:

Conversational Turn      (this plugin)
├── LLM Call             (this plugin)
└── litellm_request      (the proxy)
    └── raw_gen_ai_request

That tree is from a real run against LiteLLM, not a mock.

Why the turn root, and not the LLM Call

before_provider_headers fires before before_provider_request — transformHeaders is awaited inside prepareRequest, before the provider builds its payload. I confirmed this at runtime as well as from the call graph; the two hooks landed 9ms apart in that order.

So when the header hook runs, the generation for that call does not exist yet and the turn root is the only span available. The gateway's span becomes a sibling of LLM Call rather than its child.

Parenting exactly would mean opening the generation in the header hook instead. I did not do that here: the header hook also fires for requests that are not turn steps, and I could not convince myself that cache warming never fires it while a turn is in flight — that would create LLM Call spans that never complete. It seemed wrong to trade a real regression risk for nesting depth. Happy to look at it as a follow-up if you think it is worth it.

Guards

  • No header when no turn is in flight. Compaction, branch summaries and cache warming all reach this hook; there is nothing for them to attach to, so if (!state) return.
  • Never names a dropped root. The TraceFlags.SAMPLED check mirrors publishParentContext, for the same reason — a child pointing at a trace whose root was never exported.
  • Providers that do not understand the header ignore it.

Tests

test/helpers.ts did not record request headers, so the mock provider now does, exposed as headers() alongside the existing payloads().

The new integration test asserts every provider request carried a traceparent equal to the exported root span's ids, and separately that the generation really shares the trace the header advertises — so it cannot pass on a span tree that never linked.

I checked the test can fail: with the injection commented out it fails on request 0 must carry the turn root as its traceparent. Full suite is 129/129 with tsc --noEmit clean.

Two things worth your call

  1. This is on by default for every request. If you would rather it were opt-in behind a config field, say so and I will add one — I left it default-on because a traceparent is inert to anything that does not read it, and an opt-in default means most people never discover the problem it solves.
  2. It does not deduplicate. The plugin and the gateway now report the same tokens in one trace. They remain two independent measurements of one request, so summing across both still double-counts. I noted that in the README rather than trying to be clever about it.

When the provider endpoint is a tracing-aware proxy — LiteLLM, or anything
exporting its own OpenTelemetry spans — it traces the same request separately.
The result is two unrelated traces of one call: this plugin's LLM Call and the
proxy's generation, with nothing linking them.

Every provider request now carries a W3C traceparent naming the current turn
root, so a gateway that honours inbound trace context records its span in this
plugin's trace.

The header names the turn root rather than the matching LLM Call:
before_provider_headers fires before before_provider_request, so the generation
for that call does not exist yet and the root is the only span available. The
gateway's span is therefore a sibling of LLM Call, not its child.

Nothing is sent when no turn is in flight, which covers compaction, branch
summaries and cache warming. A root the sampler dropped is skipped, matching
publishParentContext.

The mock provider now records request headers so the integration test can
assert the header against the exported root span, and that the generation
really shares the trace the header advertises.
@CLAassistant

CLAassistant commented Sep 20, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@milanagm
milanagm requested a review from hassiebp September 23, 2026 15:55
@milanagm

Copy link
Copy Markdown
Contributor

@kevinb361 thank you for raising this, we are looking into it!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants