Conversation
When the provider endpoint is a tracing-aware proxy — LiteLLM, or anything exporting its own OpenTelemetry spans — it traces the same request separately. The result is two unrelated traces of one call: this plugin's LLM Call and the proxy's generation, with nothing linking them. Every provider request now carries a W3C traceparent naming the current turn root, so a gateway that honours inbound trace context records its span in this plugin's trace. The header names the turn root rather than the matching LLM Call: before_provider_headers fires before before_provider_request, so the generation for that call does not exist yet and the root is the only span available. The gateway's span is therefore a sibling of LLM Call, not its child. Nothing is sent when no turn is in flight, which covers compaction, branch summaries and cache warming. A root the sampler dropped is skipped, matching publishParentContext. The mock provider now records request headers so the integration test can assert the header against the exported root span, and that the generation really shares the trace the header advertises.
Contributor
|
@kevinb361 thank you for raising this, we are looking into it! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The problem
When the provider endpoint is a tracing-aware proxy — LiteLLM, or anything that exports its own OpenTelemetry spans — it traces the same request separately. You end up with two unrelated traces of one call: this plugin's
LLM Calland the proxy's generation, with nothing linking them.On my self-hosted setup that was not a corner case. Measured over a 9.7h window: 1,156 traces litellm-only, 156 pi-only, 0 shared. The agent's view and the proxy's view of the same request never met, so following a turn down into the model call was impossible and any rollup across both sources double-counted tokens.
The cause is simply that pi sends no
traceparent.The change
Every provider request now carries a W3C
traceparentnaming the current turn root:A gateway that honours inbound trace context then produces one trace:
That tree is from a real run against LiteLLM, not a mock.
Why the turn root, and not the
LLM Callbefore_provider_headersfires beforebefore_provider_request—transformHeadersis awaited insideprepareRequest, before the provider builds its payload. I confirmed this at runtime as well as from the call graph; the two hooks landed 9ms apart in that order.So when the header hook runs, the generation for that call does not exist yet and the turn root is the only span available. The gateway's span becomes a sibling of
LLM Callrather than its child.Parenting exactly would mean opening the generation in the header hook instead. I did not do that here: the header hook also fires for requests that are not turn steps, and I could not convince myself that cache warming never fires it while a turn is in flight — that would create
LLM Callspans that never complete. It seemed wrong to trade a real regression risk for nesting depth. Happy to look at it as a follow-up if you think it is worth it.Guards
if (!state) return.TraceFlags.SAMPLEDcheck mirrorspublishParentContext, for the same reason — a child pointing at a trace whose root was never exported.Tests
test/helpers.tsdid not record request headers, so the mock provider now does, exposed asheaders()alongside the existingpayloads().The new integration test asserts every provider request carried a
traceparentequal to the exported root span's ids, and separately that the generation really shares the trace the header advertises — so it cannot pass on a span tree that never linked.I checked the test can fail: with the injection commented out it fails on
request 0 must carry the turn root as its traceparent. Full suite is 129/129 withtsc --noEmitclean.Two things worth your call
traceparentis inert to anything that does not read it, and an opt-in default means most people never discover the problem it solves.