Repository navigation
Conversation
tychtjan
marked this pull request as ready for review
October 9, 2026 12:52
With GOODDATA_EVAL_JOIN_GENAI_TRACE on, every chat request of an eval item carries W3C baggage naming the item's trace and root span (langfuse_trace_id, langfuse_experiment_item_root_observation_id) and its experiment (langfuse_experiment_id/_name/_dataset_id/_item_id). A gen-ai that accepts caller baggage opens its turn under that root and reports the trace id on response_started. Both ids are derived from the conversation id, so the sending thread and the deferred scorer agree without passing state between threads. For a joined conversation, the deferred task reads the trace by id (/v2/observations?traceId=), waiting until every joined turn is ingested, then exports gd-eval's root into that trace in the sdk-experiment environment, spanning the whole run, and writes each score once on it. Langfuse then reports the item's real cost and latency, and the score rows are halved. Conversations that did not join, and every run with the switch off, keep today's path: poll by session, root in its own trace, scores on both. The item scope is set before the K runs in all twelve agentic kinds. anomaly_detection and what_if now pass conversation_id to observe, so an unlinked root there now carries session.id. jira: trivial risk: low
ChatClient.send_message passes each attempt's lines through _tap (identity by default) before parsing. A subclass that needs more of the stream than ChatResult keeps, such as tavern's obfuscation client reading an error's reason, overrides _tap instead of copying send_message, and so keeps the turn deadline and the eval-trace join. jira: trivial risk: low
Langfuse Cloud rate-limits public API calls per organization in fixed one-minute windows, and score writes share the bucket with observation and score reads (30/min on Hobby, 1,000/min on Pro). A 429 carries a Retry-After counting down to the window's reset, often 30-60 s. create_score capped every retry wait at 5 s, so both retries landed in the same exhausted window and the score was dropped. On a Hobby project a 60-request burst lost 43 scores this way. A 429 now waits up to the 60 s window; 5xx retries keep the 5 s cap. jira: trivial risk: low
Each eval item wrote every score with its own POST /api/public/scores, twice when the item was not joined to gen-ai's trace. Those requests share Langfuse's public-API rate-limit bucket with the observation and score reads, so a nightly run spent most of its budget on score writes. submit_trace_scoring now collects every score_safe call of the deferred scoring step and sends them as POST /api/public/scores arrays of up to 100. Score ids are fixed before the first attempt, so a retried batch cannot write a score twice. Langfuse validates array entries one by one and answers 207 when some are rejected; that is logged, not retried. Callers outside submit_trace_scoring (the single-shot sink, any client without create_scores) still write each score as it comes. Under a sustained 429 one item now waits at most 2 x 60 s, where per-score writes with the 60 s Retry-After cap could wait 36 minutes, past tavern's 420 s test timeout. A cancelled drain sends what was collected before the interrupt once, without waiting out a retry. jira: trivial risk: low
Langfuse groups API traffic by User-Agent, so every gooddata-eval call showed up as python-httpx/0.28.1, indistinguishable from other callers. Its exported root spans carried only service.name, so a trace did not say which build or CI run produced it. Every client from make_http_client now sends gooddata-eval/<version> (<ci|local>; run=<GITHUB_RUN_ID>; sha=<12 chars of GITHUB_SHA>), with "local" for any value unset or unsafe in a header. The OTLP resource adds service.version and, when set, github.run_id, github.workflow, github.ref_name and github.event_name. jira: trivial risk: low
tychtjan
force-pushed
the
jtyc/eval-single-trace
branch
from
October 9, 2026 12:56
096f2ec to
fa9a2c4
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Each agentic eval item becomes one Langfuse trace: gd-eval's experiment root with gen-ai's turns nested under it, and every score written once. Behind
GOODDATA_EVAL_JOIN_GENAI_TRACE, off by default; with the switch off, requests and traces are unchanged.How it works:
evaluate_agentic_*kind sets an item scope (experiment base name, dataset id, item id, whether runs are suffixed).langfuse_trace_id,langfuse_experiment_item_root_observation_idandlangfuse_experiment_{id,name,dataset_id,item_id}. Trace and root ids are derived from the conversation id, so the sending thread and the deferred scorer agree without sharing state.traceIdonresponse_started. A conversation counts as joined only when that id matches./v2/observations?traceId=), waiting until every joined turn is ingested. It then exports the root into that trace in thesdk-experimentenvironment, spanning the whole run, and writes each score once on it.One item with the switch on
sequenceDiagram autonumber participant I as gd-eval item participant G as gen-ai participant L as Langfuse participant T as gd-eval scorer I->>G: chat turn + join baggage (T, S) G->>L: send_message span under T/S G-->>I: response_started traceId=T Note over I: conversation marked joined T->>L: read trace T (all turns in) T->>L: export root S, last T->>L: one scores array, on ST and S are derived from the conversation id; the baggage also names the experiment and dataset item. A conversation whose
traceIddoes not match T (switch off, or a gen-ai without the join) keeps today's path: poll by session, root in its own trace, scores to both traces.Also in this PR:
ChatClient._tap: a hook over each attempt's SSE lines, so tavern's obfuscation client no longer copiessend_message.create_scoreon a 429 now waits for the Retry-After the server names, up to the 60 s Langfuse rate-limit window, instead of capping at 5 s and dropping the score. Langfuse Cloud limits score writes and reads together, per organization, in fixed one-minute windows.The deferred scoring step sends an item's scores as
POST /api/public/scoresarrays (up to 100), on both paths. Ids are fixed before the first attempt, so a retry cannot duplicate; Langfuse validates entries one by one and answers 207 on a partial result (logged, not retried). Under a sustained 429 an item waits at most 2 × 60 s, against 36 min for per-score writes with the longer wait. A cancelled drain sends what it collected once, without retrying.gooddata-eval names itself to Langfuse:
gooddata-eval/<version> (<ci|local>; run=<GITHUB_RUN_ID>; sha=<12 chars>)on every request, andservice.versionplus thegithub.*run context on the exported spans' resource.What changes for consumers when the switch is on:
sdk-experimentand trace namegd-eval: ….anomaly_detectionandwhat_ifnow passconversation_idtoobserve, so an unlinked root there carriessession.id.Test plan
packages/gooddata-eval: 1655 passed, 1 skipped;make format lint,ty checkcleanagent_metric_skill+agent_guardrails, gen-ai from the gdc-nas branch):tycleanRollout
gen-ai (gdc-nas) deploys first. Turning the switch on against a gen-ai without it would tag gen-ai's own traces with experiment ids under a root that never arrives. Moving eval traffic to
sdk-experimentdoes not reach InternalBI, which reads only the production project's export (confirmation pending).Risk
low: everything new is behind a switch that defaults to off. The 429 change only lengthens a retry wait.
Summary by CodeRabbit