Skip to content

feat(gooddata-eval): put gen-ai turns under the eval item root trace - #1863

Open
tychtjan wants to merge 5 commits into
masterfrom
jtyc/eval-single-trace
Open

tychtjan wants to merge 5 commits into
masterfrom
jtyc/eval-single-trace

Conversation

@tychtjan

@tychtjan tychtjan commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Each agentic eval item becomes one Langfuse trace: gd-eval's experiment root with gen-ai's turns nested under it, and every score written once. Behind GOODDATA_EVAL_JOIN_GENAI_TRACE, off by default; with the switch off, requests and traces are unchanged.

How it works:

  • Before the K runs, each evaluate_agentic_* kind sets an item scope (experiment base name, dataset id, item id, whether runs are suffixed).
  • Every chat request of the item carries W3C baggage: langfuse_trace_id, langfuse_experiment_item_root_observation_id and langfuse_experiment_{id,name,dataset_id,item_id}. Trace and root ids are derived from the conversation id, so the sending thread and the deferred scorer agree without sharing state.
  • A gen-ai that accepts caller baggage (gdc-nas PR, staging only) opens its turn under that root and reports traceId on response_started. A conversation counts as joined only when that id matches.
  • For a joined conversation, the deferred task reads the trace by id (/v2/observations?traceId=), waiting until every joined turn is ingested. It then exports the root into that trace in the sdk-experiment environment, spanning the whole run, and writes each score once on it.
  • Conversations that did not join keep today's path: poll by session, root in its own trace, scores written to both.

One item with the switch on

sequenceDiagram
    autonumber
    participant I as gd-eval item
    participant G as gen-ai
    participant L as Langfuse
    participant T as gd-eval scorer
    I->>G: chat turn + join baggage (T, S)
    G->>L: send_message span under T/S
    G-->>I: response_started traceId=T
    Note over I: conversation marked joined
    T->>L: read trace T (all turns in)
    T->>L: export root S, last
    T->>L: one scores array, on S
Loading

T and S are derived from the conversation id; the baggage also names the experiment and dataset item. A conversation whose traceId does not match T (switch off, or a gen-ai without the join) keeps today's path: poll by session, root in its own trace, scores to both traces.

Also in this PR:

  • ChatClient._tap: a hook over each attempt's SSE lines, so tavern's obfuscation client no longer copies send_message.

  • create_score on a 429 now waits for the Retry-After the server names, up to the 60 s Langfuse rate-limit window, instead of capping at 5 s and dropping the score. Langfuse Cloud limits score writes and reads together, per organization, in fixed one-minute windows.

  • The deferred scoring step sends an item's scores as POST /api/public/scores arrays (up to 100), on both paths. Ids are fixed before the first attempt, so a retry cannot duplicate; Langfuse validates entries one by one and answers 207 on a partial result (logged, not retried). Under a sustained 429 an item waits at most 2 × 60 s, against 36 min for per-score writes with the longer wait. A cancelled drain sends what it collected once, without retrying.

  • gooddata-eval names itself to Langfuse: gooddata-eval/<version> (<ci|local>; run=<GITHUB_RUN_ID>; sha=<12 chars>) on every request, and service.version plus the github.* run context on the exported spans' resource.

What changes for consumers when the switch is on:

  • The experiment item gets its real cost (the sum of all generations) and a latency covering the whole run.
  • Score rows halve.
  • Eval traces move to environment sdk-experiment and trace name gd-eval: ….
  • With the switch off, anomaly_detection and what_if now pass conversation_id to observe, so an unlinked root there carries session.id.

Test plan

  • packages/gooddata-eval: 1655 passed, 1 skipped; make format lint, ty check clean
  • New tests: id derivation, run index by send order, join detection, joined read by trace id, partial-trace wait, root export and single score write, switch off sends no new keys, every kind sends run 0 first, 429 wait
  • Local side-by-side on the docker stack (20 items × K=2, agent_metric_skill + agent_guardrails, gen-ai from the gdc-nas branch):
    • switch on: 0/40 unlinked, 1 trace per item, 312 score writes
    • switch off: 2.48 traces per item, 624 score writes
    • same pass/fail per item; the experiment item shows cost only with the switch on
  • Score batching and caller identity: 1672 passed, 1 skipped; format, lint, ty clean
  • Not yet re-run locally after the 429 fix and the batching
  • Staging dispatch on one combo after the gen-ai change is deployed

Rollout

gen-ai (gdc-nas) deploys first. Turning the switch on against a gen-ai without it would tag gen-ai's own traces with experiment ids under a root that never arrives. Moving eval traffic to sdk-experiment does not reach InternalBI, which reads only the production project's export (confirmation pending).

Risk

low: everything new is behind a switch that defaults to off. The 429 change only lengthens a retry wait.

Summary by CodeRabbit

  • New Features
    • Evaluation traces can be linked to their dataset items and conversations in Langfuse, with run-specific trace names and context.
    • Evaluation summaries now include conversation activity when a separate root span is unavailable.
    • Langfuse exports include package and CI run details.
  • Improvements
    • Scores are submitted in batches where supported, with clearer handling of partial rejections.
    • Retry delays now respect rate limits while applying shorter caps to other retryable errors.

@coderabbitai

coderabbitai Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Essentials
  • Run ID: e08b5e61-dbea-4955-aae0-8d27bbe792e5

📥 Commits

Reviewing files that changed from the base of the PR and between 8168136 and fa9a2c4.


📒 Files selected for processing (34)
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/_langfuse.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/_trace_linker.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/alert_skill.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/anomaly_detection.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/conversation.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/dashboard_skill.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/general_question.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/guardrail.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/kda_skill.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/metric_skill.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/report_skill.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/search_tool.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/visualization.py
  • packages/gooddata-eval/src/gooddata_eval/core/agentic/what_if.py
  • packages/gooddata-eval/src/gooddata_eval/core/chat/sse_client.py
  • packages/gooddata-eval/src/gooddata_eval/core/langfuse/_env.py
  • packages/gooddata-eval/src/gooddata_eval/core/langfuse/client.py
  • packages/gooddata-eval/src/gooddata_eval/core/langfuse/experiment.py
  • packages/gooddata-eval/src/gooddata_eval/core/langfuse/item_scope.py
  • packages/gooddata-eval/src/gooddata_eval/core/langfuse/observations.py
  • packages/gooddata-eval/src/gooddata_eval/core/langfuse/otlp.py
  • packages/gooddata-eval/src/gooddata_eval/core/models.py
  • packages/gooddata-eval/tests/conftest.py
  • packages/gooddata-eval/tests/test_agentic_anomaly_detection.py
  • packages/gooddata-eval/tests/test_agentic_join_trace.py
  • packages/gooddata-eval/tests/test_agentic_runner.py
  • packages/gooddata-eval/tests/test_agentic_what_if.py
  • packages/gooddata-eval/tests/test_langfuse_client.py
  • packages/gooddata-eval/tests/test_langfuse_e2e_fake_server.py
  • packages/gooddata-eval/tests/test_langfuse_env.py
  • packages/gooddata-eval/tests/test_langfuse_item_scope.py
  • packages/gooddata-eval/tests/test_langfuse_observations.py
  • packages/gooddata-eval/tests/test_langfuse_otlp.py
  • packages/gooddata-eval/tests/test_sse_client.py

Included review availability: This review used your included allowance. 3 included reviews remain after this review. Your included PR review attempts over the past 7 days set your current allowance at 5 reviews per hour.



📝 Walkthrough

Walkthrough

Agentic evaluators now create item-scoped Langfuse traces and propagate trace context through chat requests. Joined conversations use trace-ID lookups and export their experiment root into the joined trace. Score writes can be batched, and Langfuse request and OTLP resource metadata now include environment details.

Changes

Agentic Langfuse trace linking

Layer / File(s) Summary
Item scopes and evaluator integration
packages/gooddata-eval/src/gooddata_eval/core/langfuse/item_scope.py, packages/gooddata-eval/src/gooddata_eval/core/agentic/_trace_linker.py, packages/gooddata-eval/src/gooddata_eval/core/agentic/*, packages/gooddata-eval/tests/test_langfuse_item_scope.py, packages/gooddata-eval/tests/test_agentic_runner.py, packages/gooddata-eval/tests/test_agentic_join_trace.py
Agentic evaluators open item scopes and run evaluations inside them. The same identity and scope are passed to trace scoring.
Chat trace propagation and join recording
packages/gooddata-eval/src/gooddata_eval/core/chat/sse_client.py, packages/gooddata-eval/src/gooddata_eval/core/models.py, packages/gooddata-eval/tests/test_sse_client.py
The SSE client captures traceId, adds item-scope baggage to requests, assigns conversation run indexes, and records a join when the returned trace ID matches.
Joined trace lookup and root export
packages/gooddata-eval/src/gooddata_eval/core/agentic/_langfuse.py, packages/gooddata-eval/src/gooddata_eval/core/langfuse/experiment.py, packages/gooddata-eval/src/gooddata_eval/core/langfuse/observations.py, packages/gooddata-eval/tests/test_agentic_join_trace.py, packages/gooddata-eval/tests/test_langfuse_observations.py
Joined conversations are fetched by trace ID after their expected turns are ingested. Their experiment root uses the joined trace and span IDs. Trace summaries support traces without a parentless row.
Score collection, batching, and retry handling
packages/gooddata-eval/src/gooddata_eval/core/agentic/_langfuse.py, packages/gooddata-eval/src/gooddata_eval/core/agentic/_trace_linker.py, packages/gooddata-eval/src/gooddata_eval/core/langfuse/client.py, packages/gooddata-eval/tests/test_agentic_join_trace.py, packages/gooddata-eval/tests/test_langfuse_client.py, packages/gooddata-eval/tests/test_langfuse_e2e_fake_server.py
Score writes can be collected and posted in batches of up to 100. Retry delays vary by response status, and HTTP 207 partial rejections are logged without retrying.
Langfuse request and OTLP environment metadata
packages/gooddata-eval/src/gooddata_eval/core/langfuse/_env.py, packages/gooddata-eval/src/gooddata_eval/core/langfuse/otlp.py, packages/gooddata-eval/tests/test_langfuse_env.py, packages/gooddata-eval/tests/test_langfuse_otlp.py
HTTP requests include a generated User-Agent. OTLP resource attributes include service details and available GitHub environment values.

Priority: ⬇️ Low

Estimated code review effort: 4 (Complex) | ~60 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant Evaluator
  participant ItemTraceScope
  participant SSEClient
  participant ChatEndpoint
  participant JoinedRunRegistry
  Evaluator->>ItemTraceScope: Open item trace scope
  ItemTraceScope->>SSEClient: Provide trace and run baggage
  SSEClient->>ChatEndpoint: Send message with baggage
  ChatEndpoint->>SSEClient: Return response_started traceId
  SSEClient->>JoinedRunRegistry: Record conversation when traceId matches
Loading

Merge Risk: 🔵 Low · up to fa9a2

Long-running evaluations can retain joined-conversation state; reusing a conversation ID may delay trace lookup. This is a bounded risk that can be addressed with owner follow-up.

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage Warning Docstring coverage is 33.72% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 172 functions across 34 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check Passed Check skipped because no linked issues were found for this pull request.
Description Check Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check Passed The title clearly and concisely describes the main change: placing gen-AI turns under the evaluation item root trace.


  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR


  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

A rabbit checks the trace IDs bright,
And tucks each run into scope just right.
Baggage travels, turns arrive,
Batched scores keep records alive.
Joined roots bloom in Langfuse light.

Comment @coderabbitai help to get the list of available commands.

@tychtjan
tychtjan marked this pull request as ready for review October 9, 2026 12:52
With GOODDATA_EVAL_JOIN_GENAI_TRACE on, every chat request of an eval item
carries W3C baggage naming the item's trace and root span
(langfuse_trace_id, langfuse_experiment_item_root_observation_id) and its
experiment (langfuse_experiment_id/_name/_dataset_id/_item_id). A gen-ai
that accepts caller baggage opens its turn under that root and reports
the trace id on response_started. Both ids are derived from the
conversation id, so the sending thread and the deferred scorer agree
without passing state between threads.

For a joined conversation, the deferred task reads the trace by id
(/v2/observations?traceId=), waiting until every joined turn is ingested,
then exports gd-eval's root into that trace in the sdk-experiment
environment, spanning the whole run, and writes each score once on it.
Langfuse then reports the item's real cost and latency, and the score
rows are halved.

Conversations that did not join, and every run with the switch off,
keep today's path: poll by session, root in its own trace, scores on
both.

The item scope is set before the K runs in all twelve agentic kinds.
anomaly_detection and what_if now pass conversation_id to observe, so an
unlinked root there now carries session.id.

jira: trivial
risk: low
ChatClient.send_message passes each attempt's lines through _tap
(identity by default) before parsing. A subclass that needs more of the
stream than ChatResult keeps, such as tavern's obfuscation client
reading an error's reason, overrides _tap instead of copying
send_message, and so keeps the turn deadline and the eval-trace join.

jira: trivial
risk: low
Langfuse Cloud rate-limits public API calls per organization in fixed
one-minute windows, and score writes share the bucket with observation
and score reads (30/min on Hobby, 1,000/min on Pro). A 429 carries a
Retry-After counting down to the window's reset, often 30-60 s.

create_score capped every retry wait at 5 s, so both retries landed in
the same exhausted window and the score was dropped. On a Hobby project
a 60-request burst lost 43 scores this way. A 429 now waits up to the
60 s window; 5xx retries keep the 5 s cap.

jira: trivial
risk: low
Each eval item wrote every score with its own POST /api/public/scores,
twice when the item was not joined to gen-ai's trace. Those requests
share Langfuse's public-API rate-limit bucket with the observation and
score reads, so a nightly run spent most of its budget on score writes.

submit_trace_scoring now collects every score_safe call of the deferred
scoring step and sends them as POST /api/public/scores arrays of up to
100. Score ids are fixed before the first attempt, so a retried batch
cannot write a score twice. Langfuse validates array entries one by one
and answers 207 when some are rejected; that is logged, not retried.
Callers outside submit_trace_scoring (the single-shot sink, any client
without create_scores) still write each score as it comes.

Under a sustained 429 one item now waits at most 2 x 60 s, where
per-score writes with the 60 s Retry-After cap could wait 36 minutes,
past tavern's 420 s test timeout. A cancelled drain sends what was
collected before the interrupt once, without waiting out a retry.

jira: trivial
risk: low
Langfuse groups API traffic by User-Agent, so every gooddata-eval call
showed up as python-httpx/0.28.1, indistinguishable from other callers.
Its exported root spans carried only service.name, so a trace did not
say which build or CI run produced it.

Every client from make_http_client now sends
gooddata-eval/<version> (<ci|local>; run=<GITHUB_RUN_ID>; sha=<12
chars of GITHUB_SHA>), with "local" for any value unset or unsafe in a
header. The OTLP resource adds service.version and, when set,
github.run_id, github.workflow, github.ref_name and github.event_name.

jira: trivial
risk: low
@tychtjan
tychtjan force-pushed the jtyc/eval-single-trace branch from 096f2ec to fa9a2c4 Compare October 9, 2026 12:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant