Skip to content

Commit 1daa2c4

Browse files
Tomkessclaude
andauthored
feat(gooddata-eval): add the agentic anomaly-detection evaluator (#1801)
* feat(gooddata-eval): add the agentic anomaly-detection evaluator Scores the anomaly-detection skill end to end, with the granularity map aligned to what the service actually accepts and whole-conversation latency and cost rather than the goal turn's alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(gooddata-eval): let --gate decide an anomaly item, not pass@K alone `run_agentic_anomaly_detection` already computed `pass_power_k`; the evaluator never read it, asking `if not summary.pass_at_k` for the verdict, so `--gate power` on a flaky item reported a pass on the strength of one good run out of K. Wired the same way as the other multi-run kinds: the evaluator takes `gate`, the dispatch passes it, `gate_passed(...)` decides, the gate's note goes into the assertion message -- which matters because the body describes the BEST run, and under pass^K that can be one that passed -- and pass@K/pass^K/gate_passed reach Langfuse alongside the run metadata stamp. The default is unchanged: without --gate this is pass@K exactly as before, pinned by a test beside the power case. The power test fails against the previous code. Found by CodeRabbit on #1798. Forecasting is fixed there; what-if, which is already on master, in #1831. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
1 parent c3c5284 commit 1daa2c4

5 files changed

Lines changed: 1188 additions & 0 deletions

File tree

‎packages/gooddata-eval/src/gooddata_eval/cli/agentic_runner.py‎

Lines changed: 14 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -12,6 +12,7 @@
1212
from gooddata_eval.core.agentic._langfuse import make_langfuse_client
1313
from gooddata_eval.core.agentic._trace_linker import BackgroundTraceLinker, SubmitTraceLink, run_trace_link_inline
1414
from gooddata_eval.core.agentic.alert_skill import evaluate_agentic_alert_skill
15+
from gooddata_eval.core.agentic.anomaly_detection import evaluate_agentic_anomaly_detection
1516
from gooddata_eval.core.agentic.conversation import ConversationFixture, evaluate_agentic_conversation
1617
from gooddata_eval.core.agentic.dashboard_skill import evaluate_agentic_dashboard_skill
1718
from gooddata_eval.core.agentic.general_question import evaluate_agentic_general_question
@@ -48,6 +49,7 @@ class _LfKw(TypedDict, total=False):
4849
"agentic_guardrail",
4950
"agentic_conversation",
5051
"agentic_kda_skill",
52+
"agentic_anomaly_detection",
5153
"agentic_what_if",
5254
}
5355
)
@@ -278,6 +280,18 @@ def _dispatch_agentic(
278280
agent_id=agent_id,
279281
**lf_kw,
280282
)
283+
elif kind == "agentic_anomaly_detection":
284+
return evaluate_agentic_anomaly_detection(
285+
host=host,
286+
token=token,
287+
workspace_id=workspace_id,
288+
question=item.question,
289+
expected_output=eo if isinstance(eo, dict) else {},
290+
k=k,
291+
gate=gate,
292+
agent_id=agent_id,
293+
**lf_kw,
294+
)
281295
elif kind == "agentic_conversation":
282296
fixture_data = eo.get("fixture") or eo if isinstance(eo, dict) else {}
283297
return evaluate_agentic_conversation(

0 commit comments

Comments
 (0)