Repository navigation
Conversation
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #1862 +/- ##
==========================================
+ Coverage 84.19% 84.32% +0.12%
==========================================
Files 333 335 +2
Lines 23261 23439 +178
==========================================
+ Hits 19585 19764 +179
+ Misses 3676 3675 -1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
romrak
marked this pull request as ready for review
October 10, 2026 08:44
Publisher calls its output a document, so the evaluator, its test kind and its check names now say document: agentic_document_skill, document_skill.py, Document* classes and document_* scores. The agentic_report_skill test kind is gone: its one dataset has to be migrated for the Document wire names anyway, and the Tavern shim calls the evaluator function, not the kind. The Python names stay for existing callers: report_skill.py re-exports the old names with a DeprecationWarning, and core.agentic keeps the old aliases. The wire names gen-ai sends are unchanged here; the evaluator still reads the report part. jira: LX-3248 risk: low Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The evaluator now reads the names gen-ai sends after the Publisher rename: the document answer part with document_ref and the saved_document_id and base_document_id keys, the draft_document tool and the document_builder skill. A new strict check, document_wire_names, fails a run when any Report name is still on the wire: the report part type or keys, the draft_report or list_report_layouts tool, or the report_builder skill. It reads every turn, and its failure leads the list and names each old name, so a half-switched chain shows which part still sends the old one. A user context with view.report is rejected before the first request, because gen-ai reads view.document. This evaluator fails against a gen-ai without the Document names, so release it only once gen-ai sends them. jira: LX-3248 risk: high Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
romrak
force-pushed
the
rr/LX-3248-document-names
branch
from
October 10, 2026 08:50
b41dd40 to
92fe97e
Compare
hkad98
approved these changes
Oct 10, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
gd-eval's Publisher copilot evaluator now uses the Document names and fails a run when gen-ai still sends a Report name. Do not merge before gen-ai ships the Document wire names (LX-3244): against today's gen-ai every document eval fails, by design.
Jira: LX-3248, epic GDP-3485. Wire names as proposed on LX-3244.
Two commits, review them one by one:
refactor— rename only. Evaluator, test kind and score names go from report to document; the old Python names keep working.feat— the wire names flip, plus the newdocument_wire_namescheck.What this implements
agentic_report_skillagentic_document_skillcore/agentic/report_skill.pycore/agentic/document_skill.pyreport_drafted,report_ref_matches, …document_drafted,document_ref_matches, …, plusdocument_wire_namestype: "report",report,report_ref,saved_report_id,base_report_idtype: "document",document,document_ref,saved_document_id,base_document_iddraft_report,report_builderdraft_document,document_builderview.reportpassed throughview.document;view.reportrejected before the first requestWhat a run against today's gen-ai reports (from the unit test that pins it):
Decisions
A fallback would pass a half-switched chain, and the point is to prove every link was renamed.
src/gooddata_eval/core/agentic/document_skill.py—legacy_wire_names,_LEGACY_*constantsdraft_reportinstead of replying up tomax_iterationstimes.document_wire_namesis a strict check published on every run, and its failure leads the list.Without it an old gen-ai reads as "never produced a draft_document call", which hides why. It scans every turn: part type and keys, the embedded document's type, the old tools, and
report_builderin either theset_skillsarguments or its result.quality_scoremoves (4/6 → 5/7 when it passes). Scores before and after this PR are not comparable anyway, because the names changed.gdc-nas imports
gooddata_eval.core.agentic.report_skilldirectly, so its names stay. Theagentic_report_skillkind is gone: its one dataset needs migrating for the Document wire names anyway, and the Tavern shim calls the function, not the kind.core/agentic/report_skill.pyre-exports the old names with aDeprecationWarning.core/agentic/__init__.pykeeps the old aliases.agentic_report_skillis skipped as an unsupportedtest_kindand listed at the end of the run.report_*todocument_*.Score history splits at this release: old runs keep
report_*, new runs getdocument_*.What comes next
tests/tavern-e2e/ci/combo_report.py_REPORT_*_FIELDS→document_*, thereport_skill_agentic.pyimport path and the Langfuse dataset kind.test_kindagentic_report_skill→agentic_document_skill, andview.report→view.documentin their user context.Test plan
toxpy310–py314 for gooddata-eval: all green (1622 passed, 1 skipped on each).ruff format --checkandruff checkrepo-wide: clean.ty check: no new diagnostics insrc; the test file keeps its existing test-double diagnostics.risk: high — a published package that stops passing against current gen-ai.
🤖 Generated with Claude Code
Summary by CodeRabbit