feat(AIC-3363): Support inline datasets - #95
aknight-ld wants to merge 5 commits into
Conversation
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 6748eaa. Configure here.
| SUMMARY_POLL_INTERVAL_SECONDS = 2.0 | ||
| SUMMARY_POLL_TIMEOUT_SECONDS = 180.0 | ||
| MAX_INLINE_TEXT_BYTES = 1_048_576 | ||
| MAX_ROWS = 10000 |
There was a problem hiding this comment.
We have a feature flag for this ai-evaluator-dataset-max-row-count
| DEFAULT_UI_BASE_URI = "https://app.launchdarkly.com" | ||
| SUMMARY_POLL_INTERVAL_SECONDS = 2.0 | ||
| SUMMARY_POLL_TIMEOUT_SECONDS = 180.0 | ||
| MAX_INLINE_TEXT_BYTES = 1_048_576 |
There was a problem hiding this comment.
Let's use a feature flag for this.
| key="support-qa-2026-08-20", | ||
| rows=[ | ||
| DatasetRow( | ||
| row_index=0, |
There was a problem hiding this comment.
This seems unnecessary to have the user specify the row_index. This can be inferred by the index in the array.
|
|
||
| ### 3. Letting an unserializable value into an evaluation event payload | ||
|
|
||
| An SDK event buffer is drained and serialized on a background thread, so a value the JSON encoder cannot encode is not reported back to the caller — the events are simply lost, and the run ends in a polling timeout with nothing to explain it. This is only reachable through `run(rows=[...])`, where `variables`/`metadata` hold arbitrary caller objects, which is why `_validate_rows` serialization-checks them up front with `allow_nan=False` and refuses to coerce. Do not relax that into a `default=str` rescue: silently stringifying a caller's value changes what a judge renders and what LaunchDarkly stores, and is unrecoverable once the row is persisted. |
There was a problem hiding this comment.
This comment here is a good example on why validation before event ingestion makes so much sense.
| return prepared | ||
|
|
||
| @staticmethod | ||
| def _validate_rows(rows: list[DatasetRow]) -> None: |
There was a problem hiding this comment.
This method is a good example on why validation should be done in the backend. There's a lot of domain logic in here that could change over time.
|
closing in favor of alternate approach |

Server side support of inline datasets is in flight; this PR covers the corresponding SDK work to pass in dataset rows defined from code.
Sizable test suite added in, testing underway to ensure we are matching the server.
Note
Overview
Adds inline datasets so
init_evaluations().run()accepts either a hosteddatasetkey or caller-suppliedrows=[DatasetRow(...)], with exactly one required and no dataset API traffic for inline runs.Inline rows are normalized through shared
_render_row(templates, injectedinput/expected_outputinvariables) without mutating caller objects. Pre-flight validation covers row count (10k), unique non-negativerow_index, JSON-serializablevariables/metadata, and rendered UTF-8 byte limits (1 MiB per field) so bad payloads fail before generation or silent event loss.Generation events for inline runs omit
datasetId/datasetKey(five-fieldeventId) and include rowinput,expectedOutput,variables, andmetadata; hosted runs keep excluding row bodies. Run creation sends{"source": "api"}withoutdatasetIdwhen inline. Criterion payloads treat dataset identifiers as optional.Docs in
packages/aiandpackages/clientREADMEs plusagents.mddescribe the new mode and pitfalls. A largetest_evaluations_run.pysuite covers API shapes, events, rendering, limits, and validation errors.Reviewed by Cursor Bugbot for commit d9971cd. Bugbot is set up for automated code reviews on this repo. Configure here.