Skip to content

feat(AIC-3363): Support inline datasets - #95

Closed
aknight-ld wants to merge 5 commits into
mainfrom
AIC-3363-inline-datasets
Closed

aknight-ld wants to merge 5 commits into
mainfrom
AIC-3363-inline-datasets

Conversation

@aknight-ld

@aknight-ld aknight-ld commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Server side support of inline datasets is in flight; this PR covers the corresponding SDK work to pass in dataset rows defined from code.

Sizable test suite added in, testing underway to ensure we are matching the server.


Note

Overview
Adds inline datasets so init_evaluations().run() accepts either a hosted dataset key or caller-supplied rows=[DatasetRow(...)], with exactly one required and no dataset API traffic for inline runs.

Inline rows are normalized through shared _render_row (templates, injected input/expected_output in variables) without mutating caller objects. Pre-flight validation covers row count (10k), unique non-negative row_index, JSON-serializable variables/metadata, and rendered UTF-8 byte limits (1 MiB per field) so bad payloads fail before generation or silent event loss.

Generation events for inline runs omit datasetId/datasetKey (five-field eventId) and include row input, expectedOutput, variables, and metadata; hosted runs keep excluding row bodies. Run creation sends {"source": "api"} without datasetId when inline. Criterion payloads treat dataset identifiers as optional.

Docs in packages/ai and packages/client READMEs plus agents.md describe the new mode and pitfalls. A large test_evaluations_run.py suite covers API shapes, events, rendering, limits, and validation errors.

Reviewed by Cursor Bugbot for commit d9971cd. Bugbot is set up for automated code reviews on this repo. Configure here.

@aknight-ld
aknight-ld marked this pull request as ready for review September 18, 2026 18:58

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 6748eaa. Configure here.

Comment thread packages/client/src/launchdarkly_ai_server/evaluations/module.py Outdated
SUMMARY_POLL_INTERVAL_SECONDS = 2.0
SUMMARY_POLL_TIMEOUT_SECONDS = 180.0
MAX_INLINE_TEXT_BYTES = 1_048_576
MAX_ROWS = 10000

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We have a feature flag for this ai-evaluator-dataset-max-row-count

DEFAULT_UI_BASE_URI = "https://app.launchdarkly.com"
SUMMARY_POLL_INTERVAL_SECONDS = 2.0
SUMMARY_POLL_TIMEOUT_SECONDS = 180.0
MAX_INLINE_TEXT_BYTES = 1_048_576

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's use a feature flag for this.

Comment thread packages/client/README.md
key="support-qa-2026-08-20",
rows=[
DatasetRow(
row_index=0,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This seems unnecessary to have the user specify the row_index. This can be inferred by the index in the array.

Comment thread packages/client/agents.md

### 3. Letting an unserializable value into an evaluation event payload

An SDK event buffer is drained and serialized on a background thread, so a value the JSON encoder cannot encode is not reported back to the caller — the events are simply lost, and the run ends in a polling timeout with nothing to explain it. This is only reachable through `run(rows=[...])`, where `variables`/`metadata` hold arbitrary caller objects, which is why `_validate_rows` serialization-checks them up front with `allow_nan=False` and refuses to coerce. Do not relax that into a `default=str` rescue: silently stringifying a caller's value changes what a judge renders and what LaunchDarkly stores, and is unrecoverable once the row is persisted.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This comment here is a good example on why validation before event ingestion makes so much sense.

return prepared

@staticmethod
def _validate_rows(rows: list[DatasetRow]) -> None:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This method is a good example on why validation should be done in the backend. There's a lot of domain logic in here that could change over time.

@aknight-ld

Copy link
Copy Markdown
Contributor Author

closing in favor of alternate approach

@aknight-ld aknight-ld closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants