Skip to content

Flush LLM Obs intake writer synchronously in serverless environments - #12249

Open
purple4reina wants to merge 12 commits into
masterfrom
rey.abolofia/llmobs-serverless-writer-fix
Open

purple4reina wants to merge 12 commits into
masterfrom
rey.abolofia/llmobs-serverless-writer-fix

Conversation

@purple4reina

@purple4reina purple4reina commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

Context

dd-trace-java's LLM Obs writer doesn't respect the S in CBAS(H) (note that the H is silent). The aws lambda runtime gets suspended between invocations, yet LLM Obs data was being sent on a periodic schedule. This meant if the function was not actively running when the flush struck, data wouldn't be sent.

Screenshot 2026-08-26 at 13 37 04

Summary

The DD_INTAKE_WRITER_TYPE branch of WriterFactory.createWriter() — which is what actually carries LLM Obs spans, since Agent.java defaults the writer type to MultiWriter:DDIntakeWriter,DDAgentWriter whenever LLM Obs is enabled — never set alwaysFlush. In Lambda, the execution environment can freeze as soon as the handler returns, before this writer's periodic flush timer next fires, so spans were silently dropped most of the time.

This surfaced as ml_obs.trace rarely/never being emitted for java Lambda functions in serverless-e2e-tests.

What changed

Set alwaysFlush on the DDIntakeWriter branch the same way the DDAgentWriter branch already does (config.isAgentConfiguredUsingDefault() && ServerlessInfo.get().isRunningInServerlessEnvironment()), and drop the DDAgentWriter branch's separate LLM Obs DDIntakeWriter + MultiWriter — it duplicated the track the intake branch already builds, so every span was being sent twice (visible as ml_obs.trace reporting exactly 2x the invocation count).

In Lambda, TraceProcessingWorker.flush() now also waits on the secondary queue, so sampled-out traces (e.g. DD_TRACE_SAMPLE_RATE=0) aren't left behind at freeze. The check runs once, from ServerlessInfo, when the worker is built. Outside Lambda it still enqueues a single marker. The one behavior change there: enqueuing a marker now counts against the flush timeout instead of spinning until the queue has room, so a flush against a full queue returns false at the timeout.

Validation

Built and deployed custom Lambda layers to a sandbox stack to isolate each half of this fix:

  • Pre-fix code (no alwaysFlush on the intake branch): 0/8 invocations delivered an LLM Obs span.
  • Same code + only the alwaysFlush change: 8/8 invocations delivered exactly one span each, confirmed via CloudWatch logs (Successfully sent 1 traces ... LLMOBS/v2).
  • Confirmed via live invocation of the original MultiWriter fix that it sent every span twice (two independent LLMOBS/v2 payloads per invocation).

@purple4reina
purple4reina requested a review from a team as a code owner August 20, 2026 20:45
@purple4reina
purple4reina requested review from ValentinZakharov and removed request for a team August 20, 2026 20:45
@dd-octo-sts

dd-octo-sts Bot commented Aug 20, 2026

Copy link
Copy Markdown
Contributor

Hi! 👋 Thanks for your pull request! 🎉

To help us review it, please make sure to:

  • Add at least one type, and one component or instrumentation label to the pull request

If you need help, please check our contributing guidelines.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4196c2186d

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread dd-trace-core/src/main/java/datadog/trace/common/writer/WriterFactory.java Outdated
Comment thread dd-trace-core/src/main/java/datadog/trace/common/writer/WriterFactory.java Outdated
Comment thread dd-trace-core/src/main/java/datadog/trace/common/writer/WriterFactory.java Outdated

@datadog-datadog-prod-us1 datadog-datadog-prod-us1 Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Datadog Autotest: FAIL

When LLM Observability and AppSec API Security run together outside serverless mode, both writer threads can post-process the same trace. This can close one AppSec request context twice and release its limit permit twice.

Open Bits AI session

🤖 Datadog Autotest · Commit 4196c21 · What is Autotest? · @DataDog review to ask questions · Any feedback? Reach out in #autotest

Comment thread dd-trace-core/src/main/java/datadog/trace/common/writer/WriterFactory.java Outdated
@datadog-datadog-prod-us1

This comment has been minimized.

@dd-octo-sts

dd-octo-sts Bot commented Aug 20, 2026 •

Copy link
Copy Markdown
Contributor

🟢 Java Benchmark SLOs — All performance SLOs passed

Suite Status
Startup 🟢 pass

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 13.86 s 13.93 s [-1.2%; +0.3%] (no difference)
startup:insecure-bank:tracing:Agent 12.85 s 12.96 s [-1.5%; -0.1%] (maybe better)
startup:petclinic:appsec:Agent 17.68 s 17.59 s [-0.5%; +1.5%] (no difference)
startup:petclinic:iast:Agent 17.29 s 17.57 s [-2.4%; -0.7%] (maybe better)
startup:petclinic:profiling:Agent 17.19 s 17.06 s [-0.2%; +1.6%] (no difference)
startup:petclinic:sca:Agent 17.68 s 17.66 s [-0.8%; +1.0%] (no difference)
startup:petclinic:tracing:Agent 16.67 s 16.20 s [-1.5%; +7.2%] (no difference)

Commit: b85d757e · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

@purple4reina purple4reina changed the title Fix LLM Obs spans silently dropped when using DDAgentWriter (Lambda/serverless) Flush LLM Obs intake writer synchronously in serverless environments Aug 24, 2026
@purple4reina

Copy link
Copy Markdown
Contributor Author

Layer published from current commit (b83d309)

arn:aws:lambda:us-west-2:425362996713:layer:dd-trace-java-serverless-writer-fix-debug:1

E2e pipeline running this layer (executing just java on lambda-features suite) https://gitlab.ddbuild.io/DataDog/serverless-e2e-tests/-/pipelines/133218539

@purple4reina

Copy link
Copy Markdown
Contributor Author

Pipeline passed. Java numbers are all at 5's now, showing that this PR fixes the underlying issue.

Screenshot 2026-08-26 at 13 34 12

@purple4reina purple4reina added type: bug fix Bug fix comp: mlobs ML Observability (LLMObs) inst: aws lambda AWS Lambda instrumentation labels Aug 26, 2026
@AlexeyKuznetsov-DD

Copy link
Copy Markdown
Contributor

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b83d309c8f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread dd-trace-core/src/main/java/datadog/trace/common/writer/WriterFactory.java Outdated
Comment thread dd-trace-core/src/main/java/datadog/trace/common/writer/WriterFactory.java Outdated
Comment thread dd-trace-core/src/main/java/datadog/trace/common/writer/WriterFactory.java Outdated
@purple4reina
purple4reina force-pushed the rey.abolofia/llmobs-serverless-writer-fix branch from b83d309 to ced2ef1 Compare August 27, 2026 17:53
@dd-octo-sts dd-octo-sts Bot added the tag: ai generated Largely based on code generated by an AI or LLM label Aug 28, 2026
@purple4reina

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 49858a3b59

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

@dougqh dougqh left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since this change effectively only applies in serverless, I think this change is fine from performance perspective.

@bric3 bric3 left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@dougqh Indeed automatic flushing (.alwaysFlush(isServerlessDefault)) after each write is serverless-only, but I believe the change to TraceProcessingWorker method affect all callers of flush. which makes me careful on this change.

After some code search, it seems it affects in particular, and writer shutdown and CI Visibility. Which makes these callers now wait for the secondary queue as well. That it might be a good thing to properly handle this.

Since it affects more usecases, I believe my concern on the timeout needs to be addressed.


@purple4reina In my comment I give some code example which I believe to be fixing the issue with deadline. Note I also suggest a test that exercises this scenario.

@purple4reina
purple4reina requested a review from bric3 September 28, 2026 21:43
purple4reina and others added 6 commits September 28, 2026 14:43
The DDAgentWriter branch of WriterFactory never wired up an LLM Obs
DDIntakeWriter, so LLM Obs spans were silently dropped whenever the
tracer picked DDAgentWriter (e.g. in Lambda/serverless mode with CI
Visibility disabled). Broadcast to a dedicated LLM Obs DDIntakeWriter
via MultiWriter, flushed synchronously in serverless environments
just like the primary writer.
The DDAgentWriter branch's dedicated LLM Obs DDIntakeWriter duplicated the
LLM Obs track that the DD_INTAKE_WRITER_TYPE branch already builds whenever
llmObsEnabled is true (Agent.java defaults to
MultiWriter:DDIntakeWriter,DDAgentWriter for that case), so every span was
sent twice. Verified live on Lambda: the pre-fix build sent LLMOBS/v2
payloads twice per invocation.

The original bug this branch was fixing -- spans silently dropped in
Lambda -- was actually caused by the DDIntakeWriter branch never setting
alwaysFlush, so the periodic flush timer rarely beat the execution
environment freezing after the handler returns. Verified live: without
alwaysFlush, 0/8 invocations delivered a span; with it, 8/8 did, one send
each.
Sampled-out traces route to the secondary queue, which flush() never
touched, so a Lambda invocation whose execution environment freezes
right after the synchronous flush returns can lose them. Verified live:
1/15 delivered without this fix, 15/15 with it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Both the DDIntakeWriter and DDAgentWriter branches computed the same
config.isAgentConfiguredUsingDefault() && isRunningInServerlessEnvironment()
condition separately; hoist it into a single isServerlessDefault local.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 2, 2026
@purple4reina
purple4reina added this pull request to the merge queue Oct 5, 2026
@dd-octo-sts

dd-octo-sts Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-10-05 16:53:09 UTC ℹ️ Start processing command /merge


2026-10-05 16:53:13 UTC ℹ️ MergeQueue: pull request added to the queue

The expected merge time in master is approximately 1h (p90).


2026-10-05 17:27:00 UTC ❌ MergeQueue: The build pipeline contains failing jobs for this merge request

Build pipeline has failing jobs for 38a42d6:

⚠️ Do NOT retry failed jobs directly (why?).

What to do next?

  • Investigate the failures and when ready, re-add your pull request to the queue!
  • If your PR checks are green, try to rebase/merge. It might be because the CI run is a bit old.
  • Any question, go check the FAQ.
Details

Since those jobs are not marked as being allowed to fail, the pipeline will most likely fail.
Therefore, and to allow other builds to be processed, this merge request has been rejected and the pipeline got canceled.

@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 5, 2026
@purple4reina
purple4reina added this pull request to the merge queue Oct 6, 2026
@dd-octo-sts

dd-octo-sts Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-10-06 18:47:54 UTC ℹ️ Start processing command /merge


2026-10-06 18:47:59 UTC ℹ️ MergeQueue: pull request added to the queue

The expected merge time in master is approximately 1h (p90).


2026-10-06 19:16:45 UTC ❌ MergeQueue: The build pipeline contains failing jobs for this merge request

Build pipeline has failing jobs for 59fb1c5:

⚠️ Do NOT retry failed jobs directly (why?).

What to do next?

  • Investigate the failures and when ready, re-add your pull request to the queue!
  • If your PR checks are green, try to rebase/merge. It might be because the CI run is a bit old.
  • Any question, go check the FAQ.
Details

Since those jobs are not marked as being allowed to fail, the pipeline will most likely fail.
Therefore, and to allow other builds to be processed, this merge request has been rejected and the pipeline got canceled.

@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 6, 2026
@purple4reina
purple4reina enabled auto-merge October 7, 2026 20:23
@purple4reina
purple4reina added this pull request to the merge queue Oct 8, 2026
@dd-octo-sts

dd-octo-sts Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-10-08 17:34:59 UTC ℹ️ Start processing command /merge


2026-10-08 17:35:04 UTC ℹ️ MergeQueue: pull request added to the queue

The expected merge time in master is approximately 1h (p90).


2026-10-08 17:58:56 UTC ❌ MergeQueue: The build pipeline contains failing jobs for this merge request

Build pipeline has failing jobs for 347fd35:

⚠️ Do NOT retry failed jobs directly (why?).

What to do next?

  • Investigate the failures and when ready, re-add your pull request to the queue!
  • If your PR checks are green, try to rebase/merge. It might be because the CI run is a bit old.
  • Any question, go check the FAQ.
Details

Since those jobs are not marked as being allowed to fail, the pipeline will most likely fail.
Therefore, and to allow other builds to be processed, this merge request has been rejected and the pipeline got canceled.

@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 8, 2026
@purple4reina
purple4reina added this pull request to the merge queue Oct 8, 2026
@dd-octo-sts

dd-octo-sts Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-10-08 18:48:13 UTC ℹ️ Start processing command /merge


2026-10-08 18:48:17 UTC ℹ️ MergeQueue: pull request added to the queue

The expected merge time in master is approximately 1h (p90).


2026-10-08 19:23:24 UTC ❌ MergeQueue: The build pipeline contains failing jobs for this merge request

Build pipeline has failing jobs for 031fb73:

⚠️ Do NOT retry failed jobs directly (why?).

What to do next?

  • Investigate the failures and when ready, re-add your pull request to the queue!
  • If your PR checks are green, try to rebase/merge. It might be because the CI run is a bit old.
  • Any question, go check the FAQ.
Details

Since those jobs are not marked as being allowed to fail, the pipeline will most likely fail.
Therefore, and to allow other builds to be processed, this merge request has been rejected and the pipeline got canceled.

@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 8, 2026
@purple4reina
purple4reina added this pull request to the merge queue Oct 9, 2026
@dd-octo-sts

dd-octo-sts Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-10-09 18:19:11 UTC ℹ️ Start processing command /merge


2026-10-09 18:19:16 UTC ℹ️ MergeQueue: pull request added to the queue

The expected merge time in master is approximately 1h (p90).


2026-10-09 18:40:22 UTC ❌ MergeQueue: The build pipeline contains failing jobs for this merge request

Build pipeline has failing jobs for 35c776b:

⚠️ Do NOT retry failed jobs directly (why?).

What to do next?

  • Investigate the failures and when ready, re-add your pull request to the queue!
  • If your PR checks are green, try to rebase/merge. It might be because the CI run is a bit old.
  • Any question, go check the FAQ.
Details

Since those jobs are not marked as being allowed to fail, the pipeline will most likely fail.
Therefore, and to allow other builds to be processed, this merge request has been rejected and the pipeline got canceled.

@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 9, 2026
@purple4reina
purple4reina added this pull request to the merge queue Oct 9, 2026
@dd-octo-sts

dd-octo-sts Bot commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

/merge

@gh-worker-devflow-routing-ef8351

gh-worker-devflow-routing-ef8351 Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

View all feedbacks in Devflow UI.

2026-10-09 19:38:32 UTC ℹ️ Start processing command /merge


2026-10-09 19:38:37 UTC ℹ️ MergeQueue: pull request added to the queue

The expected merge time in master is approximately 1h (p90).


2026-10-09 19:59:05 UTC ❌ MergeQueue: This merge request was updated

This PR is rejected because it was updated

@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Oct 9, 2026
purple4reina and others added 2 commits October 9, 2026 12:58
The merge queue reported infra failures (git checkout timeouts) for
some attempts, but every attempt that got far enough also failed
DemoExecutorServiceTest on IBM JDK 8 (7 of 7 runs since Oct 2; the test
passes on ibm8 elsewhere). The queue cancels the pipeline at the first
failed job, so the bot comment does not always show this failure.

Cause: on IBM JDK 8 the agent delays starting the writer by at least
100ms (okHttpDelayMillis, because of IBMSASL). The ExecutorService app
exits before that, so the shutdown flush runs before the serializer
thread starts. This branch checked serializerThread.isAlive() before
the first offer, so flush() returned false at once, the writer closed,
and the trace was never sent. Before this branch, flush() enqueued the
FlushEvent and waited, and the late-starting serializer drained it.

Fix: always offer once, and check the thread only while the queue is
full. The new test fails without this change (expected true, was
false) on any JVM.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp: mlobs ML Observability (LLMObs) inst: aws lambda AWS Lambda instrumentation tag: ai generated Largely based on code generated by an AI or LLM type: bug fix Bug fix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants