Skip to content

Reduce system test sharding overhead - #12787

Draft
cgcote wants to merge 6 commits into
masterfrom
ci/measure-system-test-sharding
Draft

cgcote wants to merge 6 commits into
masterfrom
ci/measure-system-test-sharding

Conversation

@cgcote

@cgcote cgcote commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

What does this PR do?

Tunes the upstream system-test desired_execution_time using measured CI Visibility durations to minimize end-to-end workflow wall time. This is isolated from the separate Gradle clean experiment in #12785.

Selected target: 12.5 minutes (248 shards). It stays below the upstream 256-job cap while reducing the slow tail substantially.

Results

Target Pipeline Shards Median p95 Slowest Runner time <5m
15m control 37m43s 193 9m54s 18m25s 24m15s 1,913.7m 20
20m 39m36s 134 14m12s 25m52s 29m00s 1,849.9m 14
12.5m 29m46s 248 8m07s 16m20s 19m19s 1,977.4m 47
12.5m repeat 32m19s 248 8m07s 16m01s 19m39s 2,004.5m 42
12m 34m24s 256 7m38s 15m23s 20m58s 1,997.6m 43
13m 35m13s 235 8m27s 17m21s 22m03s 2,071.4m 31

The two 12.5-minute runs average 31m03s, 6m41s or 17.7% faster than the 15-minute control. Both also reduced the slowest shard by more than four minutes.

The 20-minute treatment used fewer runner-minutes but slowed the pipeline. The 12-minute treatment hit the 256-shard cap and regressed wall time. The 13-minute treatment reduced startup count but regressed wall time, p95, maximum, and runner cost. Together these bound the fastest observed target at 12.5 minutes.

The planner consumes historical duration data from system-tests time-stats.json and uses first-fit-decreasing allocation; it does not currently query live CI Visibility data. CI Visibility is used here to measure and select the target.

Inspired by DataDog/dd-trace-rb#6439, while keeping setup-cost changes outside this PR.

Validation

  • All six measured experiment workflows completed successfully.
  • The selected 12.5-minute treatment was repeated successfully.
  • Workflow YAML parses successfully.
  • git diff --check passes.

@cgcote cgcote added tag: no release notes Changes to exclude from release notes type: refactoring comp: tooling Build & Tooling tag: ai generated Largely based on code generated by an AI or LLM labels Oct 8, 2026
@datadog-datadog-prod-us1

This comment has been minimized.

@dd-octo-sts

dd-octo-sts Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

🟢 Java Benchmark SLOs — All performance SLOs passed

Suite Status
Startup 🟢 pass

SLO thresholds are defined here based on automatically generated metrics. A warning is raised when results are within 5% of the threshold.

PR vs. master results
Scenario Candidate master Δ (95% CI of mean)
startup:insecure-bank:iast:Agent 13.88 s 13.85 s [-0.4%; +0.9%] (no difference)
startup:insecure-bank:tracing:Agent 12.81 s 12.92 s [-1.6%; -0.2%] (maybe better)
startup:petclinic:appsec:Agent 17.10 s 16.99 s [-0.2%; +1.6%] (no difference)
startup:petclinic:iast:Agent 16.88 s 16.91 s [-1.1%; +0.7%] (no difference)
startup:petclinic:profiling:Agent 16.54 s 16.71 s [-1.8%; -0.2%] (maybe better)
startup:petclinic:sca:Agent 17.05 s 16.97 s [-0.5%; +1.4%] (no difference)
startup:petclinic:tracing:Agent 16.14 s 16.08 s [-0.6%; +1.3%] (no difference)

Commit: 1f529938 · CI Pipeline · Benchmarking Platform UI


Load and DaCapo benchmarks can be triggered manually in the GitLab pipeline. Results will appear in the Benchmarking Platform UI after completion.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp: tooling Build & Tooling tag: ai generated Largely based on code generated by an AI or LLM tag: no release notes Changes to exclude from release notes type: refactoring

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant