Skip to content

chore: upgrade OpenTelemetry to 1.66.0 and add a metrics benchmark - #421

Merged
merlimat merged 3 commits into
oxia-db:mainfrom
merlimat:otel-1-66-upgrade
Oct 3, 2026
Merged

merlimat merged 3 commits into
oxia-db:mainfrom
merlimat:otel-1-66-upgrade

Conversation

@merlimat

@merlimat merlimat commented Oct 3, 2026

Copy link
Copy Markdown
Collaborator

Motivation

Every operation records its latency into a LatencyHistogram, and puts, gets, lists and range scans also add their size to a Counter. These calls happen on whichever threads complete the operations, all writing into the same series. With the OpenTelemetry SDK 1.63.0 that the build uses, they get much slower as threads are added:

The client only depends on opentelemetry-api, so applications pick up the gains by running SDK 1.66.0 or later. This PR moves the build, the tests and the perf tool to 1.66.0. The perf tool bundles the SDK, so its own measurements pay much less of this overhead.

Modifications

  • Bump opentelemetry from 1.63.0 to 1.66.0, and opentelemetry-bom-alpha from 1.63.0-alpha to 1.66.0-alpha, in gradle/libs.versions.toml.
  • Add MetricsBenchmark. It records into Counter, LatencyHistogram and a plain LongAdder (the baseline) from 1, 4, 8 and 16 threads, all into the same series.
    • The instruments are backed by an SdkMeterProvider with an InMemoryMetricReader. Without a reader the SDK hands out noop instruments.
    • Each instrument has four attributes, plus the oxia.namespace the client adds (and oxia.response.status for the histogram).
    • Each thread replays latencies from 100µs to ~200ms, so the records spread across the histogram buckets.

Release notes for 1.64 through 1.66, checked against the client:

  • API changes are limited to baggage, trace state and hex decoding, which the client doesn't use.
  • Breaking changes are in the incubator API, declarative config and the PrometheusMetricReader constructors, and the Zipkin exporter is no longer published. Neither the client nor the perf tool uses any of them; the perf tool goes through autoconfigure.
  • The Prometheus exporter now converts kBy instead of KBy (Align UCUM byte unit conversions with the specification table open-telemetry/opentelemetry-java#8752). The client's byte unit is By, which still maps to bytes.
  • Runtime dependencies: only opentelemetry-api, opentelemetry-context and opentelemetry-common change, from 1.63.0 to 1.66.0. The published POM imports the 1.66.0 BOMs.

Benchmark results

MetricsBenchmark with -prof gc on an Apple M1 Max (8 performance + 2 efficiency cores) and JDK 26.0.1. Values are ns per call as seen by each thread; lower is better. The LongAdder column doesn't depend on the SDK; it shows both runs to confirm they're comparable.

Threads LongAdder (1.63.0 / 1.66.0 run) Counter 1.63.0 → 1.66.0 LatencyHistogram 1.63.0 → 1.66.0
1 7.3 / 7.4 16.5 → 18.5 20.2 → 26.2
4 8.9 / 7.3 159 → 20.9 295 → 58.8
8 8.9 / 8.3 375 → 22.2 1,138 → 203
16 21 / 15 721 → 41.9 2,293 → 312
  • Counter now stays within 3× of a bare LongAdder at every thread count. On 1.63.0 it was up to 42× slower.
  • LatencyHistogram is 5–7× faster from 4 threads up. It's ~6 ns slower single-threaded, since each record now updates several adders instead of taking one uncontended lock. Past 4 threads it still slows down: 203 ns per call at 8 threads.
  • None of the three allocates per call.

Verification

  • ./gradlew build passes locally: 403 tests (400 in client, 3 in perf), no failures or skips, and no new compiler warnings. That includes all 7 Docker IT classes against a real Oxia server: OxiaClientIT, ClientReconnectIT, OxiaClientFailFastIT, RpcProviderIT, ShardSplitSessionIT, SharedResourcesIT and NotificationIt.
  • ./gradlew :client:dependencies --configuration runtimeClasspath on 1.63.0 and 1.66.0 differs only in the three OpenTelemetry modules and the BOMs.
  • The benchmark numbers above come from ./gradlew :benchmarks:jmhJar and java -jar benchmarks/build/libs/benchmarks-*-jmh.jar MetricsBenchmark -prof gc, once per version.

Compares Counter and LatencyHistogram against a plain LongAdder from 1, 4, 8 and 16 threads, all recording into the same series through an SDK meter provider.

Signed-off-by: Matteo Merli <mmerli@apache.org>
1.64.0 stops writing AggregatorHandle.valuesRecorded on every record, and 1.66.0 removes the lock from the explicit bucket histogram. Recording from many threads into the same series gets much cheaper.

Signed-off-by: Matteo Merli <mmerli@apache.org>
Signed-off-by: Matteo Merli <mmerli@apache.org>
@merlimat
merlimat merged commit a1368a9 into oxia-db:main Oct 3, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant