chore: upgrade OpenTelemetry to 1.66.0 and add a metrics benchmark - #421
Merged
Merged
Conversation
Compares Counter and LatencyHistogram against a plain LongAdder from 1, 4, 8 and 16 threads, all recording into the same series through an SDK meter provider. Signed-off-by: Matteo Merli <mmerli@apache.org>
1.64.0 stops writing AggregatorHandle.valuesRecorded on every record, and 1.66.0 removes the lock from the explicit bucket histogram. Recording from many threads into the same series gets much cheaper. Signed-off-by: Matteo Merli <mmerli@apache.org>
Signed-off-by: Matteo Merli <mmerli@apache.org>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Every operation records its latency into a
LatencyHistogram, and puts, gets, lists and range scans also add their size to aCounter. These calls happen on whichever threads complete the operations, all writing into the same series. With the OpenTelemetry SDK 1.63.0 that the build uses, they get much slower as threads are added:AggregatorHandlewrote its volatilevaluesRecordedflag on every record. Every thread then wrote to the same cache line, even though the counter value itself is aLongAdder. 1.64.0 only writes the flag when it's still false (Only set valuesRecorded if its false open-telemetry/opentelemetry-java#8559).LongAdderbucket counts, aDoubleAddersum and CAS-based min and max (Improve explicit histogram contention performance open-telemetry/opentelemetry-java#8717).The client only depends on
opentelemetry-api, so applications pick up the gains by running SDK 1.66.0 or later. This PR moves the build, the tests and the perf tool to 1.66.0. The perf tool bundles the SDK, so its own measurements pay much less of this overhead.Modifications
opentelemetryfrom 1.63.0 to 1.66.0, andopentelemetry-bom-alphafrom 1.63.0-alpha to 1.66.0-alpha, ingradle/libs.versions.toml.MetricsBenchmark. It records intoCounter,LatencyHistogramand a plainLongAdder(the baseline) from 1, 4, 8 and 16 threads, all into the same series.SdkMeterProviderwith anInMemoryMetricReader. Without a reader the SDK hands out noop instruments.oxia.namespacethe client adds (andoxia.response.statusfor the histogram).Release notes for 1.64 through 1.66, checked against the client:
PrometheusMetricReaderconstructors, and the Zipkin exporter is no longer published. Neither the client nor the perf tool uses any of them; the perf tool goes through autoconfigure.kByinstead ofKBy(Align UCUM byte unit conversions with the specification table open-telemetry/opentelemetry-java#8752). The client's byte unit isBy, which still maps tobytes.opentelemetry-api,opentelemetry-contextandopentelemetry-commonchange, from 1.63.0 to 1.66.0. The published POM imports the 1.66.0 BOMs.Benchmark results
MetricsBenchmarkwith-prof gcon an Apple M1 Max (8 performance + 2 efficiency cores) and JDK 26.0.1. Values are ns per call as seen by each thread; lower is better. TheLongAddercolumn doesn't depend on the SDK; it shows both runs to confirm they're comparable.LongAdder(1.63.0 / 1.66.0 run)Counter1.63.0 → 1.66.0LatencyHistogram1.63.0 → 1.66.0Counternow stays within 3× of a bareLongAdderat every thread count. On 1.63.0 it was up to 42× slower.LatencyHistogramis 5–7× faster from 4 threads up. It's ~6 ns slower single-threaded, since each record now updates several adders instead of taking one uncontended lock. Past 4 threads it still slows down: 203 ns per call at 8 threads.Verification
./gradlew buildpasses locally: 403 tests (400 inclient, 3 inperf), no failures or skips, and no new compiler warnings. That includes all 7 Docker IT classes against a real Oxia server:OxiaClientIT,ClientReconnectIT,OxiaClientFailFastIT,RpcProviderIT,ShardSplitSessionIT,SharedResourcesITandNotificationIt../gradlew :client:dependencies --configuration runtimeClasspathon 1.63.0 and 1.66.0 differs only in the three OpenTelemetry modules and the BOMs../gradlew :benchmarks:jmhJarandjava -jar benchmarks/build/libs/benchmarks-*-jmh.jar MetricsBenchmark -prof gc, once per version.