perf: replace the per-operation orTimeout with a per-client timeout sweeper - #419
Merged
merlimat merged 1 commit intoOct 3, 2026
Merged
Conversation
…weeper Every put, get, delete, deleteRange and list, and every read, list and range-scan stream attempt, called CompletableFuture.orTimeout, which schedules and cancels a task on the single JVM-wide delayer, taking its lock twice per operation. TimeoutSweeper replaces it: a future is appended to a list picked by the calling thread, the adding thread drops the completed futures when its list doubles, and a periodic sweep on the client executor fails the expired ones. Timeouts are approximate: never early, and at most two sweep intervals (a tenth of the timeout, up to 100 ms) late. Fixes oxia-db#394 Signed-off-by: Matteo Merli <mmerli@apache.org>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #394
Problem
Every
put,get,delete,deleteRangeandlistcallsCompletableFuture.orTimeout. On JDK 17 and 21, that schedules and then cancels a task on the single JVM-wideCompletableFuture.Delayerthread, taking its lock twice and allocating several objects per operation. Each read, list and range-scan attempt does the same for its stream establishment barrier inGrpcRpcProvider. With many producer threads, that lock is a contention point shared by the whole JVM.Example
AsyncOxiaClientImpl.put(and the other operations) ends with:Change
Added
io.oxia.client.util.TimeoutSweeper, owned by each client and eachGrpcRpcProvider:addappends the future to a list picked by the calling thread (striped by thread id), so threads rarely contend. The adding thread drops the completed futures whenever its list doubles, so memory follows the operations in flight rather than the request rate.TimeoutExceptionas before. It runs every tenth of the timeout, at most every 100 ms.orTimeoutwith their remaining time, since the owner shuts down the executor next.Behavior notes:
Design notes: a single shared queue only matched
orTimeoutwith 8 threads (the contention moved to the queue tail), and queues striped by thread collapsed under load because the single sweep thread could not drain them as fast as they filled. Having the adding thread compact its own list avoids both.Testing
TimeoutSweeperTest(7 tests): deadline, dropping completed futures, compaction by the adding thread, sweep interval, close, adding after close, and a real-executor check that a timeout never fires early.:client:testpasses locally (400 tests, including the integration tests); spotbugs and spotless are clean.TimeoutSweeperBenchmark(registering a timeout and completing the future), 10-core Mac, ops/µs:orTimeoutorTimeoutorTimeoutThe 24 B is the
CompletableFutureitself: registering the timeout allocates nothing. On JDK 26,orTimeoutis faster on one thread but slower with 8 threads than on JDK 17/21; that 8-thread number is noisy (an earlier run measured 5.2 ops/µs).Netty's
HashedWheelTimer(default 100 ms tick) behind the sameadd(), measured with a temporary variant of the benchmark that is not committed, ops/µs:That is slower than
orTimeout, especially with 8 threads, andorTimeoutspent only about 18 ms in GC per run (the sweeper 18–56 ms). The reasons, from the Netty 4.2.18 source:newTimeoutincrements a sharedAtomicLongand appends to a shared queue, eachcancelappends to a second one, and each operation allocates the timeout, the task and the cancel callback.