Skip to content

perf: replace the per-operation orTimeout with a per-client timeout sweeper - #419

Merged
merlimat merged 1 commit into
oxia-db:mainfrom
merlimat:wt/oxia-java-approximate-timeouts-0099ca
Oct 3, 2026
Merged

merlimat merged 1 commit into
oxia-db:mainfrom
merlimat:wt/oxia-java-approximate-timeouts-0099ca

Conversation

@merlimat

@merlimat merlimat commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Fixes #394

Problem

Every put, get, delete, deleteRange and list calls CompletableFuture.orTimeout. On JDK 17 and 21, that schedules and then cancels a task on the single JVM-wide CompletableFuture.Delayer thread, taking its lock twice and allocating several objects per operation. Each read, list and range-scan attempt does the same for its stream establishment barrier in GrpcRpcProvider. With many producer threads, that lock is a contention point shared by the whole JVM.

Example

AsyncOxiaClientImpl.put (and the other operations) ends with:

return callback
        .orTimeout(requestTimeoutMs, TimeUnit.MILLISECONDS)
        .whenComplete(...);

Change

Added io.oxia.client.util.TimeoutSweeper, owned by each client and each GrpcRpcProvider:

return requestTimeouts
        .add(callback)
        .whenComplete(...);
  • add appends the future to a list picked by the calling thread (striped by thread id), so threads rarely contend. The adding thread drops the completed futures whenever its list doubles, so memory follows the operations in flight rather than the request rate.
  • A periodic sweep on the client's executor fails the expired futures with the same TimeoutException as before. It runs every tenth of the timeout, at most every 100 ms.
  • Timeouts are approximate: never early, and at most two sweep intervals late (30.0–30.2 s with the default 30 s timeout).
  • On close, the futures still pending are handed back to orTimeout with their remaining time, since the owner shuts down the executor next.

Behavior notes:

  • The callbacks of a timed-out operation now run on the client's executor thread instead of the JDK delayer thread.
  • Each client now runs two light periodic tasks (operations and RPC barriers), each waking at most every 100 ms.
  • The write stream's own timeout check and the Failsafe timeouts on session RPCs are unchanged: they are not on the per-operation path.

Design notes: a single shared queue only matched orTimeout with 8 threads (the contention moved to the queue tail), and queues striped by thread collapsed under load because the single sweep thread could not drain them as fast as they filled. Having the adding thread compact its own list avoids both.

Testing

  • New TimeoutSweeperTest (7 tests): deadline, dropping completed futures, compaction by the adding thread, sweep interval, close, adding after close, and a real-executor check that a timeout never fires early.
  • Full :client:test passes locally (400 tests, including the integration tests); spotbugs and spotless are clean.
  • New TimeoutSweeperBenchmark (registering a timeout and completing the future), 10-core Mac, ops/µs:
JDK 17 orTimeout JDK 17 sweeper JDK 21 orTimeout JDK 21 sweeper JDK 26 orTimeout JDK 26 sweeper
1 thread 8.6 71.8 8.8 71.4 12.5 76.6
8 threads 7.5 306 7.8 305 2.7 366
Allocation per op 216 B 24 B 216 B 24 B 200 B 24 B

The 24 B is the CompletableFuture itself: registering the timeout allocates nothing. On JDK 26, orTimeout is faster on one thread but slower with 8 threads than on JDK 17/21; that 8-thread number is noisy (an earlier run measured 5.2 ops/µs).

Netty's HashedWheelTimer (default 100 ms tick) behind the same add(), measured with a temporary variant of the benchmark that is not committed, ops/µs:

JDK 17 JDK 21
1 thread 7.6 7.9
8 threads 3.7 3.6
Allocation per op 184 B 184 B
GC time per 10 s run 1.4–3.1 s 1.3–3.2 s

That is slower than orTimeout, especially with 8 threads, and orTimeout spent only about 18 ms in GC per run (the sweeper 18–56 ms). The reasons, from the Netty 4.2.18 source:

  • Each newTimeout increments a shared AtomicLong and appends to a shared queue, each cancel appends to a second one, and each operation allocates the timeout, the task and the cancel callback.
  • The worker moves at most 100,000 new timeouts into the wheel per tick, about 1M/s with a 100 ms tick. Above that its queue grows without bound, hence the GC time: those rates aren't sustainable.
  • Each instance starts its own worker thread, and Netty logs an error past 64 instances, so a timer per client doesn't scale; a single shared timer would again be a JVM-wide contention point.

…weeper

Every put, get, delete, deleteRange and list, and every read, list and
range-scan stream attempt, called CompletableFuture.orTimeout, which
schedules and cancels a task on the single JVM-wide delayer, taking its
lock twice per operation.

TimeoutSweeper replaces it: a future is appended to a list picked by the
calling thread, the adding thread drops the completed futures when its
list doubles, and a periodic sweep on the client executor fails the
expired ones. Timeouts are approximate: never early, and at most two
sweep intervals (a tenth of the timeout, up to 100 ms) late.

Fixes oxia-db#394

Signed-off-by: Matteo Merli <mmerli@apache.org>
@merlimat
merlimat merged commit 518006f into oxia-db:main Oct 3, 2026
2 checks passed
@merlimat
merlimat deleted the wt/oxia-java-approximate-timeouts-0099ca branch October 3, 2026 17:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Replace the per-operation orTimeout with a cheaper timeout mechanism

1 participant