Repository navigation
Async serving: ROUTER socket, predict thread and latest-wins request queue - #34
Open
khanhnd61-vr wants to merge 1 commit into
Open
khanhnd61-vr wants to merge 1 commit into
khanhnd61-vr wants to merge 1 commit into
Conversation
…latest-wins request queue
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
vla-serverwas a single-threaded REQ/REP loop: one request, one reply, and the socket was deaf while the model ran. A robot client therefore paused for a full round trip per chunk. This PR makes the server asynchronous without changing anything for existing clients.vla::predict. Replies travel back over an inproc PUSH/PULL and carry the peer's routing envelope. A REQ client sees exactly the exchange it did before; a DEALER client can keep several requests in flight.--queue latest|fifo(src/serving/request_queue.h, header-only).latest(default) keeps one pending request per client and answers a superseded one at once witherror="superseded", so the model always works on each robot's freshest observation and several robots share a server fairly.fifoserves every request in order for benchmarks.latency_ms_queueadded toPredictResponse: how long a request waited for the predict thread. Malformed requests are rejected on the socket thread, so errors come back within milliseconds mid-predict.USAGE.md,ARCHITECTURE.md, README rollout section), CHANGELOG entry, andtests/test_request_queue.cpp.Companion PRs: the lerobot fork gains
lerobot-vla-cpp --mode=async(khanhnd61-vr/lerobot), and vla.simd gets its session lock out of the predict path (cair-vinuni/vla.simd).Test plan
test_request_queue(ctest): coalescing per client, no cross-client coalescing, FIFO order and depth cap, drain-then-stop, blocking pop.--queue fifoserves all five in order with queue waits growing by one predict each; SIGINT shuts down cleanly with the served/superseded summary.lerobot-vla-cppasync client against this server at 30 fps for 8 s: 229 of 239 ticks executed an action, the 10 starved ticks were the initial fill, 0.05 ms median control-thread work per tick.