Skip to content

Async serving: ROUTER socket, predict thread and latest-wins request queue - #34

Open
khanhnd61-vr wants to merge 1 commit into
mainfrom
feat/async-inference
Open

khanhnd61-vr wants to merge 1 commit into
mainfrom
feat/async-inference

Conversation

@khanhnd61-vr

@khanhnd61-vr khanhnd61-vr commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

vla-server was a single-threaded REQ/REP loop: one request, one reply, and the socket was deaf while the model ran. A robot client therefore paused for a full round trip per chunk. This PR makes the server asynchronous without changing anything for existing clients.

  • ROUTER socket + predict thread. The socket thread receives, validates and decodes; the worker is the only caller of vla::predict. Replies travel back over an inproc PUSH/PULL and carry the peer's routing envelope. A REQ client sees exactly the exchange it did before; a DEALER client can keep several requests in flight.
  • --queue latest|fifo (src/serving/request_queue.h, header-only). latest (default) keeps one pending request per client and answers a superseded one at once with error="superseded", so the model always works on each robot's freshest observation and several robots share a server fairly. fifo serves every request in order for benchmarks.
  • latency_ms_queue added to PredictResponse: how long a request waited for the predict thread. Malformed requests are rejected on the socket thread, so errors come back within milliseconds mid-predict.
  • Docs (USAGE.md, ARCHITECTURE.md, README rollout section), CHANGELOG entry, and tests/test_request_queue.cpp.

Companion PRs: the lerobot fork gains lerobot-vla-cpp --mode=async (khanhnd61-vr/lerobot), and vla.simd gets its session lock out of the predict path (cair-vinuni/vla.simd).

Test plan

  • test_request_queue (ctest): coalescing per client, no cross-client coalescing, FIFO order and depth cap, drain-then-stop, blocking pop.
  • End to end on an RTX 3060 with the SmolVLA LIBERO checkpoint: legacy REQ client unchanged; a 5-request DEALER burst serves the first and the newest (queue wait = one predict) and answers the three stale ones in 26 ms; two DEALER clients both served; a malformed request answered in 6 ms during a 147 ms predict; REQ and DEALER interleaved; --queue fifo serves all five in order with queue waits growing by one predict each; SIGINT shuts down cleanly with the served/superseded summary.
  • lerobot-vla-cpp async client against this server at 30 fps for 8 s: 229 of 239 ticks executed an action, the 10 starved ticks were the initial fill, 0.05 ms median control-thread work per tick.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant