[Blog] Taking vLLM Apart: A Practical Guide to Disaggregated Serving - #347
Conversation
Adds a hands on guide to running disaggregated serving in vLLM v0.31.0+. It covers splitting prefill and decode with KV connectors, moving tokenization and parsing off the GPU with the new /render and /derender frontend and how the pieces fit together. It also looks at who's running this in production and what's still neede to be done. Includes pipeline, multi-turn, transfer cost and ITL tail figures. Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
7514cd6 to
ae8f189
Compare
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
|
You should reword "Tail latency gets boring," it screams AI prose. Also I feel that Fig. 2 would be better expressed as a line graph. |
Review comment: - vllm-project#347 (comment)
Thanks @DarkLight1337 for the review, updated and ready for review again. |
Review comment: - vllm-project#347 (comment) Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
c37a0e7 to
fb01e80
Compare
NickLucche
left a comment
There was a problem hiding this comment.
Thanks @hickeyma ! I think we can shorted the PD part a bit and focus more on the gpu-less FE :)
…post
- Reframe the throughput claim around goodput instead of quoting the
docs line that says P/D doesn't help throughput
- Lead with the MoRI-IO and llm-d results and cut the L40S run down
to a short caveat about slow transfers
- Shrink the multi-turn section to a summary plus a link to the
bidirectional KV transfer post, keeping the reasoning-model warning
- Mention KV offloading/shared KV and fabric checks in the table
- Drop the redundant kv_load_failure_policy from the examples
- Note the DP=2 baseline isn't the strongest collocated setup
- Remove the two figures that are no longer used
Updates based on review comment:
- vllm-project#347 (review)
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Explain why tokens in/tokens out matters beyond saving CPU: routing on real token IDs, exact tokens for RL and eval and multimodal preprocessing on the render tier. Add a streaming derender example and a short section on deploying and sizing the render tier. The renderer_num_workers advice moves up there from "What's still left". Review comment: - vllm-project#347 (review) Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Thanks @NickLucche for the review and feedback. Trimmed the P/D part. Multi-turn is now a short summary pointing to #345 and the L40S benchmark is just a caveat. The frontend section got updated with: why tokens in/tokens out matters beyond CPU, a streaming derender example and how to deploy and size the render tier. Let me know what you think. |
Explain why tokens in/tokens out matters beyond saving CPU: routing on real token IDs, exact tokens for RL and eval and multimodal preprocessing on the render tier. Add a streaming derender example and a short section on deploying and sizing the render tier. The renderer_num_workers advice moves up there from "What's still left". Review comment: - vllm-project#347 (review) Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
182b474 to
9b1ec97
Compare
Review comments: - vllm-project#347 (review) Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Removed unnecessary dividing line from first section and updated the final sentnece in the opening to read better. Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
30568f5 to
4cb1c63
Compare
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Adds a hands on guide to disaggregated serving in vLLM.
It walks through splitting prefill and decode onto separate instances with a KV connector and moving tokenization, detokenization and tool/reasoning parsing onto a CPU-only frontend using the new
/renderand/derenderendpoints. It also shows how the two fit together, including prefill reusing conversation state from decode in multi-turn chat and covers who's running this in production and what's still missing.
Includes four figures: the serving pipeline, multi-turn flow, KV transfer cost and ITL tail latency.
cc @NickLucche @sagearc @DarkLight1337 @chaunceyjiang @njhill @vMaroon @hyeongyun0916