Skip to content

[Blog] Taking vLLM Apart: A Practical Guide to Disaggregated Serving - #347

Merged
NickLucche merged 9 commits into
vllm-project:mainfrom
hickeyma:add-disagg-serv-blog
Sep 30, 2026
Merged

NickLucche merged 9 commits into
vllm-project:mainfrom
hickeyma:add-disagg-serv-blog

Conversation

@hickeyma

@hickeyma hickeyma commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Adds a hands on guide to disaggregated serving in vLLM.

It walks through splitting prefill and decode onto separate instances with a KV connector and moving tokenization, detokenization and tool/reasoning parsing onto a CPU-only frontend using the new /render and /derender
endpoints. It also shows how the two fit together, including prefill reusing conversation state from decode in multi-turn chat and covers who's running this in production and what's still missing.

Includes four figures: the serving pipeline, multi-turn flow, KV transfer cost and ITL tail latency.

cc @NickLucche @sagearc @DarkLight1337 @chaunceyjiang @njhill @vMaroon @hyeongyun0916

Adds a hands on guide to running disaggregated serving in vLLM v0.31.0+.
It covers splitting prefill and decode with KV connectors, moving
tokenization and parsing off the GPU with the new /render and /derender
frontend and how the pieces fit together. It also looks at who's running
this in production and what's still neede to be done.

Includes pipeline, multi-turn, transfer cost and ITL tail figures.

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>

@chaunceyjiang chaunceyjiang left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks~ LGTM.

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
@DarkLight1337

Copy link
Copy Markdown
Member

You should reword "Tail latency gets boring," it screams AI prose. Also I feel that Fig. 2 would be better expressed as a line graph.

hickeyma added a commit to hickeyma/vllm-project.github.io that referenced this pull request Sep 24, 2026
@hickeyma

hickeyma commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author

You should reword "Tail latency gets boring," it screams AI prose. Also I feel that Fig. 2 would be better expressed as a line graph.

Thanks @DarkLight1337 for the review, updated and ready for review again.

Review comment:
- vllm-project#347 (comment)

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>

@NickLucche NickLucche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @hickeyma ! I think we can shorted the PD part a bit and focus more on the gpu-less FE :)

Comment thread _posts/2026-09-20-disaggregated-serving-guide.md Outdated
Comment thread _posts/2026-09-20-disaggregated-serving-guide.md Outdated
Comment thread _posts/2026-09-20-disaggregated-serving-guide.md Outdated
Comment thread _posts/2026-09-20-disaggregated-serving-guide.md Outdated
Comment thread _posts/2026-09-20-disaggregated-serving-guide.md Outdated
Comment thread _posts/2026-09-20-disaggregated-serving-guide.md Outdated
Comment thread _posts/2026-09-20-disaggregated-serving-guide.md Outdated
Comment thread _posts/2026-09-20-disaggregated-serving-guide.md Outdated
Comment thread _posts/2026-09-20-disaggregated-serving-guide.md Outdated
…post

  - Reframe the throughput claim around goodput instead of quoting the
    docs line that says P/D doesn't help throughput
  - Lead with the MoRI-IO and llm-d results and cut the L40S run down
    to a short caveat about slow transfers
  - Shrink the multi-turn section to a summary plus a link to the
    bidirectional KV transfer post, keeping the reasoning-model warning
  - Mention KV offloading/shared KV and fabric checks in the table
  - Drop the redundant kv_load_failure_policy from the examples
  - Note the DP=2 baseline isn't the strongest collocated setup
  - Remove the two figures that are no longer used

Updates based on review comment:

- vllm-project#347 (review)

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
hickeyma added a commit to hickeyma/vllm-project.github.io that referenced this pull request Sep 28, 2026
Explain why tokens in/tokens out matters beyond saving CPU: routing
on real token IDs, exact tokens for RL and eval and multimodal
preprocessing on the render tier. Add a streaming derender example
and a short section on deploying and sizing the render tier. The
renderer_num_workers advice moves up there from "What's still left".

Review comment:
- vllm-project#347 (review)

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
@hickeyma

Copy link
Copy Markdown
Contributor Author

Thanks @hickeyma ! I think we can shorted the PD part a bit and focus more on the gpu-less FE :)

Thanks @NickLucche for the review and feedback. Trimmed the P/D part. Multi-turn is now a short summary pointing to #345 and the L40S benchmark is just a caveat. The frontend section got updated with: why tokens in/tokens out matters beyond CPU, a streaming derender example and how to deploy and size the render tier.

Let me know what you think.

@hickeyma
hickeyma requested a review from NickLucche September 28, 2026 10:19
Explain why tokens in/tokens out matters beyond saving CPU: routing
on real token IDs, exact tokens for RL and eval and multimodal
preprocessing on the render tier. Add a streaming derender example
and a short section on deploying and sizing the render tier. The
renderer_num_workers advice moves up there from "What's still left".

Review comment:
- vllm-project#347 (review)

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>

@sagearc sagearc left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @hickeyma, very nice!

Comment thread _posts/2026-09-20-disaggregated-serving-guide.md Outdated
Comment thread _posts/2026-09-29-disaggregated-serving-guide.md
Review comments:
- vllm-project#347 (review)

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>

@NickLucche NickLucche left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Removed unnecessary dividing line from first section and
updated the final sentnece in the opening to read better.

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
@NickLucche
NickLucche merged commit f9792a2 into vllm-project:main Sep 30, 2026
2 checks passed
@hickeyma
hickeyma deleted the add-disagg-serv-blog branch September 30, 2026 10:55

This branch was successfully deployed

1 active deployment
Preview — be1f4f9b Deployed Sep 30, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants