Conversation
| @@ -0,0 +1,226 @@ | |||
| --- | |||
| layout: post | |||
There was a problem hiding this comment.
lets rebase to pick up the date of admission
|
|
||
| ## Acknowledgments | ||
|
|
||
| We thank Nicolò Lucchesi (Mistral AI), Mark McLoughlin (Red Hat), and the vLLM community for their continued support in reviewing and merging this feature. |
There was a problem hiding this comment.
| We thank Nicolò Lucchesi (Mistral AI), Mark McLoughlin (Red Hat), and the vLLM community for their continued support in reviewing and merging this feature. | |
| We thank Nicolò Lucchesi (Mistral), Mark McLoughlin (Red Hat), and the vLLM community for their continued support in reviewing and merging this feature. |
Just Mistral after rebranding :)
|
|
||
| Multi-turn agentic workloads (or multi-turn conversations) are sequential dialogue exchanges between users and large language model (LLM) systems that maintain contextual coherence across interaction cycles. Implementing them requires persistent conversation-history management, where each subsequent user query is concatenated with the relevant historical context: system prompts, prior user queries, and model-generated responses. | ||
|
|
||
| In non-disaggregated deployments, the LLM inference server keeps sessions resident and reuses their KV cache across turns, which improves performance. Prefill-decode (P-D) disaggregated architectures, however, present a fundamental challenge on two fronts. First, the KV cache for model-generated responses from previous turns resides exclusively on the decode nodes and is inaccessible to the prefill nodes, so the current architecture forces prefill nodes to recompute KV projections for all response tokens on every turn, resulting in a substantial waste of AI-accelerator compute. Second, in a P-D disaggregated multi-turn system, the decode node maintains the session state while the prefill nodes do not, which raises the probability of cache eviction for active sessions on the prefill nodes. On a cache miss, the prefill node must then recompute KV projections for the entire conversational context (system prompts, historical user queries, and prior model responses). |
There was a problem hiding this comment.
| In non-disaggregated deployments, the LLM inference server keeps sessions resident and reuses their KV cache across turns, which improves performance. Prefill-decode (P-D) disaggregated architectures, however, present a fundamental challenge on two fronts. First, the KV cache for model-generated responses from previous turns resides exclusively on the decode nodes and is inaccessible to the prefill nodes, so the current architecture forces prefill nodes to recompute KV projections for all response tokens on every turn, resulting in a substantial waste of AI-accelerator compute. Second, in a P-D disaggregated multi-turn system, the decode node maintains the session state while the prefill nodes do not, which raises the probability of cache eviction for active sessions on the prefill nodes. On a cache miss, the prefill node must then recompute KV projections for the entire conversational context (system prompts, historical user queries, and prior model responses). | |
| In non-disaggregated deployments, the LLM inference server keeps sessions resident and reuses their KV cache across turns. Prefill-decode (P-D) disaggregated architectures, however, present a fundamental challenge on two fronts. First, the KV cache for model-generated responses from previous turns resides exclusively on the decode nodes and is inaccessible to the prefill nodes, so the current architecture forces prefill nodes to recompute KV projections for all response tokens on every turn, resulting in a substantial waste of AI-accelerator compute. Second, the decode node maintains the session state while the prefill nodes do not, which raises the probability of cache eviction for active sessions on the prefill nodes. On a cache miss, the prefill node must then recompute KV projections for the entire conversational context (system prompts, historical user queries, and prior model responses). |
to simplify paragraph slightly
There was a problem hiding this comment.
also what do you mean with "session state on D"?
There was a problem hiding this comment.
I mean in a cache aware router architectures, the router picks the same D node for a given session_id. whereas P nodes are generally picked based on load balancing algorithm,
|
|
||
| In this post we cover how bi-directional KV cache transfer between prefill and decode nodes optimizes KV cache utilization in P-D disaggregated deployments and cuts redundant prefill recomputation on multi-turn conversations, reducing time to first token by up to ~3x on long multi-turn prompts (see the [Performance data](#performance-data) section). We built the feature on AWS Trainium instance clusters with AWS Elastic Fabric Adapter (EFA, the low latency RDMA network protocol used across all servers in AWS) and upstreamed it to vLLM as an accelerator-agnostic feature. The results here are from an AWS p5en (GPU) cluster with EFA, showing the gains carry over to GPUs. | ||
|
|
||
| **Bi-directional KV transfer ([#32553](https://github.com/vllm-project/vllm/pull/32553))**: a mechanism that allows a decode instance to return previously computed KV to a prefill instance on subsequent turns of a conversation, avoiding recomputation of shared context on the prefill node, governed by a cost-based recompute threshold. The feature is implemented in `vllm/distributed/kv_transfer/kv_connector/v1/nixl/` and interoperates with vLLM's hybrid KV cache manager for sliding-window models. It does not modify the existing standard P->D transfer path. These are purely additive extensions that activate only when bi-directional KV transfer is enabled and the multi-turn reusable token count exceeds the user-set threshold. Available with `vllm>=v0.21.0`. |
There was a problem hiding this comment.
We can probably shorten this paragraph, I feel
a mechanism that allows a decode instance to return previously computed KV to a prefill instance on subsequent turns of a conversation, avoiding recomputation of shared context on the prefill node, governed by a cost-based recompute threshold
could be blended into the previous one
|
|
||
| <figure> | ||
| <img src="/assets/figures/2026-09-21-bidirectional-kvxfer-multiturn-agentic-workload/kv-connector-remote-blocks.png" alt="KV connector components on the prefill and decode nodes, with KV loads in both directions and a KV-cache-aware router above them." style="width: 100%;"> | ||
| <figcaption><em>The prefill node gains a path to consume remote KV blocks, mirroring how the decode node already consumes the prefill node's shared KV blocks. Each node keeps its own logic for computing the metadata of local and remote blocks, while both the P->D and D->P transfers use the same NIXL READ mechanism. Each direction loads only the blocks that are not already present locally.</em></figcaption> |
There was a problem hiding this comment.
I feel this could be an extra paragraph, and the desacription something simpler
|
|
||
| ## Bi-directional KV transfer | ||
|
|
||
| Bi-directional KV cache transfer between prefill and decode nodes mitigates both computational-redundancy scenarios described earlier: nodes load KV cache from each other, which removes unnecessary recomputation. Realizing this requires the KV to survive past the turn that produced it and to travel in the reverse direction, coordinated by a KV-cache-aware router. |
There was a problem hiding this comment.
this paragraph is also slightly redundant/not adding too much
|
|
||
| ### Block retention TTL | ||
|
|
||
| Normally a decode instance frees a request's blocks as soon as generation finishes. Under bi-directional mode it instead retains them for a configurable lifetime (`decoder_kv_blocks_ttl`, default 480 seconds) and publishes their location and expiry, so the conversation's KV, both the prompt prefix and the generated response, remains available for the next turn to claim. When that turn arrives, its prefill instance reads the retained blocks back from the decode instance (a decode-to-prefill transfer) and reuses them instead of recomputing the grown context. The retention window bounds how long the KV is held: a follow-up turn that arrives within it reuses the KV, while one that arrives later falls back to normal prefill. |
There was a problem hiding this comment.
I think long TTLs are bad when a P dies; perhaps we can say next developments might make this shorter with a lease-like mechanism like the existing one on P
… agentic workloads Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>
6d47f5a to
7cc4988
Compare
No description provided.