Skip to content

[P/D][Blog]: Add a blog on bi-directional KV transfers for multi-turn agentic workloads - #345

Open
snadampal wants to merge 1 commit into
vllm-project:mainfrom
snadampal:bidirection_kvxfer_multiturn_workload
Open

snadampal wants to merge 1 commit into
vllm-project:mainfrom
snadampal:bidirection_kvxfer_multiturn_workload

Conversation

@snadampal

Copy link
Copy Markdown

No description provided.

@@ -0,0 +1,226 @@
---
layout: post

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lets rebase to pick up the date of admission


## Acknowledgments

We thank Nicolò Lucchesi (Mistral AI), Mark McLoughlin (Red Hat), and the vLLM community for their continued support in reviewing and merging this feature.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
We thank Nicolò Lucchesi (Mistral AI), Mark McLoughlin (Red Hat), and the vLLM community for their continued support in reviewing and merging this feature.
We thank Nicolò Lucchesi (Mistral), Mark McLoughlin (Red Hat), and the vLLM community for their continued support in reviewing and merging this feature.

Just Mistral after rebranding :)


Multi-turn agentic workloads (or multi-turn conversations) are sequential dialogue exchanges between users and large language model (LLM) systems that maintain contextual coherence across interaction cycles. Implementing them requires persistent conversation-history management, where each subsequent user query is concatenated with the relevant historical context: system prompts, prior user queries, and model-generated responses.

In non-disaggregated deployments, the LLM inference server keeps sessions resident and reuses their KV cache across turns, which improves performance. Prefill-decode (P-D) disaggregated architectures, however, present a fundamental challenge on two fronts. First, the KV cache for model-generated responses from previous turns resides exclusively on the decode nodes and is inaccessible to the prefill nodes, so the current architecture forces prefill nodes to recompute KV projections for all response tokens on every turn, resulting in a substantial waste of AI-accelerator compute. Second, in a P-D disaggregated multi-turn system, the decode node maintains the session state while the prefill nodes do not, which raises the probability of cache eviction for active sessions on the prefill nodes. On a cache miss, the prefill node must then recompute KV projections for the entire conversational context (system prompts, historical user queries, and prior model responses).

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
In non-disaggregated deployments, the LLM inference server keeps sessions resident and reuses their KV cache across turns, which improves performance. Prefill-decode (P-D) disaggregated architectures, however, present a fundamental challenge on two fronts. First, the KV cache for model-generated responses from previous turns resides exclusively on the decode nodes and is inaccessible to the prefill nodes, so the current architecture forces prefill nodes to recompute KV projections for all response tokens on every turn, resulting in a substantial waste of AI-accelerator compute. Second, in a P-D disaggregated multi-turn system, the decode node maintains the session state while the prefill nodes do not, which raises the probability of cache eviction for active sessions on the prefill nodes. On a cache miss, the prefill node must then recompute KV projections for the entire conversational context (system prompts, historical user queries, and prior model responses).
In non-disaggregated deployments, the LLM inference server keeps sessions resident and reuses their KV cache across turns. Prefill-decode (P-D) disaggregated architectures, however, present a fundamental challenge on two fronts. First, the KV cache for model-generated responses from previous turns resides exclusively on the decode nodes and is inaccessible to the prefill nodes, so the current architecture forces prefill nodes to recompute KV projections for all response tokens on every turn, resulting in a substantial waste of AI-accelerator compute. Second, the decode node maintains the session state while the prefill nodes do not, which raises the probability of cache eviction for active sessions on the prefill nodes. On a cache miss, the prefill node must then recompute KV projections for the entire conversational context (system prompts, historical user queries, and prior model responses).

to simplify paragraph slightly

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

also what do you mean with "session state on D"?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I mean in a cache aware router architectures, the router picks the same D node for a given session_id. whereas P nodes are generally picked based on load balancing algorithm,


In this post we cover how bi-directional KV cache transfer between prefill and decode nodes optimizes KV cache utilization in P-D disaggregated deployments and cuts redundant prefill recomputation on multi-turn conversations, reducing time to first token by up to ~3x on long multi-turn prompts (see the [Performance data](#performance-data) section). We built the feature on AWS Trainium instance clusters with AWS Elastic Fabric Adapter (EFA, the low latency RDMA network protocol used across all servers in AWS) and upstreamed it to vLLM as an accelerator-agnostic feature. The results here are from an AWS p5en (GPU) cluster with EFA, showing the gains carry over to GPUs.

**Bi-directional KV transfer ([#32553](https://github.com/vllm-project/vllm/pull/32553))**: a mechanism that allows a decode instance to return previously computed KV to a prefill instance on subsequent turns of a conversation, avoiding recomputation of shared context on the prefill node, governed by a cost-based recompute threshold. The feature is implemented in `vllm/distributed/kv_transfer/kv_connector/v1/nixl/` and interoperates with vLLM's hybrid KV cache manager for sliding-window models. It does not modify the existing standard P->D transfer path. These are purely additive extensions that activate only when bi-directional KV transfer is enabled and the multi-turn reusable token count exceeds the user-set threshold. Available with `vllm>=v0.21.0`.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We can probably shorten this paragraph, I feel

a mechanism that allows a decode instance to return previously computed KV to a prefill instance on subsequent turns of a conversation, avoiding recomputation of shared context on the prefill node, governed by a cost-based recompute threshold

could be blended into the previous one


<figure>
<img src="/assets/figures/2026-09-21-bidirectional-kvxfer-multiturn-agentic-workload/kv-connector-remote-blocks.png" alt="KV connector components on the prefill and decode nodes, with KV loads in both directions and a KV-cache-aware router above them." style="width: 100%;">
<figcaption><em>The prefill node gains a path to consume remote KV blocks, mirroring how the decode node already consumes the prefill node's shared KV blocks. Each node keeps its own logic for computing the metadata of local and remote blocks, while both the P->D and D->P transfers use the same NIXL READ mechanism. Each direction loads only the blocks that are not already present locally.</em></figcaption>

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I feel this could be an extra paragraph, and the desacription something simpler


## Bi-directional KV transfer

Bi-directional KV cache transfer between prefill and decode nodes mitigates both computational-redundancy scenarios described earlier: nodes load KV cache from each other, which removes unnecessary recomputation. Realizing this requires the KV to survive past the turn that produced it and to travel in the reverse direction, coordinated by a KV-cache-aware router.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this paragraph is also slightly redundant/not adding too much


### Block retention TTL

Normally a decode instance frees a request's blocks as soon as generation finishes. Under bi-directional mode it instead retains them for a configurable lifetime (`decoder_kv_blocks_ttl`, default 480 seconds) and publishes their location and expiry, so the conversation's KV, both the prompt prefix and the generated response, remains available for the next turn to claim. When that turn arrives, its prefill instance reads the retained blocks back from the decode instance (a decode-to-prefill transfer) and reuses them instead of recomputing the grown context. The retention window bounds how long the KV is held: a follow-up turn that arrives within it reuses the KV, while one that arrives later falls back to normal prefill.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think long TTLs are bad when a P dies; perhaps we can say next developments might make this shorter with a lease-like mechanism like the existing one on P

… agentic workloads

Signed-off-by: Sunita Nadampalli <nadampal@amazon.com>

This branch was successfully deployed

1 active deployment
Preview — 7cc49887 Deployed Sep 29, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants