Skip to content

Restore zero-context-length guards in the SM90 paged MQA scheduler - #8

Merged
zyongye merged 1 commit into
devfrom
fix-sm90-paged-mqa-zero-context-guards
Sep 15, 2026
Merged

zyongye merged 1 commit into
devfrom
fix-sm90-paged-mqa-zero-context-guards

Conversation

@zyongye

@zyongye zyongye commented Sep 15, 2026

Copy link
Copy Markdown
Member

Problem

9a86ae2b ("SM90 paged MQA logits: support kv_block=32 and next_n=4", #7) re-ported the SM90 paged-MQA work from deepseek-ai/DeepGEMM's nv_dev branch onto the 26/09 release base. nv_dev carried four guards for zero context lengths that the 26/09 base never had, and the re-port dropped all four.

With every context length zero, sum == 0, so the binary search in sm90_paged_mqa_logits_metadata finds no prefix sum greater than seg_starts and returns lo == batch_size. Used unclamped as q_idx, that reads prefix_sum[batch_size] — one past the end of a shared buffer sized to exactly align(batch_size, 32) ints (num_smem_ints = aligned_batch_size, no slack). When batch_size is already 32-aligned the read is out of range and the kernel traps:

Detected an exception of type CUDBG_EXCEPTION_WARP_OUT_OF_RANGE_ADDRESS (5)
#0  deep_gemm::sched::sm90_paged_mqa_logits_metadata<1024u, 256u, 132u, false>

vLLM hits this on every DeepSeek V4 startup on Hopper: the DSA indexer warmup runs a dummy batch of max_num_seqs (default 1024, and align(1024, 32) == 1024) rows with all context lengths zero. It took down nvidia-h100-lm-eval-kv-offload-large deterministically — 4/4 ranks, both tests — on the vLLM PR that repins to this fork.

Fix

Restore the four guards, leaving the scheduler behaviorally identical to the nv_dev version it was ported from:

  • total == 0 early-returns with one-past-the-end sentinels — avoids the prefix_sum[batch_size] read and keeps a trailing SM off a synthetic zero-KV task.
  • q_idx is clamped to batch_size - 1 before indexing prefix_sum.
  • The scheduler constructor checks exist_q_atom_idx before refresh_num_kv_and_advance dereferences context_lens/indices, and initializes current_advance/current_num_kv/last_advance for the empty range.
  • fetch_next_task skips a zero-context request's atoms rather than exposing an empty task, covering traversal that crosses one between two non-empty requests.

I diffed the result against the nv_dev scheduler at 8b1392b (the pin vLLM is moving off): identical apart from comments.

Test

test_paged_mqa_logits draws context lengths around a positive average, so it never produces an empty request — which is why this regressed silently. Added test_paged_mqa_logits_zero_context covering all-zero and interleaved-zero context lengths over block_kv 32/64, next_n 1/2/4, and batch sizes 32 and 1024 (both 32-aligned, so prefix_sum[batch_size] lands exactly one past the end).

Verified on 8x H200 (SM90), driving the kernel through vLLM's vendored deep_gemm:

scheduler header new test vLLM tests/kernels/attention/test_deepgemm_attention.py + tests/v1/attention/test_indexer_native_next_n.py
9a86ae2b (unfixed) CUDA_ERROR_ILLEGAL_ADDRESS —
this PR passes 14 passed

The reproducer is the CI shape exactly: batch_size=1024, next_n=1, block_kv=64, all context_lens zero.

AI assistance

AI assistance (Claude Code) was used for this change; the submitter has reviewed every changed line.

🤖 Generated with Claude Code

9a86ae2 re-ported the SM90 paged MQA kv_block=32/next_n=4 support from
deepseek-ai/DeepGEMM's nv_dev branch onto the 26/09 release base. nv_dev
carried four guards for zero context lengths that the 26/09 base never
had, and the re-port dropped all of them.

With every context length zero, `sum` is zero, so the binary search in
`sm90_paged_mqa_logits_metadata` finds no prefix sum greater than
`seg_starts` and returns `lo == batch_size`. Using that unclamped as
`q_idx` reads `prefix_sum[batch_size]`, which is one past the end of a
shared buffer sized to exactly `align(batch_size, 32)` ints. When
`batch_size` is already 32-aligned the read is out of range and the
kernel traps:

    CUDBG_EXCEPTION_WARP_OUT_OF_RANGE_ADDRESS (5)
    #0 deep_gemm::sched::sm90_paged_mqa_logits_metadata<1024u, 256u, 132u, false>

vLLM hits this on every DeepSeek V4 startup: the DSA indexer warmup runs
a dummy batch of max_num_seqs (default 1024) rows with all context
lengths zero.

Restore the four guards:

  - `total == 0` early-returns with one-past-the-end sentinels, avoiding
    the `prefix_sum[batch_size]` read and keeping a trailing SM off a
    synthetic zero-KV task.
  - `q_idx` is clamped to `batch_size - 1` before indexing prefix_sum.
  - The scheduler constructor checks `exist_q_atom_idx` before
    `refresh_num_kv_and_advance` dereferences context_lens/indices, and
    initializes current_advance/current_num_kv/last_advance for the
    empty range.
  - `fetch_next_task` skips a zero-context request's atoms rather than
    exposing an empty task, for the case where traversal crosses one
    between two non-empty requests.

The scheduler is now behaviorally identical to the nv_dev version it was
ported from.

tests/test_attention.py: `test_paged_mqa_logits` draws context lengths
around a positive average and so never produces an empty request, which
is why this regressed silently. Add `test_paged_mqa_logits_zero_context`
covering all-zero and interleaved-zero context lengths over block_kv
32/64, next_n 1/2/4, and batch sizes 32 and 1024 (both 32-aligned, so
`prefix_sum[batch_size]` is exactly one past the end).

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@zyongye
zyongye merged commit ad1f172 into dev Sep 15, 2026
1 of 2 checks passed
zyongye added a commit to zyongye/vllm that referenced this pull request Sep 15, 2026
The fork commit this was first pinned to, 9a86ae2b, re-ported the SM90
paged-MQA kv_block=32/next_n=4 support onto the 2.8.0 base but dropped
the zero-context-length guards that deepseek-ai's nv_dev branch carried.
With all context lengths zero the metadata kernel reads
prefix_sum[batch_size], one past the end of a shared buffer sized to
exactly align(batch_size, 32) ints, and traps with
CUDBG_EXCEPTION_WARP_OUT_OF_RANGE_ADDRESS.

The DSA indexer warmup runs a dummy batch of max_num_seqs (default 1024,
already 32-aligned) rows with all context lengths zero, so every
DeepSeek V4 startup on Hopper hit it: nvidia-h100-lm-eval-kv-offload-large
failed deterministically on all 4 ranks of both
test_gsm8k_offloading_correctness cases.

Move the pin to the dev tip, which now carries vllm-project/DeepGEMM#8
restoring the four guards and adding a zero-context regression test.

Signed-off-by: Yongye Zhu <zyy1102000@gmail.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant