Restore zero-context-length guards in the SM90 paged MQA scheduler - #8
Merged
Merged
Conversation
9a86ae2 re-ported the SM90 paged MQA kv_block=32/next_n=4 support from deepseek-ai/DeepGEMM's nv_dev branch onto the 26/09 release base. nv_dev carried four guards for zero context lengths that the 26/09 base never had, and the re-port dropped all of them. With every context length zero, `sum` is zero, so the binary search in `sm90_paged_mqa_logits_metadata` finds no prefix sum greater than `seg_starts` and returns `lo == batch_size`. Using that unclamped as `q_idx` reads `prefix_sum[batch_size]`, which is one past the end of a shared buffer sized to exactly `align(batch_size, 32)` ints. When `batch_size` is already 32-aligned the read is out of range and the kernel traps: CUDBG_EXCEPTION_WARP_OUT_OF_RANGE_ADDRESS (5) #0 deep_gemm::sched::sm90_paged_mqa_logits_metadata<1024u, 256u, 132u, false> vLLM hits this on every DeepSeek V4 startup: the DSA indexer warmup runs a dummy batch of max_num_seqs (default 1024) rows with all context lengths zero. Restore the four guards: - `total == 0` early-returns with one-past-the-end sentinels, avoiding the `prefix_sum[batch_size]` read and keeping a trailing SM off a synthetic zero-KV task. - `q_idx` is clamped to `batch_size - 1` before indexing prefix_sum. - The scheduler constructor checks `exist_q_atom_idx` before `refresh_num_kv_and_advance` dereferences context_lens/indices, and initializes current_advance/current_num_kv/last_advance for the empty range. - `fetch_next_task` skips a zero-context request's atoms rather than exposing an empty task, for the case where traversal crosses one between two non-empty requests. The scheduler is now behaviorally identical to the nv_dev version it was ported from. tests/test_attention.py: `test_paged_mqa_logits` draws context lengths around a positive average and so never produces an empty request, which is why this regressed silently. Add `test_paged_mqa_logits_zero_context` covering all-zero and interleaved-zero context lengths over block_kv 32/64, next_n 1/2/4, and batch sizes 32 and 1024 (both 32-aligned, so `prefix_sum[batch_size]` is exactly one past the end). Signed-off-by: Yongye Zhu <zyy1102000@gmail.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
zyongye
added a commit
to zyongye/vllm
that referenced
this pull request
Sep 15, 2026
The fork commit this was first pinned to, 9a86ae2b, re-ported the SM90 paged-MQA kv_block=32/next_n=4 support onto the 2.8.0 base but dropped the zero-context-length guards that deepseek-ai's nv_dev branch carried. With all context lengths zero the metadata kernel reads prefix_sum[batch_size], one past the end of a shared buffer sized to exactly align(batch_size, 32) ints, and traps with CUDBG_EXCEPTION_WARP_OUT_OF_RANGE_ADDRESS. The DSA indexer warmup runs a dummy batch of max_num_seqs (default 1024, already 32-aligned) rows with all context lengths zero, so every DeepSeek V4 startup on Hopper hit it: nvidia-h100-lm-eval-kv-offload-large failed deterministically on all 4 ranks of both test_gsm8k_offloading_correctness cases. Move the pin to the dev tip, which now carries vllm-project/DeepGEMM#8 restoring the four guards and adding a zero-context regression test. Signed-off-by: Yongye Zhu <zyy1102000@gmail.com> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
9a86ae2b("SM90 paged MQA logits: support kv_block=32 and next_n=4", #7) re-ported the SM90 paged-MQA work fromdeepseek-ai/DeepGEMM'snv_devbranch onto the 26/09 release base.nv_devcarried four guards for zero context lengths that the 26/09 base never had, and the re-port dropped all four.With every context length zero,
sum == 0, so the binary search insm90_paged_mqa_logits_metadatafinds no prefix sum greater thanseg_startsand returnslo == batch_size. Used unclamped asq_idx, that readsprefix_sum[batch_size]— one past the end of a shared buffer sized to exactlyalign(batch_size, 32)ints (num_smem_ints = aligned_batch_size, no slack). Whenbatch_sizeis already 32-aligned the read is out of range and the kernel traps:vLLM hits this on every DeepSeek V4 startup on Hopper: the DSA indexer warmup runs a dummy batch of
max_num_seqs(default 1024, andalign(1024, 32) == 1024) rows with all context lengths zero. It took downnvidia-h100-lm-eval-kv-offload-largedeterministically — 4/4 ranks, both tests — on the vLLM PR that repins to this fork.Fix
Restore the four guards, leaving the scheduler behaviorally identical to the
nv_devversion it was ported from:total == 0early-returns with one-past-the-end sentinels — avoids theprefix_sum[batch_size]read and keeps a trailing SM off a synthetic zero-KV task.q_idxis clamped tobatch_size - 1before indexingprefix_sum.exist_q_atom_idxbeforerefresh_num_kv_and_advancedereferencescontext_lens/indices, and initializescurrent_advance/current_num_kv/last_advancefor the empty range.fetch_next_taskskips a zero-context request's atoms rather than exposing an empty task, covering traversal that crosses one between two non-empty requests.I diffed the result against the
nv_devscheduler at8b1392b(the pin vLLM is moving off): identical apart from comments.Test
test_paged_mqa_logitsdraws context lengths around a positive average, so it never produces an empty request — which is why this regressed silently. Addedtest_paged_mqa_logits_zero_contextcovering all-zero and interleaved-zero context lengths overblock_kv32/64,next_n1/2/4, and batch sizes 32 and 1024 (both 32-aligned, soprefix_sum[batch_size]lands exactly one past the end).Verified on 8x H200 (SM90), driving the kernel through vLLM's vendored
deep_gemm:tests/kernels/attention/test_deepgemm_attention.py+tests/v1/attention/test_indexer_native_next_n.py9a86ae2b(unfixed)CUDA_ERROR_ILLEGAL_ADDRESSThe reproducer is the CI shape exactly:
batch_size=1024,next_n=1,block_kv=64, allcontext_lenszero.AI assistance
AI assistance (Claude Code) was used for this change; the submitter has reviewed every changed line.
🤖 Generated with Claude Code