ggml-cuda: optimize single-token MMVQ dispatch for RDNA3 architecture - #28003
Draft
maci0 wants to merge 1 commit into
Draft
ggml-cuda: optimize single-token MMVQ dispatch for RDNA3 architecture#28003maci0 wants to merge 1 commit into
maci0 wants to merge 1 commit into
Conversation
|
Hi @maci0, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR 2: ggml-cuda: optimize single-token MMVQ dispatch for RDNA3 architecture
Target:
ggml-org/llama.cppFiles:
ggml/src/ggml-cuda/mmvq.cuMotivation
During single-token autoregressive generation (
ncols_dst == 1), execution is bound by kernel launch latency and per-wave register allocation.On AMD RDNA3 (gfx1100 / Navi 31), the generalized multi-batch launch parameter checks introduce extra branching in the launch prologue, causing Clang to allocate additional VGPRs that slightly degrade warp occupancy for single-token dispatches.
Changes
Added a direct fast-path dispatch for
ncols_dst == 1when targetingMMVQ_PARAMETERS_RDNA3_0without MoE IDs. This avoids multi-batch evaluation checks and issues the optimal wave32 single-token tile configuration directly.Benchmark Results
Tested on AMD Radeon RX 7900 XTX with
Qwen3.8-27B-Q4_K_M:Q4_KGEMV average duration: 68.91 µs