Skip to content

ggml-cuda: optimize single-token MMVQ dispatch for RDNA3 architecture - #28003

Draft
maci0 wants to merge 1 commit into
ggml-org:masterfrom
maci0:perf/rdna3-mmvq-fastpath
Draft

ggml-cuda: optimize single-token MMVQ dispatch for RDNA3 architecture#28003
maci0 wants to merge 1 commit into
ggml-org:masterfrom
maci0:perf/rdna3-mmvq-fastpath