forked from halo-box/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 10
Pull requests: halo-box/strix-llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
cuda: ROCmFPx tensor types + Q4_0_ROCMFP4_FAST MMQ tiles on HIP (gfx1151)
CUDA
ggml
#40
opened Sep 10, 2026 by
baraxnaxgaming-commits
Loading…
DRAFT: ggml: optimize Flash Next inference paths
CUDA
ggml
testing
Vulkan
#39
opened Sep 10, 2026 by
gaetan-puleo
•
Draft
HIP : route Q6_K through MMQ up to ne11 = 1024 on RDNA 3.5
CUDA
ggml
#38
opened Sep 10, 2026 by
Nyovelt
Loading…
ggml : backport persistent view initialization after buffer-size splits
ggml
testing
#37
opened Sep 8, 2026 by
arc-uri-el
Loading…
server : preserve RAM-cached branches sharing a large prompt prefix
server
#36
opened Sep 8, 2026 by
arc-uri-el
Loading…
context : preserve quantized block sizes in device state I/O
testing
#35
opened Sep 8, 2026 by
arc-uri-el
Loading…
Rocmfpx/vulkan 2026 09 07
ggml
testing
Vulkan
#32
opened Sep 7, 2026 by
LaurentZuijdwijk
Member
Loading…
Fix/rocmfpx encoder review
ggml
testing
Vulkan
#31
opened Sep 7, 2026 by
LaurentZuijdwijk
Member
Loading…
vulkan: fix ROCmFPx mat-vec cost at batch 3-8 on rocmfpx/wholesale-reference
#7
opened Aug 29, 2026 by
LaurentZuijdwijk
Member
•
Draft
ProTip!
Type g i on any issue or pull request to go back to the issue listing page.