Skip to content

Pull requests: ggml-org/llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

vulkan: large top-k support via argsort ggml changes relating to the ggml tensor library for machine learning testing Everything test related Vulkan Issues specific to the Vulkan backend
#28005 opened Aug 30, 2026 by antoinezambelli Draft
ggml-cuda: optimize single-token MMVQ dispatch for RDNA3 architecture CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#28003 opened Aug 30, 2026 by maci0 Draft
dflash: pass missing NVFP4 scales to attention operations merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. model Model specific
#28000 opened Aug 30, 2026 by JamePeng Contributor Loading…
Add sm70 FlashAttention config case for DKQ=256, DV=256, ncols=64 CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#27997 opened Aug 30, 2026 by mistrjirka Draft
kv cache : optimize restoring non-contiguous cells testing Everything test related
#27991 opened Aug 29, 2026 by itsnotoger Loading…
ggml-cpu : add mirror NUMA strategy (replicate weights on each node) CUDA Related to the CUDA backend documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning server
#27986 opened Aug 29, 2026 by matteoscalabrini Loading…
On ggml-cpu: tweak ARM_NATIVE_FLAG native baseline arch when extensions need it ggml changes relating to the ggml tensor library for machine learning
#27984 opened Aug 29, 2026 by cameronelliott Loading…
quantize: add IQ2_NL and IQ3_NL types (CPU + Metal + CUDA + Vulkan) Apple Metal https://en.wikipedia.org/wiki/Metal_(API) conversion CUDA Related to the CUDA backend examples ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language testing Everything test related Vulkan Issues specific to the Vulkan backend
#27983 opened Aug 29, 2026 by EAddario Contributor Draft
CUDA: let any expert count use the fast mm_ids_helper path CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#27978 opened Aug 29, 2026 by ServeurpersoCom Contributor Loading…
qwen4exp: reduce the generation slowdown as context grows model Model specific
#27977 opened Aug 29, 2026 by ServeurpersoCom Contributor Loading…
vulkan: fuse GATED_DELTA_NET state write into recurrent cache ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend
#27973 opened Aug 29, 2026 by PrajwalMukatti Draft
CUDA + ggml: add sparse-fa for DSV4/GLM CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning model Model specific testing Everything test related
#27970 opened Aug 29, 2026 by am17an Contributor Loading…
[SYCL] Enhance to get the free memory of Intel GPU documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language
#27968 opened Aug 29, 2026 by arthw Contributor Loading…
HIP : optimize IQ2/IQ3 (__vsub4 __vcmpne4) using SWAR CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#27962 opened Aug 29, 2026 by yanjs Loading…
ggml-cpu : conditionally add SpacemiT IME kernel sources ggml changes relating to the ggml tensor library for machine learning
#27961 opened Aug 29, 2026 by alanhc Draft
1 task done
ui : add model download pipeline server/ui
#27959 opened Aug 29, 2026 by allozaur Contributor 5/5 Draft
ui : add model compatibility estimation server/ui
#27957 opened Aug 29, 2026 by allozaur Contributor 4/5 Draft
CUDA : fix divergent FlashAttention barrier CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#27955 opened Aug 29, 2026 by mistrjirka Loading…
vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend
#27952 opened Aug 29, 2026 by 0cc4m Contributor Loading…
ui : Add Hugging Face Hub data layer server/ui
#27947 opened Aug 29, 2026 by allozaur Contributor 3/5 Loading…
ProTip! Updated in the last three days: updated:>2026-08-27.