forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 7
Pull requests: z-lab/llama.cpp-fork
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
speculative: keep DFlash draft positions contiguous across mtmd vision chunks
#5
opened Aug 23, 2026 by
wangcc57
Loading…
cuda: add Blackwell (SM120) MMVQ parameter table and tune FATTN config for Blackwell
#3
opened Aug 22, 2026 by
sunagent
Loading…
common: force draft model to SPLIT_MODE_NONE when target uses SPLIT_MODE_TENSOR
#2
opened Aug 19, 2026 by
sunagent
Loading…
dflash: zero-fill draft-cache holes left by mtmd chunks and reused prefixes
#1
opened Aug 19, 2026 by
dagnarf
Loading…
ProTip!
Add no:assignee to see everything that’s not assigned.