Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
193 commits
Select commit Hold shift + click to select a range
8f91482
spec(GFX1100-TG200): commit the 200 tok/s campaign spec
ghazni101 Aug 22, 2026
471ae60
measure(GFX1100-TG200): T1 re-prices the tip -- 40.65 tok/s median, b…
ghazni101 Aug 22, 2026
a203de4
measure(GFX1100-TG200): T2a splits the GdnPostConv symbol -- the grid…
ghazni101 Aug 22, 2026
a768304
research(GFX1100-TG200): rank vLLM/SGLang mechanisms against our meas…
ghazni101 Aug 22, 2026
683ae12
record(GFX1100-TG200): reject the pointer-keyed quant cache -- alloca…
ghazni101 Aug 23, 2026
ea63841
perf(GFX1100-TG200): merged keep-quant gate_up -- one quant GEMM per …
ghazni101 Aug 23, 2026
12660f2
perf(GFX1100-TG200): T2b flips ROCm support_static_graph_mode -- deco…
ghazni101 Aug 23, 2026
f56b24d
record(GFX1100-TG200): session-state note appended to t2b evidence (h…
ghazni101 Aug 23, 2026
38846c8
perf(GFX1100-TG200): T3a ports the f32-query decode-GQA attention arm…
ghazni101 Aug 23, 2026
965cd67
record(GFX1100-TG200): t3a evidence session-state note (hindsight 500s)
ghazni101 Aug 23, 2026
d2fbfc2
test(GFX1100-TG200): T4a lands the red-first ROCm quant-dot gate
ghazni101 Aug 23, 2026
3ae7b78
perf(GFX1100-TG200): T4a adds the VT_GEMV_MMVQ=1 K-quant decode GEMV …
ghazni101 Aug 23, 2026
c099310
perf(GFX1100-TG200): T4a folds activation quant into the MMVQ GEMV pr…
ghazni101 Aug 23, 2026
f2a3392
record(GFX1100-TG200): T4a evidence -- decode GEMV lever closed negat…
ghazni101 Aug 23, 2026
a5b9d69
fix(GFX1100-TG200): T4a repairs the MMVQ arm -- m-gates the whole dis…
ghazni101 Aug 23, 2026
4bb2a6d
record(GFX1100-TG200): T4a repair-round evidence -- lever adopted at …
ghazni101 Aug 23, 2026
fa9b546
test(GFX1100-TG200): T4a repair-2 adds host-side dispatch-route count…
ghazni101 Aug 23, 2026
7c063a7
perf(GFX1100-TG200): T4a lever-B1 makes the fold crossover tunable be…
ghazni101 Aug 23, 2026
f6ae041
record(GFX1100-TG200): T4a lever-B1 evidence -- fold-crossover re-tun…
ghazni101 Aug 23, 2026
662389d
record(GFX1100-TG200): T4a lever-B2 attributes all 21.6 Cijk calls/to…
ghazni101 Aug 23, 2026
40827c5
test(GFX1100-TG200): T4a lever-B2 adds the red-first f32-out decode-s…
ghazni101 Aug 23, 2026
cbaf98f
perf(GFX1100-TG200): T4a lever-B2 adds the VT_SKINNY_BF16=1 f32-out d…
ghazni101 Aug 23, 2026
54217b0
test(GFX1100-TG200): T4a lever-B2 restores the routing-witness case p…
ghazni101 Aug 24, 2026
21558db
test(GFX1100-TG200): T4a lever-B2 corrects the routing-witness expect…
ghazni101 Aug 24, 2026
632ecc5
record(GFX1100-TG200): T4a lever-B2 adopts the VT_SKINNY_BF16=1 f32-o…
ghazni101 Aug 24, 2026
84f6043
record(GFX1100-TG200): T4a lever-B2 closes the loop -- ON-arm capture…
ghazni101 Aug 24, 2026
ea9b8bd
test(GFX1100-TG200): T4a lever-B2 makes the TRUE-unset routing window…
ghazni101 Aug 24, 2026
4b6c49d
record(GFX1100-TG200): T4a lever-B2 re-pins the default-routing witne…
ghazni101 Aug 24, 2026
8d32955
attribution(GFX1100-TG200): lever-C maps all 97 QuantizeQ8KK decode l…
ghazni101 Aug 24, 2026
e78f276
test(GFX1100-TG200): lever-C adds red-first witnesses for the fused n…
ghazni101 Aug 24, 2026
0b02e01
feat(GFX1100-TG200): lever-C fuses Q8_K activation quant into the Rms…
ghazni101 Aug 24, 2026
a3a96ae
record(GFX1100-TG200): lever-C adopts VT_NORM_QUANT_FUSED=1 at +7.3% …
ghazni101 Aug 24, 2026
af3713f
perf(GFX1100-TG200): T5a vectorizes the shared Q8_K quant superblock …
ghazni101 Aug 25, 2026
2d200e1
perf(GFX1100-TG200): T5b extends the f32-Q DecodeGqa arm to head_dim 128
ghazni101 Aug 25, 2026
31458d1
record(GFX1100-TG200): T5b evidence — attention routing hole closed, …
ghazni101 Aug 25, 2026
523994b
record(GFX1100-TG200): T5c closed negative — MMVQ nontemporal weight …
ghazni101 Aug 25, 2026
dd757c0
perf(GFX1100-TG200): T6a adds the warp-per-row cooperative GDN scan arm
ghazni101 Aug 25, 2026
15b92d7
record(GFX1100-TG200): T6a evidence — cooperative scan adopted at +4.…
ghazni101 Aug 25, 2026
f2795d2
perf(GFX1100-TG200): T6b adds the warp-per-item cooperative attn prea…
ghazni101 Aug 25, 2026
5dd5f23
record(GFX1100-TG200): T6b evidence — cooperative preamble adopted at…
ghazni101 Aug 25, 2026
a6e3fdf
record(GFX1100-TG200): session-close attribution — 76.6 tok/s median …
ghazni101 Aug 25, 2026
bfc35d2
record(GFX1100-TG200): classify the campaign env knobs kernel-internal
ghazni101 Aug 25, 2026
c8f8f00
record(GFX1100-TG200): T7 evidence — COALK load topology closed wash,…
ghazni101 Aug 25, 2026
593bd90
perf(GFX1100-TG200): T8 adds a cooperative single-row rmsnorm arm
ghazni101 Aug 25, 2026
5c048b9
record(GFX1100-TG200): T8 evidence — cooperative rmsnorm adopted at +…
ghazni101 Aug 25, 2026
32174f4
perf(GFX1100-TG200): T9 gives the gated norm a per-row block
ghazni101 Aug 25, 2026
0dccc8f
record(GFX1100-TG200): T9 evidence — cooperative gated norm adopted a…
ghazni101 Aug 25, 2026
c3e90b2
perf(GFX1100-TG200): T10 and T11 add warp postconv and row-split scan…
ghazni101 Aug 26, 2026
0386491
record(GFX1100-TG200): T12 evidence — gated-quant fusion not adopted,…
ghazni101 Aug 26, 2026
310c9b5
record(GFX1100-TG200): dispatch-gap audit names the sampling round trip
ghazni101 Aug 26, 2026
4f26f9e
record(GFX1100-TG200): retract T10/T11 engine claims — corrupted outp…
ghazni101 Aug 26, 2026
bdc39d5
test(GFX1100-TG200): close the gate gap that let the T10 stride bug ship
ghazni101 Aug 26, 2026
b20ece8
record(GFX1100-TG200): add mechanical decision rules for the T10/T11 …
ghazni101 Aug 26, 2026
9aa5bc1
record(GFX1100-TG200): refine dispatch-gap into three measured sub-ta…
ghazni101 Aug 26, 2026
87c6aae
perf(GFX1100-TG200): T14 adds a row-split greedy argmax arm
ghazni101 Aug 26, 2026
6fed04a
record(GFX1100-TG200): verify at ISA level that the dp4a core uses v_…
ghazni101 Aug 26, 2026
1472103
record(GFX1100-TG200): T10/T11 re-measured in a clean window — adopte…
ghazni101 Aug 26, 2026
97f20a3
record(GFX1100-TG200): full-config verification — 92.7-92.9 tok/s can…
ghazni101 Aug 26, 2026
19e7f95
record(GFX1100-TG200): async-serving A/B is a wash under HTTP overhea…
ghazni101 Aug 26, 2026
850c879
record(GFX1100-TG200): LDS epilogue closed negative; host-load sensit…
ghazni101 Aug 26, 2026
cae25b4
record(GFX1100-TG200): position-resolved wvSplitKSml audit corrects t…
ghazni101 Aug 26, 2026
ec95552
perf(GFX1100-TG200): T16 adds wvSplitK launch-config sweep knobs
ghazni101 Aug 26, 2026
0361697
record(GFX1100-TG200): T14 stacked engine A/B closed — adopted at +0.9%
ghazni101 Aug 26, 2026
498a3b1
record(GFX1100-TG200): core pinning does not isolate host-memory cont…
ghazni101 Aug 26, 2026
1b31865
record(GFX1100-TG200): T13 implementation plan scoped with file:line …
ghazni101 Aug 26, 2026
ecdd976
adopt(GFX1100-TG200): T16 YTILE=4 default — wins 5/5 paired, bit-iden…
ghazni101 Aug 26, 2026
bcf12b1
record(GFX1100-TG200): idle-window gate 100.4 tok/s, T13 wash, copy s…
ghazni101 Aug 26, 2026
1d5e8d8
record(GFX1100-TG200): T17 v_dot2_f32_bf16 closed not-adopted — memor…
ghazni101 Aug 26, 2026
6281c45
feat(GFX1100-TG200): T18 v_dot4 instruction selection in KQuantGemvMm…
ghazni101 Aug 26, 2026
ba65fd1
record(GFX1100-TG200): T20 full-warp cooperative GEMV closed not-adop…
ghazni101 Aug 26, 2026
66adbe1
feat(GFX1100-TG200): T21 keep-quant for V-head row-permuted GDN proje…
ghazni101 Aug 26, 2026
9ef8ea3
fix(GFX1100-TG200): T22 NormQuant bridge token survives non-matching …
ghazni101 Aug 26, 2026
82cdf28
feat(GFX1100-TG200): T24 LDS-buffered quant epilogue in RmsNormRowCoo…
ghazni101 Aug 26, 2026
dfd46be
T25: keep ssm_out as Q5_K with runtime input permutation
ghazni101 Aug 26, 2026
0c991b7
T27: warp-cooperative QuantizeQ8KK for decode (+2.06%, byte-identical)
ghazni101 Aug 27, 2026
3af58a2
feat(GFX1100-TG150): fuse Q6_K bias correction into single dot product
ghazni101 Aug 27, 2026
a53f98e
spec(GFX1100-TG150): commit the 150 tok/s campaign spec
ghazni101 Aug 22, 2026
5d766e9
spec(GFX1100-TG150): add Outcome, update Now after attempt cap
ghazni101 Aug 27, 2026
14b22bc
spec(KERNEL-QUANT-CIQ-GEMM-ROCM): commit the keep-quant provider spec
ghazni101 Aug 21, 2026
ae4401b
feat(KERNEL-QUANT-CIQ-GEMM-ROCM): land the W1 keep-quant providers on…
ghazni101 Aug 21, 2026
d0a00a1
spec(ROCM-QUANT-GEMM-BW): commit the keep-quant bandwidth spec
ghazni101 Aug 21, 2026
32caa8f
perf(ROCM-QUANT-GEMM-BW): split each super-block across the warp's lanes
ghazni101 Aug 22, 2026
aeacfd5
perf(ROCM-QUANT-GEMM-BW): branch-free Q6K scalar decode + gated qg4 a…
ghazni101 Aug 22, 2026
e8b30f7
perf(ROCM-QUANT-GEMM-BW): split-K decode arm for the keep-quant GEMM
ghazni101 Aug 22, 2026
1b3cbc6
perf(BACKEND-ROCM): f32-query decode-GQA arm behind VT_ATTN_DECODE_GQ…
ghazni101 Aug 22, 2026
4539f62
perf(BACKEND-ROCM): row-permuted keep-quant in_proj + tiny-N f32-out …
ghazni101 Aug 22, 2026
f099727
cherry-pick(KV-FP8): W6 ROCm fp8-e4m3 KV cache store and read onto TG200
ghazni101 Aug 27, 2026
8baf4a0
spec(rocm): fp8 KV cache decode attention for PagedAttnDecodeGqaF32Q …
ghazni101 Aug 27, 2026
ec706ec
feat(rocm): fp8 KV cache support in PagedAttnDecodeGqaF32Q (#7)
ghazni101 Aug 27, 2026
e757b2d
perf(ROCm): parallel random sample with shared primitives
ghazni101 Aug 27, 2026
6548658
feat(rocm): advertise fp8 KV cache dtype support and add GGUF chat te…
ghazni101 Aug 27, 2026
b68025e
fix(GFX1100-TG200): repair the record and env gates the branch carrie…
ghazni101 Aug 28, 2026
4843961
fix(GFX1100-TG200): drop fp8 KV decode-attn extras onto row/fp8-kv-de…
ghazni101 Aug 28, 2026
b0ebb0b
Port async device-mirror combine/scatter kernels to ROCm
ghazni101 Aug 29, 2026
21b7892
Fuse silu-mul with the Q8_K quant epilogue on ROCm (VT_SILU_QUANT_FUSED)
ghazni101 Aug 29, 2026
e3700a5
record(GFX1100-TG200): three-point branch audit — merges help, 103→91…
ghazni101 Aug 29, 2026
7335a06
record(GFX1100-TG200): repair the three staged-preflight gate reds
ghazni101 Aug 29, 2026
13aef15
record(GFX1100-TG200): T34 splits the launch-bound residual into in-g…
ghazni101 Aug 29, 2026
0c610e9
T36: prefill M-tiled K-quant GEMM streams weight rows once (VT_PREFIL…
ghazni101 Aug 29, 2026
976372b
record(GFX1100-TG200): T35 b/a GEMV merge closed red on token coheren…
ghazni101 Aug 29, 2026
7f5d3fb
record(GFX1100-TG200): near-tie adjudication harness — teacher-forced…
ghazni101 Aug 29, 2026
789430b
record(BACKEND-ROCM): index #9, the never-run quant gate red on row/G…
ghazni101 Aug 29, 2026
2f1841e
record(GFX1100-TG200): allowlist VT_PREFILL_TILE, the T36 same-binary…
ghazni101 Aug 29, 2026
6ceece9
fix(GFX1100-TG200): stop the q6_K bias borrow corrupting the MMVQ arm
ghazni101 Aug 29, 2026
09da055
record(GFX1100-TG200): retire the stale 132,094 gate figure, log the …
ghazni101 Aug 29, 2026
ccf9a14
record(GFX1100-TG200): re-mint the campaign reference on the fixed Q6…
ghazni101 Aug 29, 2026
4f7c3be
T35-r3: GDN b/a same-input GEMV merge, adjudicated bit-identical, clo…
ghazni101 Aug 29, 2026
452acba
T37: small-N GEMV geometry levers closed negative — warps/split-K win…
ghazni101 Aug 29, 2026
a75b085
fix(ENG-MM-INPUT-PIPELINE): store the Qwen3-VL, Gemma-4 and GGUF visi…
localai-bot Aug 28, 2026
9272600
measure(PERF-LAGUNA-FUSED-GATEUP): W3 -- the lever is worth ~4% and i…
localai-bot Aug 28, 2026
3a1d40e
spec(BENCH-C8-ADMISSIBILITY): every c=8 number this repository quotes…
localai-bot Aug 28, 2026
720800b
feat(MODEL-MM-QWEN4-EXP): W5b-3 — the PLE dilated depthwise conv is n…
localai-bot Aug 28, 2026
d62c930
fix(ENG-ATTN-OPTIN-SWEEP): sweep every remaining `vt::Attention` call…
localai-bot Aug 28, 2026
34b81cd
fix(LTX25-TEXT-PROJ-DTYPE): resolve a caption projection's storage fo…
localai-bot Aug 28, 2026
243b7ba
feat(LTX25-ORACLE-ABSOLUTE): the blockiness ratios gate against #1864…
localai-bot Aug 28, 2026
f8f25de
spec(SPEC-DFLASH2): the selector's edge kernel reads every successor …
localai-bot Aug 28, 2026
efce792
record(MODEL-MM-GLM53-FLASH): W0 -- pin transformers 5.16.1 for the g…
localai-bot Aug 28, 2026
85b915d
spec(SPEC-DFLASH2): retract the selector-edge mechanism — `sample=` t…
localai-bot Aug 28, 2026
aaa4389
feat(MODEL-MM-QWEN4-EXP): W5b-4 — Qwen Sparse Attention as two `vt::`…
localai-bot Aug 28, 2026
a2a95f6
feat(MODEL-MM-GLM53-FLASH): W2 — the KDA forget gate is the sigmoid b…
localai-bot Aug 28, 2026
b53b41b
record(ENG-LTX-RECORD-RECONCILE): the trailer walk has no merge-commi…
localai-bot Aug 28, 2026
4ec9354
feat(QUANT-EXL3): W1a — EXL3 becomes a scheme on vLLM's LinearMethod …
localai-bot Aug 28, 2026
f298ca2
fix(ENG-RECURRENT-MULTISTATE): a recurrent layer carries N states, an…
localai-bot Aug 28, 2026
610b5a7
feat(MODEL-MM-GLM53-FLASH): W4 -- the mHC wiring, and a head collapse…
localai-bot Aug 28, 2026
30b7353
record(MODEL-TEXT-GLM-MOE-DSA): the row's two upstream anchors were e…
localai-bot Aug 28, 2026
ed2751b
fix(BACKEND-TENSTORRENT-QWEN35): the host-free opt-out leg gates agai…
lu-zero Aug 28, 2026
179b052
record(BACKEND-TENSTORRENT-QWEN35): reconcile the spec's Now after W4…
lu-zero Aug 28, 2026
2db3d56
record(MODEL-MM-QWEN4-EXP): a speed denominator exists, and the targe…
localai-bot Aug 28, 2026
7b534c1
spec(SPEC-DFLASH2): the draft forward is 76% of the draft phase, and …
localai-bot Aug 28, 2026
4fddb3a
fix(SPEC-DFLASH2): route the two HOT draft forward bodies through the…
localai-bot Aug 28, 2026
e45037d
measure(PERF-LAGUNA-FUSED-GATEUP): W4 partial -- prompt 0 reproduces …
localai-bot Aug 28, 2026
2a5aa9d
record(ORACLE-LLAMA-CPP-GLM5NEXT): pin the llama.cpp that can open th…
localai-bot Aug 28, 2026
2121c99
feat(QUANT-EXL3): W1b — a stock EXL3 checkpoint loads and generates, …
localai-bot Aug 28, 2026
6d2b6b0
feat(SPEC-DFLASH2): split `fwd` into its op groups, because every lev…
localai-bot Aug 28, 2026
3cf8740
feat(MODEL-MM-QWEN4-EXP): W5c-1 — the KV-cache spec is three groups, …
localai-bot Aug 28, 2026
c1f959c
feat(MODEL-MM-GLM53-FLASH): W3 makes the NoPE MLA geometry representa…
localai-bot Aug 29, 2026
f18d623
feat(MODEL-MM-dots3-note): W5 — the MoE layer reaches the decode path…
localai-bot Aug 29, 2026
9a7d5c2
record(BENCH-C8-ADMISSIBILITY): the spec named the wrong committed ha…
localai-bot Aug 29, 2026
7d3d4d0
record(BENCH-C8-ADMISSIBILITY): the right instrument cannot express t…
localai-bot Aug 29, 2026
33474b3
docs(SPEC-DFLASH2): the batched-lane spec still refused a merge that …
localai-bot Aug 29, 2026
ca10114
feat(MODEL-MM-QWEN4-EXP): W5b-5 — the QSA indexer composition moves o…
localai-bot Aug 29, 2026
5298258
record(BACKEND-TENSTORRENT-QWEN35): index the W3 leftovers issue (#2201)
lu-zero Aug 28, 2026
7144856
fix(BACKEND-TENSTORRENT-QWEN35): count the two missing GDN d2h paths …
lu-zero Aug 28, 2026
4ebf7c8
fix(BACKEND-TENSTORRENT-QWEN35): scope the GDN cache role refusal to …
lu-zero Aug 28, 2026
fa5160d
measure(PERF-LAGUNA-FUSED-GATEUP): W4 -- 6 of 6 prompts diverge, and …
localai-bot Aug 29, 2026
e1a4f35
feat(LOADER-GGUF-IQ): port the IQ2_XS and IQ4_XS dequantizers, the tw…
localai-bot Aug 29, 2026
10824c5
feat(QUANT-EXL3): W3 — a stock EXL3 checkpoint could not run on a GPU…
localai-bot Aug 29, 2026
32bee34
fix(SPEC-DFLASH2): the draft's paged attention synchronized inside th…
localai-bot Aug 29, 2026
185ffdb
spec(PERF-LAGUNA-GROUPED-GEMV): measure what bounds the grouped Q4_K/…
localai-bot Aug 29, 2026
6ed3a90
spec(MODEL-TEXT-GLM-MOE-DSA): GLM-5.3 is 97.49% routed experts, so th…
localai-bot Aug 29, 2026
efca242
measure(LTX25-ORACLE-ABSOLUTE): #1854's reading is taken, and our ren…
localai-bot Aug 29, 2026
39423bd
fix(MODEL-MM-QWEN4-EXP): one gamma polarity for the whole architectur…
localai-bot Aug 29, 2026
be76722
fix(SPEC-DFLASH2): the capture-safe bound was a per-STEP value baked …
localai-bot Aug 29, 2026
bf1f60c
record(BACKEND-TENSTORRENT-QWEN35): index the mesh CQ staging wave (#…
lu-zero Aug 29, 2026
d86c82a
perf(BACKEND-TENSTORRENT-QWEN35): stage bf16 uploads into a per-slot …
lu-zero Aug 29, 2026
4950343
record(BACKEND-TENSTORRENT-QWEN35): W5 lands allocation-free staging,…
lu-zero Aug 29, 2026
eb1b066
fix(MODEL-MM-GLM53-FLASH): read the layer schedule out of `attention.…
localai-bot Aug 29, 2026
f638441
fix(MODEL-MM-GLM53-FLASH): read `attention.key_length` the way llama.…
localai-bot Aug 29, 2026
d85ca62
record(BACKEND-TENSTORRENT-QWEN35): index the batched staging wave (#…
lu-zero Aug 29, 2026
3898251
record(BACKEND-TENSTORRENT-QWEN35): W6 is not expressible — the trace…
lu-zero Aug 29, 2026
d24c25e
record(BACKEND-TENSTORRENT-QWEN35): repair the W6 evidence citations …
lu-zero Aug 29, 2026
2f36d2b
record(BACKEND-TENSTORRENT-QWEN35): reconcile the W6 bullet's superse…
lu-zero Aug 29, 2026
f8e8a84
fix(MODEL-DSV4-EXL3): the carried tower's FP8 half is held at bf16 (#…
localai-bot Aug 29, 2026
84d9db9
fix(MODEL-MM-GLM53-FLASH): map the published GGUF's `glm4` pre name o…
localai-bot Aug 29, 2026
0a10c8f
feat(MODEL-MM-QWEN4-EXP): W5d-2 — one interleaved-mRoPE table builder…
localai-bot Aug 29, 2026
5f2b4ca
feat(QUANT-GGUF-IQ-VECDOT): keep the IQ2_XS and IQ4_XS blocks through…
localai-bot Aug 29, 2026
1ae1f0d
measure(PERF-LAGUNA-GROUPED-GEMV): W1 -- the grouped GEMV is latency-…
localai-bot Aug 29, 2026
0f7ef4d
feat(MODEL-MM-GLM53-FLASH): W5c -- the weight tower, and the model LO…
localai-bot Aug 29, 2026
6362cbe
fix(SPEC-DFLASH2): check the paged draft block's bounds on EVERY back…
localai-bot Aug 29, 2026
1faddcc
feat(MODEL-MM-QWEN4-EXP): W5d-1 — the ungated grouped RMS norm the PL…
localai-bot Aug 29, 2026
1404327
record(PERF-LAGUNA-GROUPED-GEMV): the W11 lever list is exhausted, an…
localai-bot Aug 29, 2026
e9a42f7
spec(MODEL-DSV4-DSA-COMPOSE): the DSA composition gets the owning row…
localai-bot Aug 29, 2026
2004968
fix(SPEC-DFLASH2): read the DEVICE bounds back, which is the one clas…
localai-bot Aug 29, 2026
4cfe613
fix(ENGINE-HYBRID-PLACEMENT): refuse an unplaceable MoE arm at the se…
localai-bot Aug 29, 2026
70df35d
feat(MODEL-MM-GLM53-FLASH): W5 lands the 288+1 MoE and the KV-cache s…
localai-bot Aug 29, 2026
8d1b782
docs(MODEL-DSV4-EXL3): the Owed entries pointed at an issue W1d close…
localai-bot Aug 29, 2026
67f63dd
fix(vt): close the PermuteVHeads scopes the upstream merge unified
ghazni101 Aug 29, 2026
075350b
record(GFX1100-TG200): T38 sync gates — upstream merge, reference intact
ghazni101 Aug 29, 2026
e243e8b
fix(ENV-DOC): allowlist the three TG200 tuning knobs the checker flagged
ghazni101 Aug 29, 2026
a3b6b19
record(GFX1100-TG200): pin the rebuilt sync SHAs and the owed anchor rot
ghazni101 Aug 29, 2026
4194a8e
record(GFX1100-TG200): T38 cross_device gate + near-tie adjudication
ghazni101 Aug 30, 2026
bec723d
fix(GFX1100-TG200): three pre-existing bugs + record-anchor ratchet
ghazni101 Aug 30, 2026
4c18b16
fix(runner): guard async dispatch helpers for non-GPU builds
ghazni101 Aug 30, 2026
3f6477a
feat(MODEL-MM-QWEN4-EXP): W5d-4 — the MoE weight adapter, and the fou…
localai-bot Aug 29, 2026
864ed5b
record(ENG-MM-INPUT-PIPELINE): the runner drops mm_features, so a Qwe…
localai-bot Aug 29, 2026
6b22352
perf(PERF-QWEN35-STAGE-WEIGHTS): stage dense decode weights to the de…
localai-bot Aug 29, 2026
975a1fb
feat(MODEL-MM-GLM53-FLASH): W5b-1 lands the DSA attention block, and …
localai-bot Aug 30, 2026
f72e62a
docs(ENV-DOC): document VT_QWEN35_STAGE_MIN_FREE_FRAC so the base gat…
localai-bot Aug 30, 2026
2e3c2ad
feat(MODEL-MM-QWEN4-EXP): W5d-3 — the QSA consumer can now read the P…
localai-bot Aug 30, 2026
b5ac5da
feat(MODEL-MM-QWEN4-EXP): W5c-2 gathers EVERY published KV group's bl…
localai-bot Aug 30, 2026
de1326c
feat(MODEL-MM-GLM53-FLASH): W5b-2a threads the four-stream manifold t…
localai-bot Aug 30, 2026
7367fd9
fix(MODEL-MM-QWEN4-EXP): the llama.cpp arm's KV guard was failing ope…
localai-bot Aug 30, 2026
3c95548
fix(ci): MSVC shadowing, stale anchor, stray file, upstream sync
ghazni101 Aug 30, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .agents/backend-matrix.md

Large diffs are not rendered by default.

13 changes: 6 additions & 7 deletions .agents/benchmark-record.md
Original file line number Diff line number Diff line change
Expand Up @@ -28723,13 +28723,12 @@ killed every measured leg mid-load on the previous attempt.
0.774 GiB on disk in bf16 and 1.547 GiB resident, because `qwen3_vl.cpp`
widens it to host f32 —
[#1359](https://github.com/mudler/vllm.cpp/issues/1359), which also affects
the Qwen3.6-27B path. #1359's Qwen3-VL half has since LANDED, and the
2026-08-28 rerun recorded later in this file MEASURED the consequence:
826916864 B = 0.770 GiB, **0.499x this figure**. The HALVING IS CORRECT
rather than a regression — the flag now frees the tower the checkpoint ships
instead of the tower plus our widening. The figure recorded here stands
unaltered as what the run at `41ab550b9` measured, and is superseded for
current behaviour; `muse-glimmer-30b` still widens, blocked on
the Qwen3.6-27B path. #1359's Qwen3-VL half has since LANDED, so this leg
rerun should read about 0.774 GiB rather than 1.542, and that HALVING IS
CORRECT rather than a regression — the flag now frees the tower the
checkpoint ships instead of the tower plus our widening. The figure recorded
here stands as what the run at `41ab550b9` measured; `muse-glimmer-30b` still
widens, blocked on
[#2166](https://github.com/mudler/vllm.cpp/issues/2166).
2. **Load-time residency, not a served request.**
`ForwardQwen3VLForConditionalGeneration` `VT_CHECK`s `input.mm.has_value()`,
Expand Down
360 changes: 89 additions & 271 deletions .agents/completed/issue-index.md

Large diffs are not rendered by default.

40 changes: 20 additions & 20 deletions .agents/engine-matrix.md

Large diffs are not rendered by default.

6 changes: 3 additions & 3 deletions .agents/kernel-matrix.md

Large diffs are not rendered by default.

8 changes: 4 additions & 4 deletions .agents/model-matrix.md

Large diffs are not rendered by default.

8 changes: 4 additions & 4 deletions .agents/quantization-matrix.md

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion .agents/roadmap_v1.md
Original file line number Diff line number Diff line change
Expand Up @@ -546,7 +546,7 @@ degraded run — it is no run at all.
| `--tool-call-parser` | 42 names, **84/90 recipe uses (93%)** — the healthy axis; `inkling` (2 uses) landed 2026-08-13 under #608 W1 | [#608](https://github.com/mudler/vllm.cpp/issues/608) for the last 6 |
| `--reasoning-parser` | 10 of 28 names, **15/76 uses (20%)**; `qwen3` (18) rejected on our own gate models | [#605](https://github.com/mudler/vllm.cpp/issues/605) |
| `--enable-auto-tool-choice`, `--trust-remote-code` | no-ops for us, yet **abort startup** on 89 and 82 recipes | [#606](https://github.com/mudler/vllm.cpp/issues/606) |
| `--language-model-only`, `--limit-mm-per-prompt` | **ACCEPTED + ENFORCED 2026-08-14** (#607 L2): the 43 recipes that pass the flag reach model load, and it zeroes every modality limit so mm requests are REFUSED with upstream's message and HTTP 400. L3 landed the tower skip, and on 2026-08-24 its HOST RSS was measured on ONE model: Qwen3-VL-4B-Instruct freed 1655791616 B = 1.542 GiB at load, `--device cpu`, MET on both pairs against a threshold declared before the run (`multimodal-track.md` §1.5 L3). #1359's Qwen3-VL half has since landed, and the **2026-08-28 rerun at `525d2b991` on `dgx:gpu0` MEASURED the consequence: 826916864 B = 0.770 GiB, MET on both pairs against the live 747625881 B threshold** — a 0.499x fall that is correct rather than a regression, and no longer a prediction. That rerun also verifies #1359's Qwen3-VL half: the default arm recovered 828219392 B = 99.7% of prediction while the tower-free `--language-model-only` control arm moved +0.0077% against a 2% bound. #1359 stays OPEN for #2166 (muse-glimmer) and #2173 (Gemma-4 vision, unreached). **Freed encoder VRAM is still entirely unmeasured** — that figure is host RAM on a `VLLM_CPP_CUDA=OFF` build — as is `muse-glimmer-30b` on either axis, so the flag may be described as freeing memory only with the model, the host and the load-time window named beside it | [#607](https://github.com/mudler/vllm.cpp/issues/607) |
| `--language-model-only`, `--limit-mm-per-prompt` | **ACCEPTED + ENFORCED 2026-08-14** (#607 L2): the 43 recipes that pass the flag reach model load, and it zeroes every modality limit so mm requests are REFUSED with upstream's message and HTTP 400. L3 landed the tower skip, and on 2026-08-24 its HOST RSS was measured on ONE model: Qwen3-VL-4B-Instruct freed 1655791616 B = 1.542 GiB at load, `--device cpu`, MET on both pairs against a threshold declared before the run (`multimodal-track.md` §1.5 L3). #1359's Qwen3-VL half has since landed, so a rerun should read about 0.774 GiB and that fall is correct rather than a regression. **Freed encoder VRAM is still entirely unmeasured** — that figure is host RAM on a `VLLM_CPP_CUDA=OFF` build — as is `muse-glimmer-30b` on either axis, so the flag may be described as freeing memory only with the model, the host and the load-time window named beside it | [#607](https://github.com/mudler/vllm.cpp/issues/607) |
| `--kv-cache-dtype` | not a serve flag; residual on the `KV-FP8` row | — |
| `--speculative-config` | MTP + DFlash land; `eagle`/`eagle3` (7 uses) do not | — |
| TP / EP / multi-node (`--tensor-parallel-size`, `--enable-expert-parallel`, `--mm-encoder-tp-mode`) | absent by scope, not by defect — single-box engine | see the TP W-plan above |
Expand Down
180 changes: 12 additions & 168 deletions .agents/specs/dsv4-dsa-compose.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,9 +7,8 @@ Oracle: vLLM at the parity pin `5559679229`, `vllm/models/deepseek_v4/`.
## Now

`READY` — this spec is the deliverable of the scoping wave. No implementation
has started. W1 cannot begin until `KV-DSV4-MULTICACHE` **W5** lands, and W5 has
no owner today (#2302; W3, which an earlier revision of this spec named, landed
as `ca3dcda21` on 2026-08-27). See `## Dependencies`.
has started. W1 cannot begin until `KV-DSV4-MULTICACHE` W3 lands (see
`## Dependencies`).

## Scope

Expand All @@ -27,23 +26,6 @@ the cache topology (`KV-DSV4-MULTICACHE`, #1925); residency (#2283); the
attention sink, which is a loaded per-head weight and not cache state
(`attention.py:218-222`).

**The boundary with `KV-DSV4-MULTICACHE` W5, which this row overlapped when it
was created** (#2302). W5's scope lived in a single wave-table cell reading "the
DSA-sparse attention path reading the published caches" and removing the
`!is_indexer && !is_comp` refusal -- the ALGORITHM, not the plumbing -- and it has
never had a design section, so this row was specced over it. The split, recorded
in both specs:

| row | owns |
|---|---|
| `KV-DSV4-MULTICACHE` W5 | the caches REACHING the model and each layer routing to its own: `attn_kv` consumed rather than `(void)`-ed, and `ModelRegistry::Forward`'s `input.multi_kv` guard (`model_registry.cpp`) stopping its refusal |
| **this row** | what RUNS on those caches: the three layer shapes, the compressor's two stages, the `coff` role selection and boundary emission, the indexer's `qr`-sourced query, and the `!is_indexer && !is_comp` refusal at `deepseek_v4.cpp:786-787` |

W5 lands first and deliberately does NOT remove the DSA refusal; it makes a cache
reachable for this row. Neither row is gateable end-to-end alone: W5's
synthetic-config gate proves routing, and the token-exact oracle gate above 512
tokens belongs here.

## Upstream chain

Read at `5559679229`. **The path is `vllm/models/deepseek_v4/`, NOT
Expand Down Expand Up @@ -74,21 +56,10 @@ nothing here and reads as "upstream does not implement it", which is wrong.
| SWA-only | 0 | 5 | neither |

Counts are `config.json`'s `compress_ratios` histogram `{0: 5, 4: 21, 128: 20}`
= 46 entries.

**RESOLVED, and it was already answered elsewhere in this row family.**
`dsv4-dsa-geometry.md` read the artifact's own `config.json` and records that the
46 entries are **43 layers + 3 MTP blocks**: layers 0 and 1 are `0`, layers 2..42
alternate `4` and `128`, and the MTP tail is `0`. So `num_hidden_layers == 43`,
2 layers are dense, 21 carry an indexer and 20 a compressor-only -- which is
exactly the "41 of 43 carry a compressor, 21 carry an indexer" the row already
recorded. There is no contradiction: 43 counts LAYERS, 46 counts the config list
INCLUDING the MTP tail, and the trellis shard count matches the layers.

W1 therefore inherits 43 and does not need to reconcile anything. The entry is
kept rather than deleted because #2186 raised it as open and a reader who saw
that deserves to find the answer here, with its source, rather than a silent
deletion.
= 46 entries. **46 is not 43**, and the row's older records say "43 layers" /
"41 of 43"; 43 is the trellis shard count (`exl3-layer-000..042`). W1 must
reconcile which number each claim means rather than inherit either (#2186
raised this and it is still open).

### D2. The 3-way stream overlap is performance, not correctness

Expand Down Expand Up @@ -137,100 +108,6 @@ forward refuses because `AttentionBlock` indexes the COLLAPSED geometry. So W1
is a forward change, not a loader change — and the refusal's own text is the
specification of what to build.

### W1 design — the compressor-only shape is COMPOSITION, not new kernels

Traced before estimating, because every earlier estimate on this row family moved
once the tree was read.

**Every primitive W1 needs already exists, on CPU and CUDA.**

| the shape needs | what exists |
|---|---|
| the window pass | `vt::MlaDecodeAttention` with `window_size` (`left == sliding_window - 1`) -- landed by `KV-DSV4-MULTICACHE` W5 |
| the compressed-history pass | the SAME op's SELECTED-SLOT arm, `topk_indices` + `valid_counts` |
| combining the two | `vt::MergeAttnStates` -- an LSE merge with both `+inf` and both-`-inf` edge cases ported |
| the pool | `CompressorPoolNorm` -- per-column softmax over the window, then RMSNorm |
| the APE save | `CompressorSaveScoreApe` |
| the per-head sink | `MlaDecodeAttentionArgs::attn_sink`, landed by W5 |

So W1 composes: save state each step, pool at a boundary into the compressed
cache, then TWO attention passes merged by their LSEs -- rather than the single
fused two-cache kernel upstream calls
(`flash_mla_with_kvcache(k_cache=swa, extra_k_cache=compressed, ...)`). The
composition is mathematically the same; only the kernel fusion differs, and that
is a performance question for a later wave, not a correctness one.

**`compress_ratio == 128` FIRST because `coff == 1` there.** `overlap` is
`compress_ratio == 4`, so the 128 shape has no overlapping windows and no
`head_offset` role selection -- the mechanism W5-4 of `dsv4-dsa-compose.md`
describes. It exercises the state cache, the boundary gate and the two-pass merge
without the hardest part.

#### The `c128a` selection is ARITHMETIC, not a learned top-k

The last unknown in W1's shape, and it resolves in W1's favour. The compressed
pass needs an index list, and the name upstream gives it --
`c128a_global_decode_topk_indices` -- reads like the Lightning Indexer's output.
It is not.

`sparse_mla.py:126-129` calls the field "Pre-computed C128A metadata
(compress_ratio == 128 only). Decode: global slot ids + valid-entry counts
**(fused from positions)**", and `_build_c128a_metadata` asserts
`cm.positions is not None` because positions are its only input. The selection is
therefore arithmetic over the current position -- which compressed windows have
CLOSED -- and carries no learned component at all.

That is what makes `compress_ratio == 128` the right first shape. It needs:

- the compressor cycle (landed: `CompressorStepCycle`),
- a window pass and a compressed pass merged by LSE (proven equivalent above),
- and an index list computable from `positions` alone.

The Lightning Indexer, which DOES learn its selection, belongs only to the
`compress_ratio == 4` layers and therefore to W3. A reader who assumed "topk
implies indexer" would have pulled W3's hardest dependency into W1 for no reason.

#### THE SINK MUST ENTER EXACTLY ONE PASS

The trap, written down before anyone hits it. A sink is one extra logit in the
DENOMINATOR. `MergeAttnStates` combines two states by their LSEs, and each pass's
LSE is `log sum exp(its scores)`. **If both passes seed the denominator with the
sink, the merged denominator counts it TWICE**, and the result is a plausible,
slightly-too-small attention output that no token gate would catch.

This is the same defect family as the split double-count `MlaDecodeAttentionArgs`
already documents: a sink added per split rather than in the final reduction. It
has now appeared twice in this design, which is why it is stated as a rule --
**the sink belongs to exactly one contributor to any merged denominator** --
rather than as a note about one kernel.

The gate must therefore compare a two-pass merged result against a SINGLE-pass
reference over the union of both key sets, with a non-zero sink, at a length
where both passes are non-empty. A gate where either pass is empty cannot see a
double-count.

#### The composition primitive now exists, and the rule is executable

`MergeWindowAndCompressed` in `src/vllm/model_executor/models/deepseek_v4_dsa.cpp`
is that composition: it attends the compressed rows with NO sink, and merges
against the window pass's output and LSE. The window pass carries the sink, so
exactly one contributor seeds the denominator. `PagedCausalMlaAttention` grew an
optional `out_lse` for this, because a merge needs both sides' LSEs.

Two constraints are asserted rather than assumed. Every query sees EVERY
compressed row -- a closed window is history, so no causal bound applies among
them -- and `VT_CHECK(num_tokens == 1 || num_heads == 1)` holds the point where
the two LSE layouts coincide, since `MergeAttnStates` wants `[H, T]` and the
decode op emits `[T, H]`. A general prefill step needs a transpose there and
does not get one yet; it is listed under `## Owed`.

The gate this section demanded is `tests/vllm/models/test_deepseek_v4_paged_equiv.cpp`,
"W1: two LSE-merged passes equal one pass over the union". Three mutations prove
it discriminates: seeding the compressed pass with the same sink (the
double-count itself), returning the window output unmerged, and bounding the
compressed rows causally. Each was built before it was read -- a mutation that
fails to compile leaves a stale binary reporting a pass.

## Our baseline

What this tree has TODAY, so a later reader does not re-derive it:
Expand Down Expand Up @@ -287,45 +164,20 @@ gate is token-exactness against the pinned oracle ABOVE 512 tokens
|---|---|
| W1 (#1960, `c1e6f3fb9`) | LANDED — `SlidingWindowMLASpec`, the four `MLAAttentionSpec` fields |
| W2 (#1973, `6b18829bc`) | LANDED — all seven groups / 167 entries published; runner refuses an unallocated published group |
| W3 (#2068, `ca3dcda21`) | LANDED 2026-08-27 — the runner allocates a buffer for EVERY published cache instead of one per hidden layer, and `ModelForwardInput` gained the third channel |
| W4 | proposal, **no owner** — non-uniform `block_size` |
| W5 | proposal, **no owner** — consumption |
| W6-W7 | proposals, **no owner** |
| W3 | OWED — the third `ModelForwardInput` channel |
| W4 | OWED — non-uniform `block_size` |
| W5 | OWED — consumption |

**W1 of this row cannot start before that row's W5.** The composition writes a
**W1 of this row cannot start before that row's W3.** The composition writes a
separate compressed cache beside a sliding-window raw cache, and it cannot reach
a cache the forward is not handed. This is a hard ordering, not a preference.

**W5, not W3** (#2302). An earlier revision of this spec named W3, which had
already landed when it was written. The wall today is the one the code names
itself, in `ModelRegistry::Forward`
(`src/vllm/model_executor/models/model_registry.cpp`, the `input.multi_kv` guard at the top of `ModelRegistry::Forward`):

> `... and no registered forward consumes a cache set keyed by layer name.
> Refusing rather than discarding an allocated KV topology in silence (row
> KV-DSV4-MULTICACHE W5 owns the consuming forward; #1925, #2068)`

A DeepSeek-V4 engine therefore constructs, publishes and ALLOCATES all 167
buffers today, and refuses at the first forward. **This makes the ordering
harder than the earlier revision claimed, not softer:** W3 had an owner and
landed, while W4 through W7 are proposals with no owner at all. Nothing in this
row can begin until W5 acquires one.

The error is recorded rather than quietly corrected because its cause is
reusable: two stale records agreed with each other and neither was the tree.
#1925's index row predates W3, and `kv-dsv4-multicache.md` `## Now` opened with
"W3 (#2068) is claimed" while its own closing paragraph already said the engine
allocates all 167 buffers and refuses naming W5. AGENTS.md `## History is git`
is explicit -- "Before you conclude anything about past work, read the spec and
run `git log -S`" -- and `git log --oneline --grep '2068'` shows `ca3dcda21`
immediately. It was not run.

## Work breakdown

| wave | scope | depends on |
|---|---|---|
| W0 | this spec | — |
| W1 | reconcile 43 vs 46; layer-shape dispatch in `AttentionBlock`, SEQUENTIAL, replacing the refusal for the `compress_ratio == 128` (compressor-only) shape first | multicache W5 |
| W1 | reconcile 43 vs 46; layer-shape dispatch in `AttentionBlock`, SEQUENTIAL, replacing the refusal for the `compress_ratio == 128` (compressor-only) shape first | multicache W3 |
| W2 | the compressor's two stages: `save_partial_states`, then boundary-gated compress/norm/RoPE/quant/store | W1 |
| W3 | the overlapped window and role selection; the `compress_ratio == 4` shape; the indexer's `qr`-sourced query | W2, and both kernel `SPIKE` rows promoted |
| W4 | the stream overlap, as a measured performance wave | W3 |
Expand Down Expand Up @@ -365,18 +217,10 @@ failure mode this row is most exposed to.
- The `43` vs `46` layer-count reconciliation, raised by #2186 and still open.
- Promotion of `KERNEL-ATTN-DSA-SPARSE-INDEX` and `KERNEL-ATTN-DSA-COMPRESSOR`
out of `SPIKE`.
- The `[T, H]` to `[H, T]` LSE transpose `MergeWindowAndCompressed` refuses, so a
PREFILL step with more than one token and more than one head can compose. A
decode step is unaffected: it carries one token.
- The `AttentionBlock` compressor arm itself. The primitive above is reached only
by its gate until that arm calls it, and the `compress_ratio == 128` refusal
stays in place until then.

## Stop conditions

- Stop if `KV-DSV4-MULTICACHE` W5 does not land: W1 has no cache to READ. W3
(the allocation and the forward channel) landed as `ca3dcda21`; W5 is the
consuming forward, and it has no owner (#2302).
- Stop if `KV-DSV4-MULTICACHE` W3 does not land: W1 has no cache to write to.
- Stop before claiming any speed number. This row makes the model RUN; a
throughput comparison against SparkInfer's 44-47 tok/s additionally needs
`nvfp4_ds_mla` and K5 speculative decoding, neither of which exists here.
Loading
Loading