Add GLM-5-Next (GLM-5.3-Flash): KDA + DSA hybrid attention, MoE, lightning indexer, vision - #116
Add GLM-5-Next (GLM-5.3-Flash): KDA + DSA hybrid attention, MoE, lightning indexer, vision#116danielhanchen wants to merge 28 commits into
Conversation
Metadata and tensor loading only. The graph entry point throws, as qwen4exp did at the same stage. kda.gate_lower_bound is read as required: kimi-k3 selects the softplus branch when it is absent, which is a different function rather than a missing clamp. The absorbed MLA projections are 3D, so glm5next joins bailingmoe3 in the MXFP4 carve-out that would otherwise quantize them as expert tensors.
glm5next's mHC is DeepSeek-V4's hyper-connection block: same wide residual, same 24-row mixer split, same two activations, same Sinkhorn. Only the final collapse differs, so the graph derives from llama_model_deepseek4::graph and reuses build_hc_pre / build_hc_post / build_hc_sinkhorn rather than restating them, as graph_dsv4 already does in dflash.cpp. dsv4_hc_mean becomes a static member so both archs can reach it; the body and both deepseek4 call sites are otherwise untouched. The generated code for deepseek4 is unchanged apart from the endbr64 landing pad the helper now needs as a global symbol. The four streams start as exact copies of the token embedding and collapse to an unweighted mean after the last layer: this checkpoint has no hc_head. KDA, DSA and the MoE land in later commits, so the two sublayers throw. The mHC wiring around them is final.
Copy-adapts kimi-k3's KDA layer rather than kimi-linear's or bailingmoe3's: it already matches on the recurrence ordering, the bounded-sigmoid decay gate and its branch selection, dt_bias added per channel before the reshape, per-head A broadcast, SiLU after the conv, f/g/beta read from the pre-convolution hidden states, and the gated output RMSNorm with a plain weight. Three differences from kimi-k3. The output gate is low rank, g_b(g_a(x)) as in kimi-linear, which is what PR 1's converter emits. The q/k L2 eps is a literal 1e-6, the reference's own constant, not f_norm_rms_eps; ggml_l2_norm implements max(sqrt(sum), eps) rather than sqrt(sum + eps), which at head_dim 128 differs by about eps/(2*sum) and never trips the clamp, so it is close but not bit-exact. And the cross-layer residual, latent MoE, situ activation and MLA output gate have no counterpart here. The conv follows the reference and convolves q|k|v as one depthwise kernel, which keeps the conv state a single contiguous block so build_conv_state can snapshot it. That plus build_recurrent_attn is what makes the layer safe under recurrent-state rollback, so the arch joins llm_arch_supports_rs_rollback; without that entry the guard in llama_context silently clamps n_rs_seq to 0. build_delta_net_autoregressive reshaped a per-channel KDA gate onto ne1, but ne0 is the key axis everywhere else in that function, so it decayed along the value axis. Invisible for GDN, where the gate is scalar and both spellings produce the same [1, 1, H_v, n_seqs], and invisible to the shape checks because S_k == S_v. Fixed rather than asserted around, since glm5next reaches that path on any backend without the fused operator. llama_model_deepseek4::graph now derives from llm_build_delta_net_base so glm5next, which derives from it for the mHC residual, can reach build_delta_net. The base is a method-only mixin over llm_graph_context with no data members and no virtuals beyond the destructor llm_graph_context already has; deepseek4.cpp, dflash.cpp and kimi-k3.cpp compile to byte-identical instructions across the change. graph_max_nodes moves the arch to kimi-k3's tier. Measured on the Tiny fixture with the chunked fallback: 182 nodes plus 15/16 per token for each KDA layer and 46 per layer for the mHC mixers, so the 45-layer model needs 8.3k + 31.9 per token before DSA or the MoE are counted, which overruns the n_tokens*40 budget. test-llama-archs synthesised no MLA, hyper-connection, kpool or expert-weight keys for glm5next, so PR 1's required get_key calls threw out of the sweep and truncated it at 75 of 143 architectures. The fixture is complete now and the row is skipped explicitly while the DSA and feed-forward sublayers still throw.
The routing is DeepSeek-V3 noaux_tc exactly as build_moe_ffn already implements it: sigmoid scores, exp_probs_b added for the top-k SELECTION only, weights gathered from the unbiased scores, normalised, then scaled by routed_scaling_factor. n_group and topk_group are both 1, so the group-limited stage is degenerate and build_moe_ffn's n_expert_groups > 1 guard skips it; no group keys are written and none are needed. The clamp is the one thing that needed a change outside this arch. glm5next clamps the gate max-only and the up symmetrically, both BEFORE the SiLU, which is what the branch behind the DEEPSEEK4/DFLASH arch gate already does; the else branch clamps after the SiLU and is a different function. Adding the arch to both gates reuses it rather than restating it. The two conditions are separate because the dense path and the MoE path read different hparams arrays. The leading dense layers clamp too. The reference builds them from the same Glm5NextTextMLP as the shared expert, so swiglu_limit is not MoE-only, and the converter already writes swiglu_clamp_shexp for every layer rather than only the sparse ones. The shared expert is added unscaled.
nope-only MLA in the absorbed form, over every cached position. below index_topk + index_kpool - 1 resident tokens the indexer selects all of them, so this is exactly what the sparse path degenerates to, and it is a reference the sparse commit can be checked against. the attention half of the hybrid memory becomes the K-only variant: after absorption the cache holds the kv_lora_rank latent and V is a view of K.
both are required keys for glm5next, so a model saved without them cannot be loaded back. this is what stops test-llama-archs from round-tripping the arch.
the DSA sublayer no longer throws, so the arch can construct and run. it needs the MLA head shape as well: with n_head_kv taken from the per-layer array it would size the K cache row n_head times wider than the latent the graph writes.
index_topk + index_kpool - 1 is the number of positions the indexer keeps, and it is what makes the dense attention this branch builds exactly equal to the sparse path below that many cached tokens. an off-by-one in it is invisible to every output comparison measured so far, on both a dense and a sparse fixture, so it is checked against a second spelling of the same arithmetic instead.
The DSA layers of this model score pools of index_kpool consecutive positions
rather than single keys, and the pooled key cannot be rebuilt from the MLA
latents. llama_memory_hybrid therefore gains an optional third cache holding one
indexer key and one compressor gate per token, so the hybrid carries the KDA
conv+recurrent state, the MLA latents and the indexer keys at once.
Absent unless filter_idx is given, which defaults to null, so every existing
architecture gets exactly what it got before, state file layout included.
Two heads per cell, not one. GLM's compressor is not a mean pool: it is a
per-channel softmax over the kpool slots with logits gate + ape, where the gate
is a second projection of the hidden state of width indexer_head_size. Caching
it beside the key is the only way a pool survives its member tokens leaving the
batch. Architectures with indexer_kpool == 0 still get one head.
The indexer cache is handed the attention cache's slot layout rather than
finding its own, so the two agree cell for cell, and apply() asserts they do.
It also keeps its own dtype: -ctk q8_0 would otherwise quantise the gates, which
feed a softmax.
llama-kv-cache-kpool.{h,cpp} builds the pool <-> cell map host side. Pools are
defined on positions and cells are whatever find_slot handed out, so the
correspondence cannot be derived in the graph. Nothing here emits a negative
index: ggml_set_rows asserts i1 >= 0, so unpopulated entries are clamped into
range and neutralised by an additive -INFINITY instead.
Two things the map does that the qwen4exp shape it is ported from does not:
- the top-k budget is indexer_top_k exactly, with the always-selected tail
biased to -INFINITY so it spends none of it, and forced back in through a
host-built base mask for the scatter. indexer_top_k is a whole number of
pools, so the cut lands on a pool boundary; the reference's own output width
of indexer_top_k + kpool - 1 does not, and ggml_top_k is unordered among
equals on both CPU and CUDA.
- one map per ubatch, shared by every indexer layer, since nothing in it
depends on the layer. Measured on a 16 Ki cell cache with 512 tokens:
~4 ms once against ~4 ms x n_layers.
A unified cache with more than one sequence would let two sequences at the same
position pool each other's keys, so create_memory refuses it up front rather
than aborting mid-run.
tests/test-glm5next-memory.cpp: 74 checks, 0 failures, on both the full and the
trunk-only fixture. test-llama-archs is byte identical to the same build without
this commit at a fixed seed: 452 rows, 0 FAIL. Session state files for
qwen3next, falcon-h1, minimax-01, qwen35moe and a real Falcon-H1-0.5B are byte
identical too, across write, reload and rewrite.
Builds the pooled lightning indexer and gives the DSA layers a sparse attention path driven by it. Top-k runs over the POOL axis at select_k = index_topk/index_kpool, and the selected pools are expanded to their member cells through pool_cells. That is the reference's own two-step (modular_glm5_next.py, Glm5NextTextIndexer.forward: topk over the pool axis, then selected_indices = pool_indices[batch_idx, selected]), and it is not interchangeable with a single top-k of width index_topk over member cells. The argument for the cell-level form - a pool's members carry its score bit-exactly, so the cut must land on a pool boundary - assumes tie groups never span pools. They do: ReLU drives most pool scores to exactly 0.0, and ggml_top_k is explicitly unordered among equals, so the cut falls inside an inter-pool tie group and splits a pool. Measured on TinySparse at 512 tokens, the cell-level form leaves a partial pool on 7.51% of query rows at layer 3 and 5.93% at layer 7; this form leaves none. The indexer key and gate STORE is unconditional; only the SCORING is gated, on n_ctx > index_topk + index_kpool - 1. Gating the store the same way would leave every cell written below n_select with no indexer state, and the first ubatch to cross n_select would pool cells that were never written. Nothing here changes any other architecture: test-llama-archs produces a table byte-identical to the parent's, 300 rows over 143 archs, 0 FAIL.
The tower is the GLM-OCR ViT with a clamped SwiGLU: the gate is bounded above only, the up projection on both sides, and both before the SiLU. ggml_swiglu_oai clamps the same way but then adds one to the up branch, which is a gpt-oss detail this model does not share, so this adds an FFN_SILU_CLAMP op rather than reusing it. The clamp sits at the per-block MLP and again at the merger. Both read hparams.ffn_op, so the graph body stays the GLM-4V one and the pair is covered together. It gets its own projector type rather than a flag on glm4v because the image token limits differ (16/8000 against 8/4096, per the GLM-5.3-Flash preprocessor) and those are hardcoded per projector, and because the clamp must stay off for GLM-4V and GLM-OCR. Also writes clip.vision.spatial_merge_size. No GLM4V-family mmproj has ever carried it: Glm4VVisionModel skips the Qwen3VL parameters, which is where it is written, so clip.cpp's hardcoded 2 has been carrying it. Images only. glm5next spells video with its own token pair and distinct start/end spans, and that is not handled here.
the vision tower shipped with the shared dynamic-size preprocessor, which is a qwen-style smart_resize. the 2026-08-26 GLM-5-Next adaptation resizes differently: both edges are aligned up by ceil rather than round, an over-budget image is fitted by binary searching the content height for the largest aligned canvas still within max_pixels, and the resized content is pasted into the top-left of that canvas rather than centred and stretched to fill it. an image already at or above min_pixels is never upscaled. min_pixels/max_pixels stay in tokens. the reference scales them by temporal_factor * factor**2 and compares against aligned_frames * area, and aligned_frames equals temporal_factor for a still image, so the two cancel and hparams.image_min_pixels / image_max_pixels (16 and 8000 tokens, 12544 and 6272000 pixels) are used directly. glm4v and glm-ocr keep the dynamic-size preprocessor. images only. video has its own token pair (154855, distinct from the image token 154854) with its own start/end spans, and is out of scope here. the resize arithmetic is covered in test-mtmd-impl against values taken from the reference processor, including the 16- and 8000-token boundaries, extreme aspect ratios, and inputs where the binary search and smart_resize disagree.
glm4 / chatglm-bpe tokenizer.json files set "ignore_merges": true, meaning a pre-token that is already a vocab entry is emitted directly and the merge loop never runs. llama.cpp implements this (llama-vocab.cpp, the get_ignore_merges() short-circuit) but only enables it for a hardcoded list of pre-tokenizer names, and glm4 was never added. Without it the merges are applied - correctly - and reach a different answer, because greedy BPE cannot always reconstruct a vocab entry from its bytes. " 王" (Ġçİĭ, id 102322) is the case that exposed it: from Ġ ç İ ĭ the only merges available are (Ġ,ç)=27944, (ç,İ)=76417 and (çİ,ĭ)=239209, so the lowest rank wins first and yields Ġç İ ĭ, at which point neither (Ġç,İ) nor (İ,ĭ) exists and it stops three tokens short. Reaching Ġçİĭ needs (Ġ,çİĭ) at 242943, which requires never taking (Ġ,ç) at 27944. The trigger is whitespace immediately before a CJK character, so pure Chinese prose is unaffected and mixed Chinese-English is not: pure Chinese prose 620 vs 620 tokens, already identical mixed Chinese-English 680 -> 600 tokens, now identical to HF (-13.3%) wikitext-2 (289569 tok) one divergence -> byte-identical Found while comparing GLM-5.3-Flash perplexity against transformers, vLLM and SGLang: the mismatch bounded how many scoring windows could be compared at long context, and reads exactly like a model-port defect rather than a tokenizer one.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e1cfee7077
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| // constants are positive, so this is exactly the same function, and it is what the | ||
| // engines and the in-tree glm-dsa both do | ||
| ggml_tensor * w = ggml_mul_mat(ctx0, layer.indexer_proj, cur); | ||
| ggml_mul_mat_set_prec(w, GGML_PREC_F32); |
There was a problem hiding this comment.
Force full FP32 for the indexer projection
On NVIDIA CUDA when NVIDIA_TF32_OVERRIDE is unset, this projection still runs through the global cuBLAS CUBLAS_TF32_TENSOR_OP_MATH mode in ggml/src/ggml-cuda/common.cuh, despite GGML_PREC_F32. These scores determine the selected attention pools, so reduced-mantissa products can change the sparse mask; the commit's validation reports only 0.896 top-1 agreement without the environment override. Correct inference should not depend on callers discovering and setting a backend-wide environment variable, so this operation needs a code-enforced full-precision path.
Useful? React with 👍 / 👎.
| cb(k, "dsa_kv_latent", il); | ||
|
|
||
| if (top_k) { | ||
| cur = build_attn_sparse(inp_attn, |
There was a problem hiding this comment.
Disable the incompatible flash-attention path
For contexts above the selection threshold, this branch eventually calls build_attn_mha, which uses flash attention whenever the default LLAMA_FLASH_ATTN_TYPE_AUTO remains enabled, as it normally does on CUDA. The commit's own validation states that GLM5Next requires -fa off for correct output because this MLA path loses unacceptable precision, but nothing here or during context creation enforces that requirement. Consequently an ordinary CUDA invocation silently produces degraded logits; force flash attention off for this architecture or provide an MLA-safe implementation.
Useful? React with 👍 / 👎.
Master picked up the upstream sync to b10632, which brings in kimi-k3,
bailingmoe3, minimax-01, nemotron-h, lfm2 and dots3note. Three consequences:
* kda.gate_lower_bound and the swiglu-clamp truncation now come from
upstream, so the two commits that added them here are dropped.
* every arch-list conflict is a union: upstream's entries stay, glm5next is
added alongside them.
* llama-quant's 3D-MLA guard was bailingmoe3-only; generalised to cover both
arches rather than duplicating the branch.
graph_max_nodes now follows upstream's shape, which also means glm5next picks
up the lora node accounting the previous version here skipped.
e1cfee7 to
ceb79e2
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ceb79e2c6c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| // trunk residual is hc_mult streams wide (deepseek4); lm_head still sees | ||
| // n_embd, the streams are averaged first | ||
| hparams.n_embd_out_impl = hparams.dsv4_hc_mult * hparams.n_embd; |
There was a problem hiding this comment.
Report the post-mean embedding width
When callers enable token or pooled embeddings, llama_context allocates and copies hparams.n_embd_out() floats per output, but this graph assigns res->t_embd after build_hc_mean, where its width is only n_embd. Setting n_embd_out_impl to hc_mult * n_embd therefore makes the embedding extraction at llama-context.cpp:1887-1893 request four times the tensor's data for the released model, causing an out-of-bounds backend read/assertion and making llama_model_n_embd_out() inconsistent with the returned tensor. Keep the widened dimension only for an eventual MTP hidden-state interface, not the normal embedding output.
Useful? React with 👍 / 👎.
deepseek4 sets n_embd_out_impl to hc_mult*n_embd to size its MTP h input. glm5next inherited that, but our t_embd is build_norm(build_hc_mean(...)), which is [n_embd, n_tokens]. n_embd_out() therefore reported 4*n_embd while the tensor held n_embd, and llama-context read n_outputs*n_embd_out floats out of it, four times what is there. The assert at that site sizes the destination buffer, so nothing catches the short source. Only --embeddings and llama_get_embeddings* reach the path, which is why plain generation never showed it. Note for when the NextN graph starts consuming h: give MTP its own width rather than widening n_embd_out again.
The mHC residual mixers, the lightning indexer (selection gate, learned k-pool position table, and the three indexer projections) and the KDA recurrence gates are about 1 GiB in total on GLM-5.3-Flash, so the size cost is noise against a 100-240 GB quant. Quantizing them perturbs which pools the indexer selects and how much state each KDA step retains, and those errors compound along a sequence rather than averaging out. Both spellings are required. The compressor tensors arrived with the DeepSeek-V4 merge and use an underscore (indexer_compressor_ape / _gate), while the projections use a dot (indexer.proj / .attn_k / .attn_q_b), so a single "indexer." prefix test silently misses the compressor pair. attn_q_a, attn_kv_a_mqa, attn_k_b and attn_v_b are deliberately not listed. They are precision sensitive too, but the release recipe pins them to q8_0 via --tensor-type, and that is the configuration the shipped quants were measured in. Verified with llama-quantize --dry-run q4_k_m on the BF16: all 12 pinned families report 0 quantized (45 mHC, 12 indexer, 34 KDA each), while ffn_gate_exps 43/43, attn_q_a 12/12 and attn_output 46/46 still quantize.
The 101-line glm5next block was dropped from test-mtmd-impl.cpp when the vision work was rebased, even though the commit message still claimed the resize arithmetic was covered there. It holds the 36-case table over the 16- and 8000-token budget boundaries, including six cases annotated as ones where a naive smart_resize disagrees, so it is the guard against sliding back to stretch-to-fill instead of ceil-align plus zero pad. Restored from 29c096371. test-mtmd-impl now runs 216 assertions, of which glm5next_resize contributes 185.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d74d26cc79
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
|
||
| llama_build_and_test(test-llama-archs.cpp) | ||
| # needs a glm5next GGUF as argv[1], so it is built but not registered | ||
| llama_build(test-glm5next-memory.cpp) |
There was a problem hiding this comment.
Remove the unapproved standalone test target
Unless explicit maintainer approval accompanies this change, do not add tests/test-glm5next-memory.cpp: llama_build creates a default executable, so every test build compiles this 1,217-line binary even though it is not registered with CTest and requires an external GGUF to run. Move the coverage into existing infrastructure or obtain and document approval.
AGENTS.md reference: AGENTS.md:L84-L87
Useful? React with 👍 / 👎.
| pre_type = LLAMA_VOCAB_PRE_TYPE_CHATGLM4; | ||
| special_bos_id = LLAMA_TOKEN_NULL; | ||
| // glm4 tokenizer.json sets "ignore_merges": true. without it greedy BPE | ||
| // cannot reach some vocab entries: " 王" (Ġçİĭ, 102322) needs (Ġ,çİĭ)=242943 |
There was a problem hiding this comment.
Encode the tokenizer example using ASCII
Replace the literal CJK and mojibake glyphs in this added comment with an ASCII representation such as escaped UTF-8 bytes or token IDs. The repository explicitly prohibits Unicode characters, so this comment violates the required source convention.
AGENTS.md reference: AGENTS.md:L69-L73
Useful? React with 👍 / 👎.
add_vision_swiglu_limit was inserted directly above the next method with no blank line between them, which flake8 flags as E301. Caught by ggml-org CI.
llama-embedding turns -np 1 into kv_unified with n_seq_max 256, so every --embeddings run hit the refusal in create_memory and llama-embedding then dereferenced the null context. Two fixes. The pool map is now per SEQUENCE rather than per stream. A non-unified cache already gives one sequence per stream, so nothing changes there. A unified cache puts every sequence of the ubatch in stream 0, and the stream's pool table is cut into one contiguous run per sequence, each rebased on its own lowest resident pool. pool_bias is -INFINITY outside the query's own run, so a query never spends budget on a foreign pool, and cand_mask already kept foreign cells out of the attention mask. Packed runs, not one full-width table per sequence: the indexer scores every pool slot against every query, so a full-width table per sequence multiplies the score tensor by the sequence count and graph_reserve asked for 286 GB at n_seq_max 256. The table is n_kv/kpool shared plus 2 slots per sequence for rebasing, which is exact while the sequences' cells are disjoint. A prefix shared through seq_cp can oversubscribe it, and then a sequence keeps its newest pools -- the same cut a large hole in the cache already forces. The top-k stays over POOLS and pool_cells still holds whole pools, so pool integrity is untouched. examples/embedding also checked only the model for null, not the context.
test-glm5next-memory asserted that create_memory REFUSES -kvu with n_seq_max 2, which was the contract before the pool map became per sequence. Assert the new one: the cache is built, and one ubatch holding both sequences is driven through llama_kv_cache_set_input_kpool. Three checks replace the guard. llama_kpool_n_pools is n_kv/kpool plus 2 slots per sequence, so the table is a shared budget and not one full-width table each. Every pool a query may spend budget on holds only that query's own visible cells -- the invariant a shared cells array breaks if the map is keyed per stream. The two sequences get disjoint runs and neither run is empty. Both cell-level checks fail if the runs are made to overlap, so they are not tautologies. cell_pool is not requested here: it has one row per stream and a cell that two sequences share has nowhere to put its second pool.
sel_mask and cand_mask are KQ-mask shaped and hold only 0.0f and -INFINITY, both exact in f16, so storing them in half the bytes is lossless. At n_ctx = 1 Mi, n_ubatch = 512 that is 2 GiB saved per mask plus 1 GiB on the per-layer ggml_dup. ggml_add gives its result src0's type and f16 + f32 -> f16 is a supported bin_bcast on CUDA and on the CPU, so the f16 selection mask absorbs the f32 KQ mask that flash-attention-off builds, and ggml_soft_max_ext takes an f16 mask as readily as an f32 one. Under flash attention the KQ mask is already f16 and the per-layer ggml_cast disappears. llama_kv_cache_set_input_kpool now writes either width and asserts the two masks share a type instead of asserting f32.
Replaces the 7-node score chain (mul_mat, cont/permute, relu, mul, sum_rows, cont/permute, add) with one ggml_lightning_indexer, as glm-dsa, deepseek4, deepseek32 and dots3note already do. The op needs an f16 mask, so pool_bias is cast once per graph in build_inp_kpool rather than once per DSA layer. pool_k is left in f32 so the CUDA op takes its f32 vector path, not the f16 wmma path, which would undo the GGML_PREC_F32 on the head weights. The unfused chain stays behind cparams.fused_lid, plus a LLAMA_FUSED_LID_DISABLE escape hatch.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 65fc13e53a
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| """ | ||
|
|
||
| model_arch = gguf.MODEL_ARCH.GLM5NEXT | ||
| supports_mtp_export = True |
There was a problem hiding this comment.
Do not advertise an unusable MTP draft export
When a user passes --mtp for GLM5Next, this flag makes the converter accept the request and emit an MTP-only GGUF; the CLI documents that output as suitable for speculative drafting (convert_hf_to_gguf.py:121-123). However, loading it as an MTP context always reaches the assertion in llama_model_glm5next::build_arch_graph at src/models/glm5next.cpp:867-870 because no NextN graph is implemented, while loading the head-only file as a normal model lacks the trunk tensors. Remove MTP export support until the draft graph is usable, or implement that graph.
Useful? React with 👍 / 👎.
| if (n_want > n_pools) { | ||
| int64_t rem = n_pools; | ||
|
|
||
| for (int64_t ps = 0; ps < n_ps; ++ps) { | ||
| run_len[ps] = std::min(run_len[ps], rem/(n_ps - ps)); | ||
| rem -= run_len[ps]; | ||
| } |
There was a problem hiding this comment.
Preserve all pools when sequences share a prefix
With a unified KV cache after llama_memory_seq_cp, the same physical prefix cells belong to multiple sequences, so the sum of their logical pool ranges can exceed n_kv / kpool. This branch responds by shortening every sequence's run, and the later b_base calculation keeps only the newest pools; older shared-prefix cells consequently get pool_of == -1 and are excluded by cand_mask, even though the reference indexer may rank those pools in its top-k. Parallel or speculative decoding that copies a sufficiently long prefix therefore silently attends over only part of the valid history; the pool representation must retain all logical per-sequence pools rather than truncate them.
Useful? React with 👍 / 👎.
|
Superseded by the ggml-org#27754 pin in scripts/unsloth/pr-set.json, which #159 moved to 5796547. The nightly builds the upstream base tag plus the pins and never compiles fork master, so this PR does not affect what ships. Its branch is also 199 commits behind the pinned commit: it still carries the "glm5next NextN graph not implemented yet" assert, so no MTP, and its seq_add predates the pooled-key cache. Keeping the glm5next/public branch. tests/test-glm5next-memory.cpp lives only there, and was dropped from the upstream PR when tests were stripped for size. |
Adds support for GLM-5-Next (released as GLM-5.3-Flash), a 321.3B hybrid
linear/sparse-attention MoE, plus its vision tower.
Architecture
Every hparam below was read out of three reference implementations
(transformers, vLLM, SGLang) rather than one, and taken as settled only where
at least two agreed.
at indices 3, 7, ... 43. Layer 44 is KDA.
gate_lower_bound * sigmoid(exp(A_log) * (f_b(f_a(x)) + dt_bias)).gate_lower_boundis -5.0 and is a multiplicative scale, not a clamp, whichis the one detail most likely to be implemented backwards.
f_norm_rms_eps.kv_lora_rank512,q_lora_rank1536,qk_nope_head_dimand
v_head_dim256, andqk_rope_head_dim0.mla_use_nopeis true, sothere is no RoPE anywhere in the text tower.
rope.dimension_countiswritten explicitly as 0 rather than omitted, otherwise the generic loader
defaults it non-zero.
index_topk2048,index_kpool4.ReLU sits between the QK dot product and head weighting,
k_normis aLayerNorm with bias at a hardcoded 1e-6, and
weights_projissign-unconstrained and runs in fp32.
where the bias affects selection only,
routed_scaling_factor2.5 appliedafter
norm_topk_prob, andfirst_k_dense_replace3.Indexer selection is over pools, not cells
The reference picks
index_topk // index_kpoolwhole pools and then expandseach to its members. Selecting the same number of individual cells is not
equivalent and is an easy mistake to make: the ReLU drives many distinct pools
to exactly 0.0, and
ggml_top_kis unordered among equal keys, so a cell-leveltop-k splits pools apart.
That bug is also close to invisible under the obvious metric. A cell-level
implementation scored 0.9958 and 0.9764 Jaccard against the reference selection,
against a bf16-vs-bf16 floor of 0.9779 and 0.8385, so it sat above the noise
floor while being wrong. What does catch it is counting partially selected
pools, which is reference-free and reads 0 for the pooled implementation and 19
rows of 70 pools for the cell-level one.
Running it
Two flags are currently required for correct output:
NVIDIA_TF32_OVERRIDE=0.ggml-cuda/common.cuhsetsCUBLAS_TF32_TENSOR_OP_MATHunconditionally, so every fp32 GEMM otherwiseruns at 10 mantissa bits. On a fixture this moved top-1 agreement from 0.896
to 0.9995.
-fa off.build_attn_mhacasts the F32 latent to F16 beforeggml_flash_attn_ext, which is the one place MLA cannot afford it.KV cache type is not a correctness requirement. f16 costs +0.0005 PPL at ctx
2048 and is 0.0018 lower at 4096, both within engine-to-engine noise.
What is verified, and what is not
Verified on this branch, CPU only: it builds clean, and
test-llama-archspasses for glm5next at 0.00e+00 with the whole suite green.
The numerical work below was done on the same code before this rebase, on
CUDA, against the real 641 GB BF16 GGUF:
sparse selection, measured by transformers against itself, is 0.0011 PPL
(2.988491 sparse vs 2.987348 forced dense), which is smaller than the 0.0015
spread between transformers, vLLM and SGLang. Treat that as the resolution
limit of any parity claim here, not as a pass mark.
including two recall probes that are unanswerable without the history.
Not yet done: CUDA verification of this branch specifically, CPU/CUDA agreement
for the indexer, perplexity at 16k, and imatrix or quantized builds.
Notes for review
decoding.
bailingmoe3, minimax-01, nemotron-h, lfm2 and dots3note, which supersedes two
commits this PR used to carry:
kda.gate_lower_boundand the swiglu-clamptruncation in
llama_model_savernow come from upstream, so both weredropped rather than duplicated.
glm5next is added alongside.
llama-quant's 3D-MLA guard was bailingmoe3-onlyand is generalised to cover both arches instead of duplicating the branch.
tests/test-mtmd-impl.cpp, which along with its CMake entry and the mtmdsymbol-export flag is not present in this repo, and porting that harness felt
like unrelated scope for this PR.