Repository navigation
fix: clean up paged decode and SSM precision - #2270
Merged
Merged
Conversation
Remove the unused dense and rotating paged compatibility paths while preserving pooled decode coverage with an explicit float32 dense oracle. The oracle retains its absolute RMS checks and adds normalized RMS checks so a one-percent output perturbation is rejected. Promote both dt operands before the compiled SSM softplus so bf16 activation parameters cannot round the float32 recurrence. Add the paired/null serving benchmark driver and analyzer for the pending Gemma 3 and Llama 4 storage decision. Refs #2230
Record both the configured MLX pin and the release build's actual MLX checkout, rejecting a mismatch before measurement. Check that the serving port is free and reuse the repository hostgate before every fresh-server arm, preserving per-arm wait, CI, process, memory, load, and driver evidence for the multi-hour run. Refs #2230
Retry the host gate when activity appears after the quiet wait, and refuse to record any nonquiet arm. Require complete provenance, exact ordered schedules and logs, quiet evidence, and backend markers before computing the paired/null storage decision. Add lightweight fixtures for valid evidence plus missing, dirty, reordered, nonmonotonic, and substituted protocol artifacts. Refs #2230
Record the complete 60-arm paired/null storage sweep and retain paged storage for Gemma 3 and Llama 4 because no cell clears its repeated-paged null spread. Preserve the raw host-gate, client, server, schedule, artifact, derived binary-source provenance, SSM stage, and Granite checkpoint evidence with checksums. Cite the dated decision from both model flags and the continuous-batching guide. The original measured scripts and their recorded hashes remain unchanged. Validation: analyzer accepted all 60 arms and 180 cells; nine protocol fixtures passed; both artifact manifests verified. Refs #2230
Commit the release-build transcript referenced by the storage sweep's derived binary-source proof. Clarify that the transcript proves the successful 19m43s build but does not print HEAD; the production-source association remains reconstructed from session context, artifact identity, ancestry, and the protocol-only intervening diff. Validation: the release-log SHA-256 matches the provenance record and the refreshed evidence manifest verifies. Refs #2230
Preserve the measured client, SSM stage, and shell-command records byte-for-byte while disabling only their known blank-EOF or trailing-space checks. Refresh both evidence manifests to cover the scoped attributes. Validation: both manifests verify and the branch diff passes Git's whitespace check. Refs #2230
inureyes
added a commit
that referenced
this pull request
Oct 11, 2026
Choose the server prefill chunk from live scheduler state: 2048 tokens while no other sequence is decoding and 512 while a decode batch is active. Rebuild parked continuations from their live cursor so requests can change chunk size safely, while explicit native, llama-compatible, and environment settings continue to pin both states. The change also adds an allocation-free next-piece calculation for repeated paged-block capacity checks, propagates the two-value policy through both server front ends, and documents the prompt-cache partition implications and measured tradeoff. ## Validation - CUDA core prefill-plan tests: 17 passed. - CUDA server CLI tests: 148 passed; runtime-settings tests: 8 passed; props compatibility: passed. - CUDA block-reclaim tests: 20 passed. - Actual scheduler regression passed exact forward sequences `[512, 512, 512, 1]`, `[2048, 952]`, and `[512, 2048, 2048, 392]`, and required request completion. - Shared-capability source audit: 3 passed, including a negative mutation fixture. - Workspace all-targets Clippy, `cargo check --profile test-fast --features cuda --lib`, formatting, and diff checks passed. - Fresh-release admission p95: 54.1, 54.8, and 55.2 ms; all below 62.0 ms. - Five-round idle-server TTFT passed on Qwen3-1.7B and Llama-3.2-1B: each adaptive-policy median stayed inside its explicit-2048 null range. - After PR #2270 and PR #2269 merged, workspace test compilation, all-targets Clippy, formatting, all three contract suites, and `make verify-test-cuda` passed with zero failed CUDA summaries, including the BF16 and F32 SSM update parity tests. Closes #2228
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue #2230 left unused paged compatibility APIs, an unresolved model-owned storage default, and a CUDA bf16 SSM recurrence whose compiled preprocessing rounded
dt'to bf16.Changes
dtanddt_biasto float32 in the shared compiled preprocessing before Metal/CUDA/HIP dispatch. The pinned MLX promotion table had selected bf16 for float32 plus bf16; no tolerance changed.supports_paged_decode_backend() == true.Measured evidence
For Granite seed 2069, normalized RMS/max fell from
1.957140831e-3 / 5.875451502e-3to2.890811210e-8 / 2.002812756e-7at compileddt', from3.008783834e-3 / 1.075881018e-2to6.809246822e-10 / 4.717584582e-9at decay, and from2.532402801e-3 / 4.127457397e-2to2.838613609e-8 / 1.701009653e-6at state.The first Llama 4 512-token/concurrency-1 paged arm was an observed anomaly: 6.5 aggregate tok/s versus 11.9 dense and 11.8 repeated-paged, yielding +83.08% paired and +81.54% null deltas. Every record is retained; its cause is unproven, and the preregistered rule leaves the cell unresolved.
Validation
TECHNICAL_REPORTS/for PR fix: clean up paged decode and SSM precision #2270.Closes #2230