Skip to content

fix: clean up paged decode and SSM precision - #2270

Merged
inureyes merged 7 commits into
mainfrom
update/issue-2230-post-epic-cleanup
Oct 11, 2026
Merged

inureyes merged 7 commits into
mainfrom
update/issue-2230-post-epic-cleanup

Conversation

@inureyes

@inureyes inureyes commented Oct 11, 2026 •

Copy link
Copy Markdown
Member

Issue #2230 left unused paged compatibility APIs, an unresolved model-owned storage default, and a CUDA bf16 SSM recurrence whose compiled preprocessing rounded dt' to bf16.

Changes

  • Remove the dense-pointer and rotating paged compatibility surface, cache metadata, validators, and self-referential tests. Four pooled decode tests now use an independent explicit float32 attention oracle, retain the original absolute RMS bound, and add normalized RMS coverage.
  • Cast both dt and dt_bias to float32 in the shared compiled preprocessing before Metal/CUDA/HIP dispatch. The pinned MLX promotion table had selected bf16 for float32 plus bf16; no tolerance changed.
  • Commit a fail-closed paired/null storage driver, the complete 60-arm/180-cell GB10 sweep, exact raw output and provenance, and dated guidance. No cell cleared its repeated-paged null spread, so Gemma 3 and Llama 4 keep supports_paged_decode_backend() == true.

Measured evidence

For Granite seed 2069, normalized RMS/max fell from 1.957140831e-3 / 5.875451502e-3 to 2.890811210e-8 / 2.002812756e-7 at compiled dt', from 3.008783834e-3 / 1.075881018e-2 to 6.809246822e-10 / 4.717584582e-9 at decay, and from 2.532402801e-3 / 4.127457397e-2 to 2.838613609e-8 / 1.701009653e-6 at state.

The first Llama 4 512-token/concurrency-1 paged arm was an observed anomaly: 6.5 aggregate tok/s versus 11.9 dense and 11.8 repeated-paged, yielding +83.08% paired and +81.54% null deltas. Every record is retained; its cause is unproven, and the preregistered rule leaves the cell unresolved.

Validation

  • Required removal search is empty; four pooled CUDA tests pass, and a temporary 1.01 output mutation made all four fail.
  • Five focused CUDA SSM parity cases pass. Granite fused-kernel and graph paths produced identical 64/64 greedy IDs under deterministic SDPA.
  • Workspace/all-target CUDA clippy, formatting, the three fast contracts, release build, and the full GPU-locked CUDA gate pass; the full gate reports zero failures.
  • The measured analyzer validates exactly 60 arms and 180 successful cells; nine protocol fixtures, both artifact manifests, and the branch-wide whitespace check pass.
  • Canonical English and Korean technical reports are included under TECHNICAL_REPORTS/ for PR fix: clean up paged decode and SSM precision #2270.

Closes #2230

Remove the unused dense and rotating paged compatibility paths while preserving pooled decode coverage with an explicit float32 dense oracle. The oracle retains its absolute RMS checks and adds normalized RMS checks so a one-percent output perturbation is rejected.

Promote both dt operands before the compiled SSM softplus so bf16 activation parameters cannot round the float32 recurrence. Add the paired/null serving benchmark driver and analyzer for the pending Gemma 3 and Llama 4 storage decision.

Refs #2230
Record both the configured MLX pin and the release build's actual MLX checkout, rejecting a mismatch before measurement. Check that the serving port is free and reuse the repository hostgate before every fresh-server arm, preserving per-arm wait, CI, process, memory, load, and driver evidence for the multi-hour run.

Refs #2230
Retry the host gate when activity appears after the quiet wait, and refuse to record any nonquiet arm. Require complete provenance, exact ordered schedules and logs, quiet evidence, and backend markers before computing the paired/null storage decision.

Add lightweight fixtures for valid evidence plus missing, dirty, reordered, nonmonotonic, and substituted protocol artifacts.

Refs #2230
Record the complete 60-arm paired/null storage sweep and retain paged storage for Gemma 3 and Llama 4 because no cell clears its repeated-paged null spread. Preserve the raw host-gate, client, server, schedule, artifact, derived binary-source provenance, SSM stage, and Granite checkpoint evidence with checksums.

Cite the dated decision from both model flags and the continuous-batching guide. The original measured scripts and their recorded hashes remain unchanged.

Validation: analyzer accepted all 60 arms and 180 cells; nine protocol fixtures passed; both artifact manifests verified.

Refs #2230
Commit the release-build transcript referenced by the storage sweep's derived binary-source proof. Clarify that the transcript proves the successful 19m43s build but does not print HEAD; the production-source association remains reconstructed from session context, artifact identity, ancestry, and the protocol-only intervening diff.

Validation: the release-log SHA-256 matches the provenance record and the refreshed evidence manifest verifies.

Refs #2230
Preserve the measured client, SSM stage, and shell-command records byte-for-byte while disabling only their known blank-EOF or trailing-space checks. Refresh both evidence manifests to cover the scoped attributes.

Validation: both manifests verify and the branch diff passes Git's whitespace check.

Refs #2230
@inureyes inureyes added status:review Under review type:chore Maintenance tasks (build, CI, etc.) priority:low Low priority area:models Model architectures, weights, loading, metadata area:core mlxcel-core: MLX FFI, primitives, KV cache, layers labels Oct 11, 2026
Record the compatibility removal, paired/null storage decision, SSM precision diagnosis, raw provenance boundary, and final CUDA/Granite validation in English and Korean.

Refs #2230, #2270
@inureyes inureyes added status:done Completed and removed status:review Under review labels Oct 11, 2026
@inureyes
inureyes merged commit 8dcc9c4 into main Oct 11, 2026
77 of 79 checks passed
@inureyes
inureyes deleted the update/issue-2230-post-epic-cleanup branch October 11, 2026 07:53
inureyes added a commit that referenced this pull request Oct 11, 2026
Choose the server prefill chunk from live scheduler state: 2048 tokens while no other sequence is decoding and 512 while a decode batch is active. Rebuild parked continuations from their live cursor so requests can change chunk size safely, while explicit native, llama-compatible, and environment settings continue to pin both states.

The change also adds an allocation-free next-piece calculation for repeated paged-block capacity checks, propagates the two-value policy through both server front ends, and documents the prompt-cache partition implications and measured tradeoff.

## Validation

- CUDA core prefill-plan tests: 17 passed.
- CUDA server CLI tests: 148 passed; runtime-settings tests: 8 passed; props compatibility: passed.
- CUDA block-reclaim tests: 20 passed.
- Actual scheduler regression passed exact forward sequences `[512, 512, 512, 1]`, `[2048, 952]`, and `[512, 2048, 2048, 392]`, and required request completion.
- Shared-capability source audit: 3 passed, including a negative mutation fixture.
- Workspace all-targets Clippy, `cargo check --profile test-fast --features cuda --lib`, formatting, and diff checks passed.
- Fresh-release admission p95: 54.1, 54.8, and 55.2 ms; all below 62.0 ms.
- Five-round idle-server TTFT passed on Qwen3-1.7B and Llama-3.2-1B: each adaptive-policy median stayed inside its explicit-2048 null range.
- After PR #2270 and PR #2269 merged, workspace test compilation, all-targets Clippy, formatting, all three contract suites, and `make verify-test-cuda` passed with zero failed CUDA summaries, including the BF16 and F32 SSM update parity tests.

Closes #2228
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:core mlxcel-core: MLX FFI, primitives, KV cache, layers area:models Model architectures, weights, loading, metadata priority:low Low priority status:done Completed type:chore Maintenance tasks (build, CI, etc.)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

chore: post-epic cleanup of compat paged kernels, the model-owned paged flag, and the CUDA SSM bf16 parity failures

1 participant