Skip to content

fix: contain decode readback failures and preserve sampler state - #2269

Merged
inureyes merged 5 commits into
mainfrom
fix/issue-2258-decode-readback-errors
Oct 11, 2026
Merged

inureyes merged 5 commits into
mainfrom
fix/issue-2258-decode-readback-errors

Conversation

@inureyes

@inureyes inureyes commented Oct 11, 2026 •

Copy link
Copy Markdown
Member

Summary

Decode host reads and feedback samplers now return backend errors through the fallible CXX boundary. Failed lookahead submits commit their pending token before synchronous fallback, and failed row draws restore mirostat/adaptive-p feedback.

  • Guard fused, per-row, plain-drafter and scheduler readbacks and adaptive-p/mirostat host reads; retain existing infallible sampler signatures for non-engine callers.
  • Restore a failed fused submission's thread-local RNG key using the pinned MLX KeySequence copy operations. This adds no GPU evaluation or scheduling call; successful-path throughput has not been measured. Discarded per-row draws retain their documented seeded-key limitation.
  • Add thread-local faults and regressions for seeded parity, feedback rollback, committed-prefix cache state, EOS/callback teardown, sequence release and thread isolation.
  • Add English and Korean technical reports for the final implementation and validation record.

Test plan

  • CUDA engine: 63 passed; sampling: 202 passed; server batch: 526 passed, 9 ignored. GPU suites ran serialized under gpu-lock.
  • Core lib/tests clippy, formatting and three contract checks passed.
  • Plain-drafter readback and thread-local fault tests passed.
  • Mutations removing pending commitment, feedback restoration or fused RNG restoration failed the corresponding regressions. Seeded EOS and callback mutations failed separately for each pipeline.
  • Restored implementation: all 13 fault regressions passed again.
  • GitHub Actions CI passed on the final commit.

Closes #2258

@inureyes inureyes added status:in-progress Currently being worked on type:bug Bug fixes, error corrections, or issue resolutions priority:medium Medium priority area:inference Generation, sampling, decoding (incl. speculative, DRY) labels Oct 11, 2026
inureyes added a commit that referenced this pull request Oct 11, 2026
inureyes added a commit that referenced this pull request Oct 11, 2026
Remove the completed report-finalization follow-up while retaining the unverified CI status and the deferred plain-drafter scope boundary in both languages.

Validation:
- Documentation diff inspected; implementation formatting, clippy, contracts and required test suites passed before this report-only update.

Refs #2258
@inureyes inureyes added status:review Under review status:done Completed and removed status:in-progress Currently being worked on status:review Under review labels Oct 11, 2026
Route decode and feedback-sampler host waits through fallible evaluation boundaries, classify sampler failures as row evaluation errors, and unwind failed reads through existing sequence teardown.

Commit pending draws before retrying failed lookahead submits and restore feedback after failed row draws. Preserve fused sampling keys with a scoped thread-local MLX RNG snapshot, keeping the existing single asynchronous schedule.

Add injected submit, draw, sampler-read and readback regressions covering seeded parity, feedback rollback, committed-prefix KV state, discarded sequence release and thread isolation. Mutation checks detect removal of pending commits, feedback restoration and RNG restoration.

Refs #2258.
Greedy EOS and callback traces can still match after discarding and redrawing a pending token. Exercise seeded fused and penalty streams with a failed submit immediately before the pending stop is committed, comparing callback deliveries and cache state with the synchronous run.

Formatting and diff checks pass. Runtime and focused mutation checks are pending the coordinated CUDA validation window.

Refs #2258
Factor the sampler tuple outputs into aliases so fallible distribution and core APIs pass clippy's type-complexity gate without changing their values or error behavior.

Formatting and diff checks pass; central CUDA clippy and runtime verification will rerun with the aliases.

Refs #2258
Remove the completed report-finalization follow-up while retaining the unverified CI status and the deferred plain-drafter scope boundary in both languages.

Validation:
- Documentation diff inspected; implementation formatting, clippy, contracts and required test suites passed before this report-only update.

Refs #2258
@inureyes
inureyes force-pushed the fix/issue-2258-decode-readback-errors branch from 5c504d1 to 9caca41 Compare October 11, 2026 07:53
@inureyes
inureyes merged commit 02cbd82 into main Oct 11, 2026
27 checks passed
@inureyes
inureyes deleted the fix/issue-2258-decode-readback-errors branch October 11, 2026 08:12
inureyes added a commit that referenced this pull request Oct 11, 2026
Choose the server prefill chunk from live scheduler state: 2048 tokens while no other sequence is decoding and 512 while a decode batch is active. Rebuild parked continuations from their live cursor so requests can change chunk size safely, while explicit native, llama-compatible, and environment settings continue to pin both states.

The change also adds an allocation-free next-piece calculation for repeated paged-block capacity checks, propagates the two-value policy through both server front ends, and documents the prompt-cache partition implications and measured tradeoff.

## Validation

- CUDA core prefill-plan tests: 17 passed.
- CUDA server CLI tests: 148 passed; runtime-settings tests: 8 passed; props compatibility: passed.
- CUDA block-reclaim tests: 20 passed.
- Actual scheduler regression passed exact forward sequences `[512, 512, 512, 1]`, `[2048, 952]`, and `[512, 2048, 2048, 392]`, and required request completion.
- Shared-capability source audit: 3 passed, including a negative mutation fixture.
- Workspace all-targets Clippy, `cargo check --profile test-fast --features cuda --lib`, formatting, and diff checks passed.
- Fresh-release admission p95: 54.1, 54.8, and 55.2 ms; all below 62.0 ms.
- Five-round idle-server TTFT passed on Qwen3-1.7B and Llama-3.2-1B: each adaptive-policy median stayed inside its explicit-2048 null range.
- After PR #2270 and PR #2269 merged, workspace test compilation, all-targets Clippy, formatting, all three contract suites, and `make verify-test-cuda` passed with zero failed CUDA summaries, including the BF16 and F32 SSM update parity tests.

Closes #2228
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:inference Generation, sampling, decoding (incl. speculative, DRY) priority:medium Medium priority status:done Completed type:bug Bug fixes, error corrections, or issue resolutions

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix: surface backend errors at every decode host readback instead of aborting

1 participant