Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
# Technical Report: PR #2269 — Contain Decode Readback Failures and Preserve Sampler State

**Date**: 2026-10-11

**Status**: Implementation and required local validation completed. CI remains pending at report preparation.

**Languages**: Rust, C++

**Risk Level**: High. Changes cross the lazy GPU graph, CXX error boundary, speculative KV teardown, and mutable sampler/RNG state.

## Executive Summary

Decode host reads now wait through fallible CXX entrypoints so backend evaluation errors can fail the affected request instead of escaping an infallible bridge. Failed lookahead submissions commit their already-drawn pending token before synchronous fallback, while failed row draws restore mirostat and adaptive-p feedback.

Seeded fused parity additionally requires restoring the RNG key consumed while constructing the failed next submission. A snapshot of the pinned MLX thread-local key sequence supplies that rollback without adding a GPU evaluation or scheduling call. Focused fault tests, the broader suites and the newly strengthened stopping-boundary mutation checks passed; the restored implementation passed all 13 fault regressions again.

## 1. Problem Statement

An asynchronous backend error may first surface when the host waits for a token, selected probability, or sorted sampler distribution. Several decode paths reached non-`Result` CXX reads before a fallible wait, allowing a C++ exception to terminate the process. Guarding only the final synchronous token read did not protect mirostat or adaptive-p, which read earlier inside sampling.

A second fault was state advancement after speculative failure. Dropping a pending token and sampling it again could consume another random key or advance feedback twice. The emitted stream could diverge from force-sync execution even though the request appeared to recover. Retiring the wrong number of speculative appends could also leave the pool offset and KV length inconsistent with the committed prefix.

## 2. Technical Review

The change confines evaluation errors to existing request-level error and teardown paths. Scheduler read failures record an evaluation failure and enter the established finishing teardown before retrying through guarded synchronous execution. This scheduler branch is reviewed in code, as the core test-only fault seam is not available to the server crate.

Token waits use per-array `try_eval`, preserving overlap with the next in-flight forward. Feedback snapshots copy only mirostat/adaptive-p state, not penalty history caches. The fused stochastic path adds a host-side snapshot allocation and key-array ownership copy; it adds no evaluation or scheduling operation. No successful-path throughput result is claimed in this draft.

Existing infallible public sampler entrypoints retain their signatures and wrap the fallible core with the documented panic message. Engine sampling uses the new fallible entrypoints. There are no new external dependencies or data migrations. CUDA is the locally validated backend; Metal and ROCm runtime validation is not claimed.

## 3. Technical Decisions

### Guard the actual host-read boundary

`try_tokens_to_host` and `try_read_token` first perform the fallible per-array wait. Adaptive-p and mirostat guard every required evaluation/copy and return errors before publishing feedback updates. `DrawError::Mask` retains structured-output classification, while `DrawError::Eval` becomes an evaluation error.

Eagerly evaluating all logits before sampling was rejected because it would serialize penalty and DRY draws behind the forward. It also would not cover every synchronous sampler entrypoint.

### Commit pending state before fallback

Both B=1 pipelines share their normal read-and-commit helper with the failed-submit branch. That branch unwinds only the failed submission's append, commits `pending` with no `next`, and enters synchronous fallback only if generation continues. EOS and callback stops therefore use the same stopping and cache-retirement rules as normal iterations. A readback failure retires the pending append plus any existing next append.

### Restore only the failed fused draw's RNG advancement

Fused graph construction consumes a random key before `schedule`. Moving the injected fault earlier would conceal the real ordering defect; adding another schedule would risk successful-path throughput. The selected implementation snapshots `random::KeySequence::default_()` after forward construction and restores it only when the existing submission schedule fails. Successful submissions keep their key advancement.

This relies on the public copy construction/assignment of `KeySequence` in MLX pin `81ba1c6a0e50a9268b931579c2d4f1158b9aab5a`. The CXX snapshot owns its copied key array; its lifetime stays inside one synchronous call on the owning thread. A regression verifies restoration and thread isolation. Per-row failed draws retain the documented limitation that a discarded stochastic draw can shift the seeded stream by its consumed key.

## 4. Implementation Details

| Area | Result |
|---|---|
| Fused, per-row, plain-drafter and scheduler token reads | Backend read errors follow fallible request-error paths |
| Mirostat v1/v2/greedy and adaptive-p | All required sampler host reads propagate errors |
| In-loop submit failure | Retire failed append, commit pending token, then conditionally fall back |
| Row draw failure | Restore mirostat/adaptive-p feedback before caller teardown |
| Fused RNG | Restore failed submission's thread-local key without an extra schedule |
| Fault seam | One-shot, thread-local Submit/Draw/Readback/SamplerRead plans; constant `Ok(())` outside tests |

The production fault seam is `#[inline(always)]` and returns `Ok(())` under `cfg(not(test))`; it has no runtime fault plan. This conclusion is based on source configuration, not release-binary disassembly.

## 5. Learning Points

A lazy graph can mutate sampler state before GPU execution succeeds. Correct recovery must distinguish graph construction, scheduling, readback, and token commitment. A final output comparison alone can miss a discarded greedy draw, because redrawing may produce the same token; seeded failure tests and targeted mutations expose that gap.

## 7. Change Summary

Before report additions, the branch changes 16 files with 868 insertions and 116 deletions relative to `24e7e7ff`. Implementation commits are `6f3c384b` (readback/state handling), `4669eeec` (seeded stopping boundaries), and `7facb4fa` (sampler output type aliases).

## 8. Follow-up Actions

- Confirm PR CI results before central merge.
- Keep the separate plain-drafter failed-submit/retraction problem outside this change, as specified by #2258.

## Appendix: Validation Recorded So Far

| Check | Observed result |
|---|---|
| Formatting and diff checks | Passed |
| Core CUDA lib/tests clippy, `-D warnings` | Passed after factoring three complex return types into aliases |
| Core CUDA test-fast test build | Passed |
| Original engine fault suite | 11 passed, GPU lock held, one test thread |
| Plain-drafter readback fault | 1 passed |
| Thread-local fault-plan regression | 1 passed |
| Remove pending commitment | Four separate seeded fused/mirostat/adaptive/penalty stream tests failed as expected |
| Remove feedback restoration | Feedback regression failed as expected |
| Remove fused RNG restoration | Seeded fused regression failed as expected |
| Restored implementation after mutations | All 13 fault regressions passed again, including seeded EOS/callback tests |
| New seeded EOS/callback tests | Both passed in the engine suite; each failed when pending commitment was removed independently from either pipeline (four negative checks) |
| Full engine/sampling/server suites and contracts | Engine 63 passed; sampling 202 passed; server batch 526 passed, 9 ignored; three contracts passed |
| Successful-path throughput measurement | Not measured for this draft |

Focused evidence is in `/tmp/test2258-final.log`, `/tmp/plain2258.log`, `/tmp/seam2258.log`, and `/tmp/mutation2258-{discard-pending,feedback,fused-rng}.log`. Central verification writes `/tmp/mlxcel-auto-plan/2258-*.log`. Stopping-boundary mutations are recorded in `/tmp/mutation2258-stops-summary.log` and `/tmp/mutation2258-pending-{eos,callback}-{fused,row}.log`; the final positive rerun is `/tmp/test2258-restored-stops.log`. These are session-local evidence paths, not committed artifacts.

Related work: issue #2258; request-local backend error contract #822; per-row pipeline #2256 and #2229.
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
# 기술 보고서: PR #2269 — 디코드 읽기 오류 격리와 샘플러 상태 보존

**작성일**: 2026-10-11

**상태**: 구현 및 필수 로컬 검증 완료. 보고서 작성 시점에 CI는 진행 중이다.

**언어**: Rust, C++

**위험도**: 높음. 지연 GPU 그래프, CXX 오류 경계, 추측 실행 KV 정리, 변경 가능한 샘플러/RNG 상태를 함께 다룬다.

## 요약

디코드의 호스트 읽기가 실패를 반환할 수 있는 CXX 진입점을 통해 기다리도록 바꿨다. 백엔드 평가 오류가 실패를 표현하지 못하는 브리지를 넘어가는 대신 해당 요청의 오류로 전달된다. 미리 실행한 제출이 실패하면 이미 추출한 pending 토큰을 커밋한 뒤 동기 실행으로 전환하며, 행별 추출 실패 시 mirostat와 adaptive-p 피드백을 복원한다.

시드가 있는 fused 실행의 동등성을 유지하려면 실패한 다음 제출을 구성하면서 소비한 RNG 키도 복원해야 한다. 고정 MLX 버전의 스레드별 키 시퀀스를 스냅샷으로 보관해 GPU 평가나 스케줄 호출을 추가하지 않고 복구한다. 집중 오류 테스트, 전체 검증 및 새로 강화한 종료 경계 변이 검증은 통과했으며, 복원한 구현은 오류 회귀 테스트 13개를 모두 재통과했다.

## 1. 문제 정의

비동기 백엔드 오류는 호스트가 토큰, 선택 확률, 정렬된 샘플러 분포를 기다릴 때 처음 드러날 수 있다. 일부 디코드 경로는 실패 가능한 대기 전에 `Result`가 아닌 CXX 읽기를 호출해 C++ 예외가 프로세스를 종료할 수 있었다. 최종 동기 토큰 읽기만 보호하면 샘플링 내부에서 먼저 읽는 mirostat와 adaptive-p는 보호되지 않는다.

추측 실행 실패 후 상태가 중복 진행되는 문제도 있었다. pending 토큰을 버리고 다시 추출하면 무작위 키를 한 번 더 소비하거나 피드백을 두 번 진행할 수 있다. 요청이 복구된 것처럼 보여도 force-sync 실행과 출력 스트림이 달라질 수 있었다. 추측 실행으로 추가한 KV 항목을 잘못된 개수만큼 정리하면 풀 오프셋과 KV 길이도 커밋된 접두 구간과 어긋날 수 있다.

## 2. 기술적 검토 사항

평가 오류는 기존 요청 단위 오류와 정리 경로로 전달한다. 스케줄러 읽기 실패는 평가 실패를 기록하고 기존 종료 정리 경로에 진입한 다음 보호된 동기 실행으로 재시도한다. 코어의 테스트 전용 오류 주입 기능은 서버 크레이트에서 사용할 수 없으므로 이 스케줄러 분기는 코드 검토로 확인한다.

토큰 대기는 배열별 `try_eval`을 사용해 다음 forward와의 중첩을 유지한다. 피드백 스냅샷은 mirostat/adaptive-p 상태만 복사하며 페널티 이력 캐시는 복사하지 않는다. fused 확률적 경로에는 호스트 스냅샷 할당과 키 배열 소유권 복사가 추가되지만 평가나 스케줄 호출은 추가되지 않는다. 이 초안은 정상 경로 처리량 측정 결과를 주장하지 않는다.

기존 infallible 공개 샘플러 진입점은 시그니처를 유지하고 정해진 panic 메시지로 fallible 코어를 감싼다. 엔진 샘플링은 새 fallible 진입점을 사용한다. 외부 의존성이나 데이터 마이그레이션은 추가하지 않았다. 로컬에서 검증한 백엔드는 CUDA이며 Metal 및 ROCm 실행 검증을 수행했다고 주장하지 않는다.

## 3. 기술적 선택과 그 이유

### 실제 호스트 읽기 경계 보호

`try_tokens_to_host`와 `try_read_token`은 먼저 실패 가능한 배열별 대기를 수행한다. Adaptive-p와 mirostat는 필요한 모든 평가 및 복사를 보호하고 피드백 변경을 게시하기 전에 오류를 반환한다. `DrawError::Mask`는 구조화 출력 오류 분류를 유지하고 `DrawError::Eval`은 평가 오류로 전달한다.

샘플링 전에 모든 logits를 평가하는 방법은 페널티와 DRY 추출을 forward 뒤로 직렬화하므로 선택하지 않았다. 이 방법으로 모든 동기 샘플러 진입점을 보호할 수도 없다.

### 동기 전환 전에 pending 상태 커밋

두 B=1 파이프라인은 정상 반복과 제출 실패 분기에서 같은 읽기 및 커밋 헬퍼를 사용한다. 실패 분기는 해당 제출이 추가한 항목만 되돌리고 `next` 없이 pending을 커밋한 뒤, 생성이 계속될 때만 동기 실행으로 전환한다. 따라서 EOS와 콜백 중지는 정상 반복과 같은 종료 및 캐시 정리 규칙을 따른다. 읽기 실패 시 pending 항목과 이미 존재하는 next 항목을 함께 정리한다.

### 실패한 fused 추출의 RNG 진행만 복원

Fused 그래프 구성은 `schedule` 전에 무작위 키를 소비한다. 주입 오류 위치를 앞으로 옮기면 실제 순서 문제를 숨기고, 스케줄 호출을 하나 더 넣으면 정상 경로 처리량에 영향을 줄 수 있다. 선택한 구현은 forward 구성 후 `random::KeySequence::default_()`를 스냅샷으로 보관하고 기존 제출 스케줄이 실패한 경우에만 복원한다. 성공한 제출은 키 진행을 유지한다.

이 구현은 MLX pin `81ba1c6a0e50a9268b931579c2d4f1158b9aab5a`의 `KeySequence`가 제공하는 공개 복사 생성과 대입에 의존한다. CXX 스냅샷은 복사한 키 배열을 소유하며, 수명은 해당 스레드의 동기 호출 하나 안에 한정된다. 회귀 테스트는 복원과 스레드 격리를 확인한다. 행별 추출 실패에는 버린 확률적 추출이 소비한 키만큼 시드 스트림이 이동할 수 있다는 기존 문서화된 제한이 남는다.

## 4. 구현 상세

| 영역 | 결과 |
|---|---|
| Fused, 행별, plain-drafter, 스케줄러 토큰 읽기 | 백엔드 읽기 오류를 fallible 요청 오류 경로로 전달 |
| Mirostat v1/v2/greedy 및 adaptive-p | 필요한 모든 샘플러 호스트 읽기에서 오류 전달 |
| 반복 중 제출 실패 | 실패한 추가 항목 정리, pending 토큰 커밋, 조건부 동기 전환 |
| 행별 추출 실패 | 호출 측 정리 전에 mirostat/adaptive-p 피드백 복원 |
| Fused RNG | 추가 스케줄 없이 실패한 제출의 스레드별 키 복원 |
| 오류 주입 | 스레드별 일회성 Submit/Draw/Readback/SamplerRead 계획; 테스트 외에는 상수 `Ok(())` |

프로덕션 오류 주입 진입점은 `#[inline(always)]`이며 `cfg(not(test))`에서 `Ok(())`를 반환하므로 실행 시 오류 계획이 없다. 이는 소스 설정으로 확인한 결과이며 릴리스 바이너리 역어셈블리 결과는 아니다.

## 5. 학습 포인트

지연 그래프는 GPU 실행 성공 전에 샘플러 상태를 변경할 수 있다. 올바른 복구를 위해 그래프 구성, 스케줄, 읽기, 토큰 커밋을 구분해야 한다. 최종 출력 비교만으로는 버린 greedy 추출을 발견하지 못할 수 있다. 다시 추출해도 같은 토큰이 나올 수 있기 때문이다. 시드 기반 실패 테스트와 대상 변이 검증으로 이 공백을 확인할 수 있다.

## 7. 변경 요약

보고서 추가 전 기준으로 `24e7e7ff` 대비 16개 파일, 868줄 추가, 116줄 삭제다. 구현 커밋은 `6f3c384b`(읽기 및 상태 처리), `4669eeec`(시드 기반 종료 경계), `7facb4fa`(샘플러 출력 타입 별칭)다.

## 8. 후속 조치

- 중앙 병합 전에 PR CI 결과를 확인한다.
- #2258에서 지정한 대로 별도의 plain-drafter 제출 실패 및 draft 철회 문제는 이번 변경 범위 밖에 둔다.

## 부록: 현재까지 기록한 검증

| 검사 | 확인한 결과 |
|---|---|
| 포맷 및 diff 검사 | 통과 |
| 코어 CUDA lib/tests clippy, `-D warnings` | 복잡한 반환 타입 세 곳을 별칭으로 정리한 뒤 통과 |
| 코어 CUDA test-fast 테스트 빌드 | 통과 |
| 기존 엔진 오류 테스트 묶음 | GPU 잠금과 단일 테스트 스레드로 11개 통과 |
| Plain-drafter 읽기 오류 | 1개 통과 |
| 스레드별 오류 계획 회귀 | 1개 통과 |
| Pending 커밋 제거 | 별도 시드 fused/mirostat/adaptive/penalty 스트림 테스트 4개가 예상대로 실패 |
| 피드백 복원 제거 | 피드백 회귀 테스트가 예상대로 실패 |
| Fused RNG 복원 제거 | 시드 fused 회귀 테스트가 예상대로 실패 |
| 변이 후 구현 복원 | 시드 EOS/콜백 테스트를 포함한 오류 회귀 테스트 13개 모두 재통과 |
| 새 시드 EOS/콜백 테스트 | 엔진 테스트에서 둘 다 통과; 각 파이프라인의 pending 커밋을 독립적으로 제거했을 때 각각 실패(음성 검증 4건) |
| 전체 엔진/샘플링/서버 테스트와 계약 검사 | 엔진 63개 통과, 샘플링 202개 통과, 서버 배치 526개 통과 및 9개 무시, 계약 검사 3개 통과 |
| 정상 경로 처리량 측정 | 이 초안에서 측정하지 않음 |

집중 검증 자료는 `/tmp/test2258-final.log`, `/tmp/plain2258.log`, `/tmp/seam2258.log`, `/tmp/mutation2258-{discard-pending,feedback,fused-rng}.log`에 있다. 중앙 검증은 `/tmp/mlxcel-auto-plan/2258-*.log`에 기록한다. 종료 경계 변이는 `/tmp/mutation2258-stops-summary.log`와 `/tmp/mutation2258-pending-{eos,callback}-{fused,row}.log`, 마지막 정상 재실행은 `/tmp/test2258-restored-stops.log`에 기록했다. 이 경로들은 세션 내 검증 자료이며 커밋된 산출물이 아니다.

관련 작업: 이슈 #2258, 요청별 백엔드 오류 계약 #822, 행별 파이프라인 #2256 및 #2229.
9 changes: 9 additions & 0 deletions src/lib/mlxcel-core/cpp/mlx_cxx_bridge.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -51,6 +51,15 @@ namespace mlx_cxx {

using namespace mlx::core;

std::unique_ptr<MlxRngSnapshot> rng_snapshot() {
return std::make_unique<MlxRngSnapshot>();
}

void restore_rng_snapshot(const MlxRngSnapshot& snapshot) {
random::KeySequence::default_() = snapshot.inner;
}


namespace {

using CompiledArrayFn =
Expand Down
10 changes: 10 additions & 0 deletions src/lib/mlxcel-core/cpp/mlx_cxx_bridge.h
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,16 @@

namespace mlx_cxx {

// A copy of the pinned MLX KeySequence retains only its lazy key array.
// Its public copy/assignment operations let failed fused submits restore the
// calling thread's RNG without evaluating a graph or adding a GPU schedule.
struct MlxRngSnapshot {
mlx::core::random::KeySequence inner;
MlxRngSnapshot() : inner(mlx::core::random::KeySequence::default_()) {}
};
std::unique_ptr<MlxRngSnapshot> rng_snapshot();
void restore_rng_snapshot(const MlxRngSnapshot& snapshot);

// Opaque wrapper struct to hold mlx::core::array
// This allows cxx to manage the lifetime without exposing the complex internals
struct MlxArray {
Expand Down
Loading
Loading