diff --git a/TECHNICAL_REPORTS/2268-adaptive-prefill-chunks-20261011.en.md b/TECHNICAL_REPORTS/2268-adaptive-prefill-chunks-20261011.en.md new file mode 100644 index 000000000..4c86dc878 --- /dev/null +++ b/TECHNICAL_REPORTS/2268-adaptive-prefill-chunks-20261011.en.md @@ -0,0 +1,136 @@ +# PR #2268: Adaptive Server Prefill Chunks + +**Date**: 2026-10-11 +**Status**: Complete; performance acceptance and full integration validation passed +**Languages**: Rust, Markdown +**Risk**: Medium (changes live batch scheduling, prefill partitioning, and CLI precedence) + +## Executive Summary + +PR #2268 changes server prefill chunking from one fixed value to a scheduler-state policy. A request uses 2048-token chunks while no other sequence is decoding and 512-token chunks while the active batch contains decode work. Explicit native, llama-compatible, or environment configuration pins both states to the requested value. + +The implementation also makes prefill plans resumable from their live cursor. This is required because a parked request can start under contention at a 512-token boundary and resume after the active batch drains, when a newly built 2048-token partition would otherwise have no piece at that cursor. Focused CUDA tests cover the actual forward lengths and completion behavior. Fresh-release quiet-host admission and TTFT acceptance passed on GB10, and the full integrated workspace and CUDA gates passed after the two preceding PRs merged. + +## 1. Problem Statement + +PR #2205 selected a fixed 2048-token prefill chunk from single-stream TTFT measurements. PR #2226 then measured a conflicting concurrent-serving result: with four streams decoding, admitting an 8192-token request raised live-stream ITL p95 to 127.9–128.9 ms at chunk 2048, while chunk 512 held it to 55.8–56.4 ms. The smaller chunk increased the admitted request's TTFT from 1903–2060 ms to 3768–4067 ms. + +Neither fixed value serves both states well. Keeping 2048 harms active streams during admission; keeping 512 pays the longer TTFT even when no stream needs protection. Selecting a value once at admission is also insufficient because decode work can finish while a long prompt is parked between pieces. + +The existing plan constructor made a live switch unsafe. Its piece boundaries were derived from the original adopted or history boundary, so a cursor reached on the 512 grid was not necessarily a valid start in a freshly constructed 2048 plan. The scheduler treated the missing piece as an abort condition. + +## 2. Technical Decisions + +### 2.1 Select the chunk at every scheduler tick + +PrefillChunkPolicy holds alone and contended values and resolves the current chunk from whether active_batch is empty. The defaults are 2048 and 512. The parked prefill itself is kept separately from active_batch, so it does not incorrectly classify its own continuation as contended. + +An explicit --prefill-chunk-size, its --batch-size / LLAMA_ARG_BATCH alias, or MLXCEL_PREFILL_CHUNK produces {n, n}. This preserves operator control and makes an explicit value equal to the default distinguishable from no flag. Native flag precedence remains above the compatibility alias, which remains above the environment default. + +Fixing the policy per sequence was rejected because a sequence admitted next to a stream would stay on 512 after that stream completed. Switching only new requests was rejected for the same reason. + +### 2.2 Resume from the live cursor + +PrefillPlan::resume_at retains the metadata that a full plan would report, but cuts continuation pieces from the current prefill_offset. The first continuation therefore starts exactly at the cursor regardless of the prior chunk size. Padding, history-boundary behavior, prompt-cache metadata, and model capability checks share the same helpers as the original constructor. + +For a 5000-token request, the tested state transition is: + + active decode present: 0..512 + active batch drained: 512..2560, 2560..4608, 4608..5000 + forward lengths: [512, 2048, 2048, 392] + +The continuation reports chunked progress even when its remaining suffix fits in one piece. This preserves reservation and progress accounting after an earlier piece established the chunked-prefill lifecycle. + +### 2.3 Preserve alone-value behavior outside the batch scheduler + +Direct engine generation, the engine benchmark, and speculative/MTP prefill paths continue to use the 2048-token alone value unless explicitly overridden. The existing batched-prefill token budget is also derived from the alone value; its per-row cap already limits each row to 512 tokens. + +## 3. Implementation Details + +The core planner now defines the two-value policy, a process-wide parsed environment override, shared metadata and piece construction, and cursor-resumable planning. An allocation-free next-piece accessor reuses the same chunk eligibility, boundary, range, and padding helpers for capacity checks that only need the next padded length. The scheduler stores the policy, reads live decode state before each start or continuation, builds the plan at the live cursor for execution, reserves the blocks for the piece it will actually execute, and records the selected chunk in its tracing span. + +CLI fields changed from a defaulted usize to Option, allowing explicit 2048 to remain explicit. Both server binaries use the same precedence and conflict rules. Startup, runtime settings, probes, server configuration, model workers, and scheduler builders carry both policy values. Startup diagnostics report both values. + +Documentation updates explain the adaptive default, override semantics, cursor repartitioning, prompt-cache partition implications, and the PR #2226 tradeoff that motivated the policy. No dependency, persistence schema, authentication, or network protocol changed. + +## 4. Technical Review + +### Correctness and compatibility + +- Chunk 0 still disables chunking in both policy states. +- Prompts shorter than the selected chunk remain one forward. +- Embedding-input/VLM prefills and unsplittable history-boundary segments keep their existing behavior. +- Explicit native, compatibility-alias, and environment values pin both states. +- Existing direct and speculative call sites retain the alone value. +- No new dependencies or public response fields were added. + +The highest-risk path is a chunk-size transition after a request has parked. The integration regression executes real scheduler prefill calls against the tiny test model, records lengths immediately before Engine::prefill, consumes every continuation, rejects error or timeout events, and requires the request to emit Done. + +### Performance + +The policy is intended to recover PR #2226's 512-token admission behavior while retaining 2048-token single-stream TTFT. Execution planning is linear in the pieces it will run and no longer creates and discards a full-prefix vector before rebuilding a suffix. Repeated paged-block capacity checks compute only the next piece in constant space instead of allocating the complete remaining suffix. + +Fresh release measurements passed every acceptance gate: + +| Gate | Acceptance | Result | +|---|---|---| +| Live-stream admission, 3 quiet-host rounds | ITL p95 at or below 62.0 ms in every round | **Passed:** 54.1, 54.8, and 55.2 ms | +| Qwen 3 1.7B server TTFT, 5 rounds | Policy delta inside explicit-2048 null range | **Passed:** policy median +0.01%; null range -4.73% to +3.26% | +| Llama 3.2 1B server TTFT, 5 rounds | Policy delta inside explicit-2048 null range | **Passed:** policy median -0.20%; null range -1.12% to -0.03% | + +### Security + +The change adds no external trust boundary. CLI and environment input use the existing positive-integer-or-zero parser and invalid environment input falls back to the adaptive defaults with a warning. Logs expose only numeric policy values. + +## 5. Verification + +Completed on the implementation worktree: + +- cargo check --profile test-fast --features cuda --lib: passed. +- cargo fmt --all -- --check: passed. +- git diff --check: passed. +- CUDA mlxcel-core prefill-plan tests: 17 passed. +- CUDA server CLI tests: 148 passed. +- CUDA adaptive scheduler execution regression: passed. +- CUDA runtime-settings tests: 8 passed. +- CUDA props compatibility regression: passed. +- Independent security/performance review found that repeated paged-block capacity checks eagerly built every remaining plan piece. Commit 88c2e6a0964d4a3a144c2ff1cc3c80dfd2893752 replaces that hot-path allocation with the O(1) accessor and adds equivalence coverage across boundaries, cursors, chunks, padding modes, model opt-outs, and embedding inputs. +- Workspace all-targets Clippy passed after the fix. Focused CUDA runtime checks passed: 17 prefill-plan tests, 20 block-reclaim tests, and the adaptive scheduler transition regression. +- The first full CUDA gate exposed a source-audit gap after capability-helper extraction. The audit now uses brace-counted interprocedural recognition, retains direct-guard detection, and includes a mutation fixture; its three tests pass. + +The acceptance artifact used production commit `88c2e6a0964d4a3a144c2ff1cc3c80dfd2893752` and measured clean worktree HEAD `b2394566a71fc1f6db83ad40b44380b9c129b57c`. The driver verified that only two test files differed. Raw records, server logs, the exact environment, a machine-readable summary, and the reproducible driver and analyzer are archived under `docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/`. + +The driver's default build commit is valid only on the unsquashed measured lineage. A replay from a squash-merged or newer clean checkout must rebuild both release binaries from that checkout and set `MLXCEL_BUILD_COMMIT="$(git rev-parse HEAD)"`; the benchmark report gives the complete command. + +After PR #2270 and PR #2269 merged, the complete series passed workspace test compilation with `--no-run`, workspace all-targets Clippy, formatting, the kernel dtype-key, kernel port-dispatch, and llama-server compatibility contract suites, and the full single-threaded CUDA workspace run through `make verify-test-cuda`. Every CUDA test summary reported zero failures, including the BF16 and F32 SSM update parity tests. The validated pre-rebase tree `ab8ff77ade8d6b1b4fb494707320cebc95150c8e` remained identical after rebasing onto current `origin/main`; the final changes after that comparison are documentation and PR status only. + +## 6. Change Summary + +| Item | Value | +|---|---| +| Files changed | 47 | +| Lines added | 1879 | +| Lines deleted | 210 | +| Measured branch head | b2394566a71fc1f6db83ad40b44380b9c129b57c | + +| Category | Summary | +|---|---| +| Core planning | Two-state policy, cached override parsing, resumable cursor-based plans | +| Scheduling | Live-state chunk selection, safe continuation/reservation, O(1) capacity checks | +| Configuration | Explicit-value detection and two-value propagation through both server front ends | +| Testing | Planner properties, CLI precedence, schema compatibility, actual scheduler forward sequences | +| Documentation | Adaptive defaults, override behavior, measurement rationale, cache partition note | + +The initial implementation commit is 7c43726a0fb792594a69de0fda1e31e78eef4fa1. Commit 7b2d17241ea9c16612766ccf37524cbe8be9f913 satisfies the CI Clippy expression policy. Commit 88c2e6a0964d4a3a144c2ff1cc3c80dfd2893752 removes the review-discovered reservation allocation, and commit 808bcccf387957a6da31b961666aadd5bf7cbe65 adds the test-only type annotation required by all-targets Clippy. Commits 6f07f7563e74c5466b5aace6f3e8ee98ec8347e2 and b2394566a71fc1f6db83ad40b44380b9c129b57c conservatively update and strengthen the source audit for the shared capability guard. + +## 7. Completion + +No implementation, acceptance, or integration-validation work remains for issue #2228. + +## References + +- Issue #2228: choose the prefill chunk by whether other sequences are decoding. +- PR #2226 and docs/benchmark_results/unified-engine-final-gb10-2026-10-08.md: motivating admission measurements. +- ADR 0007: unified batch/native engine decisions and the amended prefill-chunk policy. +- docs/CONTINUOUS_BATCHING.md: server policy, planning invariants, and operational controls. +- docs/benchmark_results/adaptive-prefill-chunk-gb10-2026-10-11.md: passing acceptance results, raw data, and reproduction protocol. diff --git a/TECHNICAL_REPORTS/2268-adaptive-prefill-chunks-20261011.ko.md b/TECHNICAL_REPORTS/2268-adaptive-prefill-chunks-20261011.ko.md new file mode 100644 index 000000000..07606a4c4 --- /dev/null +++ b/TECHNICAL_REPORTS/2268-adaptive-prefill-chunks-20261011.ko.md @@ -0,0 +1,136 @@ +# PR #2268: 서버 프리필 청크 적응형 선택 + +**작성일**: 2026-10-11 +**상태**: 완료, 성능 승인 및 전체 통합 검증 통과 +**언어**: Rust, Markdown +**위험도**: 중간 (실시간 배치 스케줄링, 프리필 분할, CLI 우선순위 변경) + +## 요약 + +PR #2268은 서버 프리필 청크를 하나의 고정값에서 스케줄러 상태에 따른 정책으로 바꾼다. 다른 시퀀스가 디코딩 중이 아니면 2048토큰 청크를 사용하고, 활성 배치에 디코딩 작업이 있으면 512토큰 청크를 사용한다. native 옵션, llama 호환 옵션 또는 환경 변수로 값을 명시하면 두 상태 모두 해당 값으로 고정한다. + +프리필 계획은 현재 커서에서 다시 시작할 수 있게 되었다. 경합 중인 요청이 512토큰 경계에서 시작한 뒤 활성 배치가 비워졌을 때, 기존 방식으로 새 2048토큰 계획을 만들면 현재 커서에서 시작하는 조각이 없을 수 있기 때문이다. 집중 CUDA 테스트는 실제 forward 길이와 요청 완료를 검증한다. GB10에서 새 release를 사용한 quiet-host admission 및 TTFT 승인이 통과했고, 선행 PR 2건을 merge한 뒤 전체 workspace 및 CUDA 통합 gate도 통과했다. + +## 1. 문제 정의 + +PR #2205는 단일 스트림 TTFT 측정을 근거로 프리필 청크를 2048로 고정했다. 이후 PR #2226은 동시 요청 환경에서 반대 결과를 측정했다. 네 스트림이 디코딩 중일 때 8192토큰 요청을 추가하면 청크 2048에서 기존 스트림의 ITL p95가 127.9–128.9 ms로 상승했지만, 청크 512에서는 55.8–56.4 ms에 머물렀다. 반면 새 요청의 TTFT는 2048에서 1903–2060 ms였고 512에서는 3768–4067 ms로 늘었다. + +어느 고정값도 두 상태를 함께 만족하지 못한다. 2048을 유지하면 새 요청을 받는 동안 기존 스트림이 지연되고, 512를 유지하면 보호할 스트림이 없는 경우에도 긴 TTFT를 지불한다. 긴 프롬프트가 조각 사이에서 대기하는 동안 디코딩 작업이 끝날 수 있으므로 요청 수락 시 한 번만 값을 선택하는 방식도 충분하지 않다. + +기존 계획 생성자는 실행 중 값을 바꾸기에 안전하지 않았다. 조각 경계가 최초 adopted offset 또는 history boundary를 기준으로 정해지므로, 512 격자에서 도달한 커서가 새 2048 계획의 조각 시작점이 아닐 수 있었다. 스케줄러는 이 경우를 남은 조각 없음으로 보고 요청을 중단했다. + +## 2. 기술적 선택과 그 이유 + +### 2.1 매 스케줄러 tick에서 청크 선택 + +PrefillChunkPolicy는 alone과 contended 값을 보관하고 active_batch가 비었는지에 따라 현재 청크를 선택한다. 기본값은 각각 2048과 512다. 대기 중인 프리필은 active_batch와 별도로 보관되므로 자기 자신 때문에 contended 상태로 잘못 판정되지 않는다. + +--prefill-chunk-size, --batch-size / LLAMA_ARG_BATCH 호환 별칭 또는 MLXCEL_PREFILL_CHUNK를 명시하면 {n, n}을 만든다. 따라서 운영자가 정책을 고정할 수 있고, 기본값과 같은 2048을 명시한 경우도 옵션이 없는 경우와 구분한다. 우선순위는 native 옵션, 호환 별칭, 환경 변수 순으로 유지한다. + +요청 단위로 정책을 고정하는 방안은 채택하지 않았다. 디코딩 스트림 옆에서 수락된 요청은 해당 스트림이 끝난 뒤에도 512를 계속 사용해 불필요한 TTFT 비용을 내기 때문이다. 새 요청에만 상태를 적용하는 방식도 같은 문제가 있다. + +### 2.2 현재 커서에서 계획 재개 + +PrefillPlan::resume_at은 전체 계획과 같은 메타데이터를 유지하면서 현재 prefill_offset부터 남은 조각을 자른다. 따라서 이전 청크 크기와 관계없이 첫 continuation은 정확히 현재 커서에서 시작한다. 패딩, history boundary, prompt cache 메타데이터, 모델 기능 검사는 기존 생성자와 같은 helper를 공유한다. + +5000토큰 요청에서 검증한 상태 전환은 다음과 같다. + + 활성 디코딩 존재: 0..512 + 활성 배치 종료: 512..2560, 2560..4608, 4608..5000 + forward 길이: [512, 2048, 2048, 392] + +남은 suffix가 한 조각에 들어가더라도 continuation은 chunked progress를 보고한다. 앞선 조각에서 시작된 chunked-prefill lifecycle의 reservation 및 진행률 계산을 유지하기 위해서다. + +### 2.3 배치 스케줄러 밖에서는 alone 값 유지 + +Direct engine 생성, engine benchmark, speculative/MTP prefill 경로는 명시적 override가 없으면 기존과 같이 2048 alone 값을 사용한다. 기존 batched-prefill token budget도 alone 값에서 계산한다. 각 row는 이미 512토큰으로 제한되어 있다. + +## 3. 구현 상세 + +Core planner에는 두 값 정책, 프로세스 단위 환경 변수 파싱 캐시, 공통 메타데이터 및 조각 생성 helper, 커서 기반 재개 계획을 추가했다. 다음 padded 길이만 필요한 capacity check는 같은 chunk eligibility, boundary, range, padding helper를 재사용하는 allocation-free next-piece accessor를 사용한다. 스케줄러는 정책을 보관하고 각 시작 또는 continuation 직전에 실시간 디코딩 상태를 확인한다. 실행 시 현재 커서에서 계획을 만들고 실제 실행할 조각에 필요한 block을 예약하며 tracing span에 선택한 청크를 기록한다. + +CLI 필드는 기본값이 들어간 usize에서 Option로 바뀌었다. 이로써 2048을 명시한 경우도 명시값으로 유지한다. 두 서버 바이너리는 같은 우선순위 및 충돌 규칙을 사용한다. Startup, runtime settings, probe, server configuration, model worker, scheduler builder가 두 정책값을 전달하며 startup 진단 로그도 두 값을 모두 표시한다. + +문서는 적응형 기본값, override 의미, 커서 재분할, prompt cache partition 영향, 정책의 근거가 된 PR #2226의 trade-off를 설명한다. dependency, 영속 스키마, 인증, 네트워크 protocol 변경은 없다. + +## 4. 기술적 검토 사항 + +### 정확성과 호환성 + +- 청크 0은 두 정책 상태 모두에서 청크 분할을 비활성화한다. +- 선택한 청크보다 짧은 프롬프트는 하나의 forward로 유지된다. +- Embedding-input/VLM prefill과 분할할 수 없는 history-boundary 구간은 기존 동작을 유지한다. +- Native 옵션, 호환 별칭, 환경 변수 명시값은 두 상태를 같은 값으로 고정한다. +- 기존 direct 및 speculative 호출 지점은 alone 값을 유지한다. +- 새 dependency나 공개 응답 field를 추가하지 않았다. + +가장 위험한 경로는 요청이 대기한 뒤 청크 크기가 바뀌는 경우다. Integration regression은 작은 테스트 모델에서 실제 스케줄러 prefill을 실행하고 Engine::prefill 직전에 길이를 기록한다. 모든 continuation을 소비하며 error와 timeout event를 거부하고 최종 Done event를 요구한다. + +### 성능 + +이 정책은 PR #2226에서 확인한 512토큰 admission 성능을 회복하면서 2048토큰의 단일 스트림 TTFT를 유지하는 것이 목적이다. 실행 계획은 실제 실행할 조각 수에 선형이며, 전체 prefix vector를 만든 뒤 버리고 suffix를 다시 만드는 할당을 제거했다. 반복되는 paged-block capacity check는 남은 suffix 전체를 할당하지 않고 다음 조각만 상수 공간에서 계산한다. + +새 release 측정은 모든 승인 기준을 통과했다. + +| 검증 항목 | 승인 기준 | 결과 | +|---|---|---| +| Live-stream admission, quiet-host 3회 | 모든 회차 ITL p95 62.0 ms 이하 | **통과:** 54.1, 54.8, 55.2 ms | +| Qwen 3 1.7B server TTFT, 5회 | policy delta가 explicit-2048 null 범위 안 | **통과:** policy median +0.01%, null 범위 -4.73%~+3.26% | +| Llama 3.2 1B server TTFT, 5회 | policy delta가 explicit-2048 null 범위 안 | **통과:** policy median -0.20%, null 범위 -1.12%~-0.03% | + +### 보안 + +새로운 외부 trust boundary를 만들지 않는다. CLI와 환경 변수 입력은 기존의 양의 정수 또는 0 parser를 사용하고, 올바르지 않은 환경 변수는 경고 후 적응형 기본값으로 돌아간다. 로그에는 숫자 정책값만 기록한다. + +## 5. 검증 + +구현 worktree에서 완료한 검증: + +- cargo check --profile test-fast --features cuda --lib: 통과. +- cargo fmt --all -- --check: 통과. +- git diff --check: 통과. +- CUDA mlxcel-core prefill-plan 테스트: 17개 통과. +- CUDA server CLI 테스트: 148개 통과. +- CUDA adaptive scheduler 실행 regression: 통과. +- CUDA runtime-settings 테스트: 8개 통과. +- CUDA props 호환성 regression: 통과. +- 독립 security/performance review에서 반복되는 paged-block capacity check가 남은 모든 계획 조각을 즉시 생성하는 문제를 발견했다. Commit 88c2e6a0964d4a3a144c2ff1cc3c80dfd2893752는 이를 O(1) accessor로 교체하고 boundary, cursor, chunk, padding mode, model opt-out, embedding input 전반의 동등성 테스트를 추가한다. +- 수정 후 workspace all-targets Clippy가 통과했다. Focused CUDA runtime 검증은 prefill-plan 17개, block-reclaim 20개, adaptive scheduler transition regression이 모두 통과했다. +- 최초 full CUDA gate에서 capability helper 분리 후 source audit가 공통 guard를 인식하지 못한 문제를 발견했다. Audit는 brace-counted interprocedural 인식으로 보완했고 direct-guard 검사를 유지했으며 mutation fixture를 추가했다. 해당 테스트 3개는 모두 통과했다. + +Acceptance artifact는 production commit `88c2e6a0964d4a3a144c2ff1cc3c80dfd2893752`를 사용했고 clean worktree HEAD `b2394566a71fc1f6db83ad40b44380b9c129b57c`에서 측정했다. Driver는 두 test file만 다르다는 점을 검증했다. Raw record, server log, 정확한 환경, machine-readable summary, 재현 driver 및 analyzer는 `docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/`에 보관한다. + +Driver의 기본 build commit은 squash 전 측정 lineage에서만 유효하다. Squash merge 이후 또는 더 새로운 clean checkout에서 재실행하려면 해당 checkout에서 release binary 2개를 다시 빌드하고 `MLXCEL_BUILD_COMMIT="$(git rev-parse HEAD)"`를 지정해야 한다. 전체 명령은 benchmark report에 기록했다. + +PR #2270과 PR #2269를 merge한 뒤 전체 series는 workspace test `--no-run` compile, workspace all-targets Clippy, formatting, kernel dtype-key, kernel port-dispatch, llama-server compatibility contract suite, 그리고 `make verify-test-cuda`의 전체 single-threaded CUDA workspace run을 모두 통과했다. CUDA test summary는 전부 실패 0건이었고 BF16 및 F32 SSM update parity test도 통과했다. Rebase 전 검증 tree `ab8ff77ade8d6b1b4fb494707320cebc95150c8e`는 현재 `origin/main`으로 rebase한 뒤에도 정확히 같았다. 해당 비교 이후 변경은 문서와 PR 상태뿐이다. + +## 6. 변경 요약 + +| 항목 | 값 | +|---|---| +| 변경 파일 | 47 | +| 추가 라인 | 1879 | +| 삭제 라인 | 210 | +| 측정 대상 branch head | b2394566a71fc1f6db83ad40b44380b9c129b57c | + +| 분류 | 요약 | +|---|---| +| Core planning | 두 상태 정책, override parsing cache, 커서 기반 재개 계획 | +| Scheduling | 실시간 상태에 따른 청크 선택, 안전한 continuation/reservation, O(1) capacity check | +| Configuration | 명시값 식별과 두 서버 front end 전체의 두 값 전달 | +| Testing | Planner property, CLI 우선순위, schema 호환성, 실제 scheduler forward 순서 | +| Documentation | 적응형 기본값, override 동작, 측정 근거, cache partition 설명 | + +최초 구현 commit은 7c43726a0fb792594a69de0fda1e31e78eef4fa1이다. 7b2d17241ea9c16612766ccf37524cbe8be9f913은 CI Clippy 표현식 정책을 만족한다. 88c2e6a0964d4a3a144c2ff1cc3c80dfd2893752는 review에서 발견된 reservation 할당을 제거하며, 808bcccf387957a6da31b961666aadd5bf7cbe65는 all-targets Clippy에 필요한 테스트 전용 type annotation을 추가한다. 6f07f7563e74c5466b5aace6f3e8ee98ec8347e2와 b2394566a71fc1f6db83ad40b44380b9c129b57c는 공통 capability guard를 위한 source audit를 보수적으로 갱신하고 fixture를 강화한다. + +## 7. 완료 상태 + +Issue #2228의 구현, 성능 승인, 통합 검증에 남은 작업은 없다. + +## 참고 자료 + +- Issue #2228: 다른 시퀀스의 디코딩 여부에 따른 프리필 청크 선택. +- PR #2226 및 docs/benchmark_results/unified-engine-final-gb10-2026-10-08.md: 정책 근거가 된 admission 측정. +- ADR 0007: unified batch/native engine 결정과 개정된 프리필 청크 정책. +- docs/CONTINUOUS_BATCHING.md: 서버 정책, 계획 invariant, 운영 제어. +- docs/benchmark_results/adaptive-prefill-chunk-gb10-2026-10-11.md: 통과한 승인 결과, raw data, 재현 protocol. diff --git a/docs/CONTINUOUS_BATCHING.md b/docs/CONTINUOUS_BATCHING.md index 3988af1f2..b7fa4c6f2 100644 --- a/docs/CONTINUOUS_BATCHING.md +++ b/docs/CONTINUOUS_BATCHING.md @@ -9,10 +9,11 @@ serving roles that split prefill and decode across processes. ## Continuous batching The scheduler keeps up to `--parallel` sequences active at once. Each step it -either admits and prefills queued prompts (chunked at `--prefill-chunk-size` so a -long prompt does not stall decode) or advances the active batch by one decode -token, then streams the new tokens out. The scheduler decides what runs; the -engine runs it. `mlxcel_core::engine::Engine` owns the model and the KV pool, +either admits and prefills queued prompts (chunked at 2048 tokens while no other +sequence is decoding and 512 while a decode batch is live, unless overridden) +or advances the active batch by one decode token, then streams the new tokens +out. The scheduler decides what runs; the engine runs it. +`mlxcel_core::engine::Engine` owns the model and the KV pool, and every model forward, sampler draw and finish step the scheduler needs goes through `Engine::prefill`, `Engine::step` (one entry for every row count: a lone request is a batch of one) and the pipelined `submit` / `finish_rows` pair @@ -31,7 +32,7 @@ the speculative burst generators stay in the scheduler. Relevant flags: | `--max-batch-prefill N` | 4 | Requests batched into one prefill forward pass (families that support it). | | `--max-batch-prefill-tokens N` | (derived) | Padded-token budget bounding one batched prefill's transient memory. Unset derives `2 * max_batch_prefill * min(prefill_chunk_size, 512)`; `0` disables the cap. | | `--max-queue-depth N` | 32 | Maximum queued (not yet admitted) requests. | -| `--prefill-chunk-size N` | 2048 | Token chunk size for prefill; bounds prefill's effect on decode latency. One policy with the CLI's `MLXCEL_PREFILL_CHUNK` (ADR 0007). | +| `--prefill-chunk-size N` | 2048 alone; 512 with live decode | Token chunk size for both scheduler states when explicitly set. `--batch-size` / `LLAMA_ARG_BATCH` and `MLXCEL_PREFILL_CHUNK` are equivalent overrides. Without one, the scheduler chooses per tick to preserve single-stream TTFT while bounding admission's effect on active decode latency (#2228). | | `--prefill-grant-interval N` | 16 | Decode ticks a parked chunked prefill yields before it is granted one; bounds an admitted long prompt's time to first token. `0` disables the grant (unbounded wait). | | `--kv-admission-watermark F` | 0.01 | Fraction of the paged KV block budget (`0.0` to `0.5`) a prefill admission keeps free while rows are decoding, so they can grow without being preempted. `0` disables it. See below. | | `--enable-preemption` | off | Allow evicting a lower-priority sequence to admit a waiting one. | @@ -177,11 +178,11 @@ context-sizing note in [environment-variables.md](environment-variables.md)). ### One prefill plan, and when a prompt-cache hit reproduces a miss -Since #2170 the CLI generator and the scheduler prefill through one `PrefillPlan` (`mlxcel_core::prefill_plan`): the adopted prompt-cache prefix is skipped, a snapshot family's chat prompt is split once at its history boundary (#1143, one unpadded forward, the point the prompt cache snapshots), the rest is cut into `--prefill-chunk-size` pieces (default 2048, the same policy value as the CLI's `MLXCEL_PREFILL_CHUNK`, decided in [ADR 0007](adr/0007-unified-batch-native-engine.md)), and a piece is tile-padded on M5+ hardware only where the model opts in. The scheduler runs one piece per tick and rebuilds the plan from the sequence each tick, so chunked prefill keeps interleaving with decode; the Gemma 4 MTP burst takes its row-wise prefill ranges from the same plan. +Since #2170 the CLI generator and the scheduler prefill through one `PrefillPlan` (`mlxcel_core::prefill_plan`): the adopted prompt-cache prefix is skipped, a snapshot family's chat prompt is split once at its history boundary (#1143, one unpadded forward, the point the prompt cache snapshots), the rest is cut into chunk pieces, and a piece is tile-padded on M5+ hardware only where the model opts in. The direct engine and an otherwise idle server use 2048. While another sequence is decoding, the server uses 512; an explicit `--prefill-chunk-size` (or `--batch-size` / `LLAMA_ARG_BATCH`) or `MLXCEL_PREFILL_CHUNK` pins both states. The scheduler runs one piece per tick and rebuilds the continuation from the live cursor, so the chunk may change safely when the decode batch empties or becomes active. The Gemma 4 MTP burst remains on the alone value and takes its row-wise prefill ranges from the same plan. Two prefills of the same prompt that forward the same pieces from the same KV state are bitwise identical under `MLXCEL_SDPA_DETERMINISTIC=1`. Two that partition the prompt differently are not: the CUDA quantized matmul routes a forward of fewer than 8 rows to the per-row `qmv` kernel and larger ones to the tiled `qmm` kernel, attention over a short query block tiles differently from the full causal pass, and the two reduce in different orders, so KV and logits differ in the last bit and a greedy stream can flip at a near-tie token. The split at the history boundary, `--prefill-chunk-size`, and a prompt-cache hit (which forwards only the suffix after the adopted prefix) are all partitions. -So a prompt-cache hit reproduces the cold (miss) run of the same prompt exactly when its adopted prefix ends on one of the cold plan's split points (`PrefillPlan::reproduces`): the history boundary of a snapshot family, or a chunk edge. That is the plan half of the condition; the other half is that the adopted rows were themselves written by the same pieces, as a boundary snapshot of an earlier turn or a primed prefix forwarded as one piece are. Rows a donor wrote by decode steps, or inside a longer forward, are a different partition even when the adopted length is a split point. `make engine-parity` measures both shapes on GB10 (2026-10-07, 64 greedy and seeded tokens, the hit primed with the prompt's history prefix): +So a prompt-cache hit reproduces the cold (miss) run of the same prompt exactly when its adopted prefix ends on one of the cold plan's split points (`PrefillPlan::reproduces`): the history boundary of a snapshot family, or a chunk edge. That is the plan half of the condition; the other half is that the adopted rows were themselves written by the same pieces, as a boundary snapshot of an earlier turn or a primed prefix forwarded as one piece are. Rows a donor wrote by decode steps, or inside a longer forward, are a different partition even when the adopted length is a split point. A prefill whose chunk changes with scheduler state is likewise a different partition; prefills that run in the same scheduler states keep the existing reproduction invariant. `make engine-parity` measures both shapes on GB10 (2026-10-07, 64 greedy and seeded tokens, the hit primed with the prompt's history prefix): | model | miss partition | hit partition | miss vs hit | |---|---|---|---| @@ -196,7 +197,14 @@ The dense-KV rows diverge because their cold plan has no split point at 45 or 71 Known limitations of the plan, recorded for the epic's end-of-run measurement: - The history segment of a snapshot family's prefill (the span before the history boundary, #1143) is always one unchunked forward, whatever `--prefill-chunk-size` says. This predates the plan (it has been the case since #1143), and the plan keeps it because the segment is where the prompt cache snapshots the model state. A long chat history on a snapshot family therefore prefills as one forward that does not interleave with concurrent decode ticks; only the pieces after the boundary do. -- What the 2048 default (up from 512 on the server) does to the inter-token latency of concurrent decode streams is not measured yet. A larger chunk lengthens each prefill tick that a live decode batch waits behind. The epic's end-of-run benchmark measures it; `--prefill-chunk-size` and `MLXCEL_PREFILL_CHUNK` lower it per deployment in the meantime. +- PR #2226 measured the tradeoff that the adaptive default resolves: with four Llama-3.2-1B 4-bit streams decoding and an 8192-token request admitted, 2048 produced 127.9–128.9 ms stream ITL p95 during admission and 1903–2060 ms admitted-request TTFT; 512 produced 55.8–56.4 ms ITL p95 and 3768–4067 ms TTFT (three rounds). The scheduler therefore uses 512 only while `active_batch` is non-empty and returns to 2048 when the decode streams finish (#2228). + +The shipped policy passed its fresh-release GB10 acceptance run: three +quiet-host admission rounds measured 54.1, 54.8, and 55.2 ms live-stream ITL +p95, below the 62.0 ms limit. Five-round idle-server TTFT comparisons placed +the policy median inside the explicit-2048 null range on both Qwen3-1.7B and +Llama-3.2-1B. See the +[raw results and protocol](benchmark_results/adaptive-prefill-chunk-gb10-2026-10-11.md). ### Decode headroom under a tight budget (`--kv-admission-watermark`) diff --git a/docs/adr/0007-unified-batch-native-engine.md b/docs/adr/0007-unified-batch-native-engine.md index fd9fdd79b..bac9d4df0 100644 --- a/docs/adr/0007-unified-batch-native-engine.md +++ b/docs/adr/0007-unified-batch-native-engine.md @@ -91,6 +91,14 @@ Phase 5 makes the CLI a client of the server engine, so the server's llama-serve | Prefill chunk | 2048 (`DEFAULT_PREFILL_CHUNK`, `src/lib/mlxcel-core/src/generate.rs:282`; `MLXCEL_PREFILL_CHUNK`) | 512 (`src/server/config.rs:1172`, `src/server/startup.rs:698`, `--prefill-chunk-size`) | 2048, 512 | TTFT at the 8192-token prompt, Qwen3-1.7B 4-bit and Llama-3.2-1B 4-bit, on both paths (`--prefill-chunk 2048` vs `512`), interleaved rounds with a null arm. Lower median TTFT wins when its per-round delta clears the null spread on both models; if neither clears it, 512 wins (bounded prefill transient and finer interleaving with live decode streams, ADR 0005 and #1011). | **2048.** TTFT at 8192 tokens, 512 vs 2048 paired: Qwen3 CLI +8.4 %, server +22.9 %; Llama CLI +29.4 %, server +14.6 %; every null range within ±1.9 %. Decode unchanged. ([results](../benchmark_results/unified-engine-baseline-gb10-2026-10-07.md)) | | Single-sequence KV storage | dense `KVCache`, MLX fused SDPA | paged when the worker can serve it (`effective_decode_storage_backend`, `src/server/batch/scheduler/mod.rs:210`: `auto` resolves to paged when `max_batch_size > 1` and the model supports batching and paged decode), else dense | dense, paged | B=1 decode tok/s on the server path, `--decode-storage dense` vs `paged`, both models, 256 and 8192-token prompts, interleaved rounds with a null arm. Both models keep their KV in the scheduler's pool, so a lone paged sequence runs the pooled paged kernel (`src/models/qwen3.rs:354`, `src/models/llama3.rs:746`); every paged row must show `paged_decode_launches > 0`. The result does not transfer to model-owned KV families, which decode a lone sequence dense under either setting. Higher median wins when it clears the null spread in every cell; a split result keeps both storages selectable and the long-context winner is the default, since decode at 8K keys is where storage layout matters. | **dense.** B=1 decode, paged vs dense paired: Qwen3 −23.0 % (256) and −6.5 % (8192); Llama −8.2 % and −8.1 %; every null range within ±1.1 %; paged rows ran the paged kernel (3612 and 2064 launches). Paged stays the batched backend. ([results](../benchmark_results/unified-engine-baseline-gb10-2026-10-07.md)) | +Amendment (#2228): PR #2226 measured the server's 2048 choice under live decode. Across three Llama-3.2-1B 4-bit admission rounds, stream ITL p95 was 127.9–128.9 ms at 2048 and 55.8–56.4 ms at 512, while the admitted request's TTFT was 1903–2060 ms and 3768–4067 ms respectively. No fixed value wins both measures, so the server now chooses per scheduler tick: 2048 with no other sequence decoding and 512 while `active_batch` is non-empty. `--prefill-chunk-size`, `--batch-size` / `LLAMA_ARG_BATCH`, and `MLXCEL_PREFILL_CHUNK` pin both states. The direct engine remains at the 2048 alone value. + +The fresh-release acceptance run on GB10 confirmed the decision: live-stream +admission ITL p95 was 54.1, 54.8, and 55.2 ms in three quiet-host rounds, and +the adaptive policy's five-round idle-server TTFT median lay inside the +explicit-2048 null range for both Qwen3-1.7B and Llama-3.2-1B +([results](../benchmark_results/adaptive-prefill-chunk-gb10-2026-10-11.md)). + ### Regression threshold | Row | Value | How it is set | diff --git a/docs/benchmark_results/adaptive-prefill-chunk-gb10-2026-10-11.md b/docs/benchmark_results/adaptive-prefill-chunk-gb10-2026-10-11.md new file mode 100644 index 000000000..582e872ff --- /dev/null +++ b/docs/benchmark_results/adaptive-prefill-chunk-gb10-2026-10-11.md @@ -0,0 +1,120 @@ +# Adaptive server prefill chunks (GB10, 2026-10-11) + +Acceptance measurement for issue #2228 and PR #2268. The server now uses a +2048-token prefill chunk while no other sequence is decoding and a 512-token +chunk while a decode batch is live. The acceptance gates check both sides of +that policy: concurrent admission must preserve live-stream latency, and an +idle server must retain the TTFT of an explicit 2048-token chunk. + +## Host and method + +NVIDIA GB10, driver 580.178.04, host `spark-102`, release profile with +`--features cuda`. The run started at 2026-10-11 04:18 UTC under `gpu-lock`. +Every admission round passed the quiet-host gate before its server started; +every TTFT arm used the benchmark harness's `--hostgate`. No other build, test, +or GPU job ran during the measurement. + +The release binaries contain production commit +`88c2e6a0964d4a3a144c2ff1cc3c80dfd2893752`. The clean measurement worktree was +at `b2394566a71fc1f6db83ad40b44380b9c129b57c`; the driver verified that the only +differences were two test files (a prefill-plan type annotation and the source +audit fixtures), so the measured production code exactly matches the release +commit. The declared MLX pin and fetched MLX HEAD were both +`81ba1c6a0e50a9268b931579c2d4f1158b9aab5a`. The fetched source carried the +project's recorded patch overlay; its complete status is in `environment.txt`. + +Admission used Llama-3.2-1B-Instruct 4-bit, `--parallel 8 --ignore-eos`, four +4096-token live streams, and one admitted 8192-token request. It ran three +rounds through `scripts/bench_mixed_step_admission.py`; acceptance required +live-stream ITL p95 at or below 62.0 ms in every admission window. + +Idle-server TTFT used Qwen3-1.7B 4-bit and Llama-3.2-1B-Instruct 4-bit with an +8192-token prompt. Five interleaved rounds compared the adaptive default +(`policy`) with an explicit `--prefill-chunk 2048` reference (`c2048`) and a +repeat of that reference (`c2048-null`). Acceptance required the paired policy +median to lie inside the paired null range, using unrounded measurements. + +## Concurrent admission + +| Round | Quiet ITL p95 | Admission ITL p95 | Admission ITL mean | Admitted TTFT | +|---|---:|---:|---:|---:| +| 1 | 15.1 ms | 54.1 ms | 13.1 ms | 3347 ms | +| 2 | 16.0 ms | 54.8 ms | 13.4 ms | 3331 ms | +| 3 | 18.8 ms | 55.2 ms | 13.3 ms | 3555 ms | + +**Passed:** all three admission p95 values are below 62.0 ms. Each round +recorded 14 prefill chunks and 13 prefill grants while 4095 decode steps ran; +the admitted request produced its first token before the live streams ended. + +## Idle-server TTFT + +| Model | Explicit 2048 median | Policy median | Policy paired delta | Null paired range | Result | +|---|---:|---:|---:|---:|---| +| Qwen3-1.7B 4-bit | 688.76 ms | 663.41 ms | +0.014712% (-3.794710..+0.975952%) | -4.729402..+3.261029% | Pass | +| Llama-3.2-1B 4-bit | 416.15 ms | 413.29 ms | -0.202854% (-1.606240..+1.595997%) | -1.115018..-0.029672% | Pass | + +Both policy medians fall inside their model's null range. On an idle server, +the adaptive default is therefore indistinguishable from explicitly fixing +the chunk to 2048 at the resolution of this run. + +## Decision + +The adaptive default passes both acceptance gates. Under live decode it keeps +admission ITL p95 at 54.1 to 55.2 ms, matching the 512-token behavior that +motivated the policy. With no live decode, its TTFT stays within the explicit +2048 noise floor on both tested model families. The shipped defaults remain +2048 alone and 512 contended; an explicit CLI, compatibility, or environment +override continues to pin both states. + +## Integrated validation + +After PR #2270 and PR #2269 merged, the complete issue #2228 series was +validated with both preceding changes present. Workspace test compilation +(`--no-run`), workspace all-targets Clippy, formatting, the kernel dtype-key, +kernel port-dispatch, and llama-server compatibility contract suites, and the +full single-threaded CUDA workspace run (`make verify-test-cuda`) all passed. +The CUDA run reported zero failed test summaries, including +`ssm_update_kernel_matches_graph_step_bf16` and +`ssm_update_kernel_matches_graph_step_f32`. + +The validated pre-rebase tree was +`ab8ff77ade8d6b1b4fb494707320cebc95150c8e`; rebasing the series onto the +current `origin/main` preserved that tree exactly. The final edits after that +comparison update documentation and PR status only. + +## Reproduce + +The exact raw records, server logs, environment provenance, and machine-readable +summary are in +[`data/adaptive-prefill-chunk-gb10-2026-10-11/`](data/adaptive-prefill-chunk-gb10-2026-10-11/). +Run its `analyze.py` against that directory to recheck the archived acceptance +without rounding. `run.sh` reruns the full protocol from a clean worktree, +acquires `gpu-lock`, checks port 18080, validates the release artifact against +`MLXCEL_BUILD_COMMIT`, and writes a new result directory under `/tmp`. +`MLXCEL_BIN`, `MLXCEL_BENCH_BIN`, and `MLXCEL_MODELS` can point it at relocated +artifacts and checkpoints. + +The default `MLXCEL_BUILD_COMMIT` is the archived production commit and is only +valid on an unsquashed PR lineage where that commit is an ancestor and every +later production source is unchanged. It is not expected to work unchanged from +a squash-merged checkout. For a safe replay of the current clean revision, +build both release binaries from that revision and identify it explicitly: + +```bash +cargo build --release --features cuda --bin mlxcel --bin mlxcel-bench-engine +MLXCEL_BUILD_COMMIT="$(git rev-parse HEAD)" \ + docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/run.sh +``` + +To replay with the archived release binaries instead, check out an unsquashed +PR revision that contains this driver and descends from +`88c2e6a0964d4a3a144c2ff1cc3c80dfd2893752`; keep the worktree clean and leave +`MLXCEL_BUILD_COMMIT` at its default. The script rejects unrelated production +changes in that mode rather than attributing them to the archived binary. + +The archived data came from the original driver invocation: + +```text +/tmp/mlxcel-auto-plan/2228-acceptance.sh /tmp/mlxcel-issue-2228-acceptance-20261011 +python3 /tmp/mlxcel-auto-plan/2228-analyze.py /tmp/mlxcel-issue-2228-acceptance-20261011 +``` diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/.gitattributes b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/.gitattributes new file mode 100644 index 000000000..1a1e2ec80 --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/.gitattributes @@ -0,0 +1 @@ +admit-policy-r*.txt whitespace=-blank-at-eof diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/acceptance-summary.json b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/acceptance-summary.json new file mode 100644 index 000000000..4f88963c5 --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/acceptance-summary.json @@ -0,0 +1,50 @@ +{ + "admission": [ + { + "round": 1, + "p95_ms": 54.1, + "pass": true + }, + { + "round": 2, + "p95_ms": 54.8, + "pass": true + }, + { + "round": 3, + "p95_ms": 55.2, + "pass": true + } + ], + "ttft": [ + { + "model": "qwen3-1.7b-4bit", + "policy_median_delta_pct": 0.014711996169936015, + "policy_range_pct": [ + -3.7947101163357533, + 0.9759518988692761 + ], + "null_range_pct": [ + -4.729402218560718, + 3.2610291206550457 + ], + "null_median_delta_pct": -1.1428652324241395, + "pass": true + }, + { + "model": "llama-3.2-1b-instruct-4bit", + "policy_median_delta_pct": -0.20285407850114678, + "policy_range_pct": [ + -1.6062401209527066, + 1.5959968179777517 + ], + "null_range_pct": [ + -1.115018402057133, + -0.02967215291590497 + ], + "null_median_delta_pct": -0.21802043638052826, + "pass": true + } + ], + "pass": true +} diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/admit-policy-r1.txt b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/admit-policy-r1.txt new file mode 100644 index 000000000..cadc63bd8 --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/admit-policy-r1.txt @@ -0,0 +1,20 @@ +# mixed prefill/decode admission bench (issue #908) +model=llama-3.2-1b-instruct-4bit streams=4 settle=5.0s +stream_prompt~128 tok admit_prompt~8192 tok stream_max_tokens=4096 +expect=any + +## scheduler tick attribution + mixed_steps_total: +0 + prefill_grants_total: +13 + prefill_chunks_total: +14 + decode_steps_total: +4095 + +## per-stream inter-token latency + quiet window (before admission): p50 10.4 ms p95 15.1 ms mean 11.1 ms n=1335 + admission window (closed by the admitted request's first token): p50 9.6 ms p95 54.1 ms mean 13.1 ms n=805 + p95 inflation during admission: 3.59x + +## admitted request + TTFT: 3347 ms + admitted_first_token_after_last_stream: False + diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/admit-policy-r2.txt b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/admit-policy-r2.txt new file mode 100644 index 000000000..f8ce45cad --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/admit-policy-r2.txt @@ -0,0 +1,20 @@ +# mixed prefill/decode admission bench (issue #908) +model=llama-3.2-1b-instruct-4bit streams=4 settle=5.0s +stream_prompt~128 tok admit_prompt~8192 tok stream_max_tokens=4096 +expect=any + +## scheduler tick attribution + mixed_steps_total: +0 + prefill_grants_total: +13 + prefill_chunks_total: +14 + decode_steps_total: +4095 + +## per-stream inter-token latency + quiet window (before admission): p50 11.1 ms p95 16.0 ms mean 12.0 ms n=1335 + admission window (closed by the admitted request's first token): p50 10.1 ms p95 54.8 ms mean 13.4 ms n=805 + p95 inflation during admission: 3.44x + +## admitted request + TTFT: 3331 ms + admitted_first_token_after_last_stream: False + diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/admit-policy-r3.txt b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/admit-policy-r3.txt new file mode 100644 index 000000000..5777a5863 --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/admit-policy-r3.txt @@ -0,0 +1,20 @@ +# mixed prefill/decode admission bench (issue #908) +model=llama-3.2-1b-instruct-4bit streams=4 settle=5.0s +stream_prompt~128 tok admit_prompt~8192 tok stream_max_tokens=4096 +expect=any + +## scheduler tick attribution + mixed_steps_total: +0 + prefill_grants_total: +13 + prefill_chunks_total: +14 + decode_steps_total: +4095 + +## per-stream inter-token latency + quiet window (before admission): p50 11.5 ms p95 18.8 ms mean 12.5 ms n=1335 + admission window (closed by the admitted request's first token): p50 9.9 ms p95 55.2 ms mean 13.3 ms n=805 + p95 inflation during admission: 2.94x + +## admitted request + TTFT: 3555 ms + admitted_first_token_after_last_stream: False + diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/analyze.py b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/analyze.py new file mode 100755 index 000000000..c5f6d758b --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/analyze.py @@ -0,0 +1,46 @@ +#!/usr/bin/env python3 +"""Check issue 2228 acceptance without rounding measurements.""" +import json +import math +from pathlib import Path +import re +import statistics +import sys + +root = Path(sys.argv[1]) +result = {'admission': [], 'ttft': []} +for round_no in range(1, 4): + path = root / f'admit-policy-r{round_no}.txt' + match = re.search(r'admission window .*?p95 ([0-9.]+) ms', path.read_text()) + if not match: + raise SystemExit(f'missing admission p95: {path}') + p95 = float(match[1]) + result['admission'].append({'round': round_no, 'p95_ms': p95, 'pass': p95 <= 62.0}) +for model in ['qwen3-1.7b-4bit', 'llama-3.2-1b-instruct-4bit']: + rows = [json.loads(line) for line in (root / f'ttft-{model}.jsonl').read_text().splitlines()] + arms = {name: {} for name in ['c2048', 'policy', 'c2048-null']} + for row in rows: + if row.get('kind') == 'header': + continue + if row['path'] != 'server' or row['prompt_target_len'] != 8192: + raise SystemExit(f'unexpected benchmark cell: {row}') + arm = arms[row['arm']] + if row['round'] in arm: + raise SystemExit('duplicate round; do not append a rerun to original measurements') + value = row['ttft_ms'] + if not math.isfinite(value) or value <= 0: + raise SystemExit('invalid TTFT') + arm[row['round']] = value + if any(set(arm) != set(range(5)) for arm in arms.values()): + raise SystemExit(f'incomplete rounds: {model}') + delta = lambda name: [(arms[name][r] / arms['c2048'][r] - 1) * 100 for r in range(5)] + policy, null = delta('policy'), delta('c2048-null') + median = statistics.median(policy) + result['ttft'].append({'model': model, 'policy_median_delta_pct': median, + 'policy_range_pct': [min(policy), max(policy)], 'null_range_pct': [min(null), max(null)], + 'null_median_delta_pct': statistics.median(null), + 'pass': min(null) <= median <= max(null)}) +result['pass'] = all(row['pass'] for group in ['admission', 'ttft'] for row in result[group]) +(root / 'acceptance-summary.json').write_text(json.dumps(result, indent=2) + '\n') +print(json.dumps(result, indent=2)) +sys.exit(0 if result['pass'] else 1) diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/environment.txt b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/environment.txt new file mode 100644 index 000000000..44f848b8d --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/environment.txt @@ -0,0 +1,46 @@ +2026-10-11T04:18:12Z +spark-102 +NVIDIA GB10, 580.178.04 +b2394566a71fc1f6db83ad40b44380b9c129b57c +release_production_commit=88c2e6a0964d4a3a144c2ff1cc3c80dfd2893752 + src/lib/mlxcel-core/src/prefill_plan_tests.rs | 2 +- + tests/vlm_wrapper_capability_delegation.rs | 161 +++++++++++++++++++------- + 2 files changed, 121 insertions(+), 42 deletions(-) +mlx_git_tag=81ba1c6a0e50a9268b931579c2d4f1158b9aab5a +mlx_build_dir=/home/inureyes/Development/backend.ai/wt-auto-20261011T0224-issue-2228/target/release/build/mlxcel-core-4c8821fbbcc3d8f4 +mlx_fetched_head=81ba1c6a0e50a9268b931579c2d4f1158b9aab5a + M mlx/backend/cpu/quantized.cpp + M mlx/backend/cuda/binary/binary.cuh + M mlx/backend/cuda/cuda_utils.h + M mlx/backend/cuda/device/binary_ops.cuh + M mlx/backend/cuda/device/gemm_sm70.cuh + M mlx/backend/cuda/device/qmm_naive.cuh + M mlx/backend/cuda/device/qmm_sm80.cuh + M mlx/backend/cuda/gemms/grouped_gemm.h + M mlx/backend/cuda/gemms/grouped_gemm_unaligned.cu + M mlx/backend/cuda/jit_module.cpp + M mlx/backend/cuda/jit_module.h + M mlx/backend/cuda/matmul.cpp + M mlx/backend/cuda/quantized/qmm/qmm_naive.cu + M mlx/backend/cuda/quantized/qmm/qmm_sm80.cu + M mlx/backend/cuda/quantized/qmm/qmv.cu + M mlx/backend/cuda/quantized/quantized.cpp + M mlx/backend/cuda/reduce/all_reduce.cu + M mlx/backend/cuda/reduce/col_reduce.cu + M mlx/backend/cuda/reduce/init_reduce.cu + M mlx/backend/cuda/reduce/reduce_ops.cuh + M mlx/backend/cuda/reduce/row_reduce.cu + M mlx/backend/cuda/scaled_dot_product_attention.cpp + M mlx/backend/cuda/scaled_dot_product_attention.cu + M mlx/backend/metal/compiled.cpp + M mlx/backend/metal/device.cpp + M mlx/backend/metal/kernels/utils.h + M mlx/backend/metal/quantized.cpp + M mlx/dtype.cpp + M mlx/fast.cpp + M mlx/ops.cpp +?? mlx/backend/cuda/gemms/grouped_gemm_arch.h +?? mlx/backend/cuda/quantized/qmm/qmm_naive_tile.h +e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855 - +2026-10-11 13:00:13.602698582 +0900 /home/inureyes/Development/backend.ai/wt-auto-20261011T0224-issue-2228/target/release/mlxcel +2026-10-11 12:58:36.481966866 +0900 /home/inureyes/Development/backend.ai/wt-auto-20261011T0224-issue-2228/target/release/mlxcel-bench-engine diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/hostgate-admit-r1.txt b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/hostgate-admit-r1.txt new file mode 100644 index 000000000..e69de29bb diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/hostgate-admit-r2.txt b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/hostgate-admit-r2.txt new file mode 100644 index 000000000..e69de29bb diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/hostgate-admit-r3.txt b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/hostgate-admit-r3.txt new file mode 100644 index 000000000..e69de29bb diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/run.sh b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/run.sh new file mode 100755 index 000000000..462948ae4 --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/run.sh @@ -0,0 +1,148 @@ +#!/usr/bin/env bash +set -euo pipefail + +SCRIPT_DIR=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd) +WT=$(git -C "$SCRIPT_DIR" rev-parse --show-toplevel) +BIN=${MLXCEL_BIN:-"$WT/target/release/mlxcel"} +BENCH_BIN=${MLXCEL_BENCH_BIN:-"$WT/target/release/mlxcel-bench-engine"} +MODELS=${MLXCEL_MODELS:-"$WT/models/mlx"} +BUILD_COMMIT=${MLXCEL_BUILD_COMMIT:-88c2e6a0964d4a3a144c2ff1cc3c80dfd2893752} +OUT=${1:-"/tmp/mlxcel-issue-2228-rerun-$(date -u +%Y%m%dT%H%M%SZ)"} + +# The policy arm must exercise the adaptive defaults, independent of the +# operator's shell environment. +unset MLXCEL_PREFILL_CHUNK LLAMA_ARG_BATCH + +if [[ ${MLXCEL_ISSUE_2228_LOCKED:-0} != 1 ]]; then + exec gpu-lock run --tag issue-2228-acceptance -- env MLXCEL_ISSUE_2228_LOCKED=1 "$0" "$OUT" +fi + +mkdir -p "$OUT" +[[ -x "$BIN" ]] || { echo "missing fresh release binary: $BIN" >&2; exit 2; } +[[ -x "$BENCH_BIN" ]] || { echo "missing fresh release benchmark: $BENCH_BIN" >&2; exit 2; } +[[ -d "$MODELS/llama-3.2-1b-instruct-4bit" ]] || { echo "missing llama model link" >&2; exit 2; } +[[ -d "$MODELS/qwen3-1.7b-4bit" ]] || { echo "missing qwen model link" >&2; exit 2; } +# The archived default identifies the measured production commit on the +# unsquashed PR lineage. A fresh build, including one from a squash-merged +# checkout, must set MLXCEL_BUILD_COMMIT to the exact revision that produced it. +if ! git -C "$WT" merge-base --is-ancestor "$BUILD_COMMIT" HEAD; then + echo "build commit $BUILD_COMMIT is not an ancestor of HEAD; use the unsquashed measured lineage or rebuild HEAD and set MLXCEL_BUILD_COMMIT=$(git -C "$WT" rev-parse HEAD)" >&2 + exit 2 +fi +while IFS= read -r changed; do + case "$changed" in + src/lib/mlxcel-core/src/prefill_plan_tests.rs|tests/vlm_wrapper_capability_delegation.rs|docs/*|TECHNICAL_REPORTS/*) ;; + *) echo "production changed since release build: $changed" >&2; exit 2 ;; + esac +done < <(git -C "$WT" diff --name-only "$BUILD_COMMIT" HEAD) +[[ -z $(git -C "$WT" status --porcelain --untracked-files=all) ]] || { echo "measurement requires a clean worktree" >&2; exit 2; } +head_time=$(git -C "$WT" show -s --format=%ct "$BUILD_COMMIT") +(( $(stat -c %Y "$BIN") >= head_time )) || { echo "mlxcel predates release production commit" >&2; exit 2; } +(( $(stat -c %Y "$BENCH_BIN") >= head_time )) || { echo "mlxcel-bench-engine predates release production commit" >&2; exit 2; } + +core_build= +while IFS= read -r candidate; do + if [[ -d "$candidate/out/build/include/cccl" && -d "$candidate/out/build/_deps/mlx-src/.git" ]]; then + core_build=$candidate + break + fi +done < <(find "$WT/target/release/build" -maxdepth 1 -type d -name 'mlxcel-core-*' -printf '%T@ %p\n' | sort -nr | cut -d' ' -f2-) +[[ -n $core_build ]] || { echo "missing release MLX/CCCL build tree" >&2; exit 2; } +mlx_src="$core_build/out/build/_deps/mlx-src" +export MLXCEL_CCCL_DIR="$core_build/out/build/include/cccl" +export MLXCEL_CUTLASS_DIR="$core_build/out/build/include" +mlx_tag=$(awk '$1 == "GIT_TAG" { gsub(/[\")]/, "", $2); print $2 }' "$WT/src/lib/mlx-cpp/CMakeLists.txt") +mlx_head=$(git -C "$mlx_src" rev-parse HEAD) +[[ $mlx_head == "$mlx_tag" ]] || { echo "fetched MLX HEAD $mlx_head != declared pin $mlx_tag" >&2; exit 2; } + +{ + date -u +%FT%TZ + hostname + nvidia-smi --query-gpu=name,driver_version --format=csv,noheader + git -C "$WT" rev-parse HEAD + echo "release_production_commit=$BUILD_COMMIT" + git -C "$WT" diff --stat "$BUILD_COMMIT" HEAD + echo "mlx_git_tag=$mlx_tag" + echo "mlx_build_dir=$core_build" + echo "mlx_fetched_head=$mlx_head" + git -C "$mlx_src" status --porcelain --untracked-files=all + git -C "$WT" status --porcelain --untracked-files=all + git -C "$WT" diff --binary | sha256sum + stat -c '%y %n' "$BIN" "$BENCH_BIN" +} > "$OUT/environment.txt" + +server_pid= +cleanup() { + if [[ -n $server_pid ]]; then + kill "$server_pid" 2>/dev/null || true + wait "$server_pid" 2>/dev/null || true + server_pid= + fi +} +trap cleanup EXIT INT TERM + +wait_ready() { + local port=$1 + for _ in $(seq 1 240); do + if curl -sf "http://127.0.0.1:$port/health" >/dev/null 2>&1 || curl -sf "http://127.0.0.1:$port/v1/models" >/dev/null 2>&1; then + return 0 + fi + sleep 1 + done + return 1 +} + +for round in 1 2 3; do + PYTHONPATH="$WT/docs/benchmark_results/data/sdpa-plan-bucket-gb10-2026-09-12/harness" \ + python3 -c 'import hostgate; hostgate.wait_quiet()' \ + >"$OUT/hostgate-admit-r$round.txt" 2>&1 + if ss -H -ltn 'sport = :18080' | grep -q .; then + echo "port 18080 is already in use; refusing to probe an unrelated server" >&2 + exit 3 + fi + log="$OUT/server-admit-policy-r$round.log" + "$BIN" serve -m "$MODELS/llama-3.2-1b-instruct-4bit" --port 18080 --metrics --parallel 8 --ignore-eos >"$log" 2>&1 & + server_pid=$! + wait_ready 18080 || { echo "admission server failed to start in round $round" >&2; exit 3; } + timeout 3600 python3 "$WT/scripts/bench_mixed_step_admission.py" --port 18080 --streams 4 --stream-max-tokens 4096 --admit-prompt-tokens 8192 --expect any >"$OUT/admit-policy-r$round.txt" 2>&1 + cleanup + sleep 5 +done + +python3 - "$OUT" <<'PY' +import pathlib +import re +import sys + +out = pathlib.Path(sys.argv[1]) +values = [] +for path in sorted(out.glob("admit-policy-r*.txt")): + text = path.read_text() + match = re.search(r"admission window .*?p95 ([0-9.]+) ms", text) + if match is None: + raise SystemExit(f"missing admission p95 in {path}") + value = float(match.group(1)) + values.append(value) + print(f"{path.name}: admission ITL p95 {value:.1f} ms") +if len(values) != 3: + raise SystemExit(f"expected three admission rounds, found {len(values)}") +if any(value > 62.0 for value in values): + raise SystemExit(f"admission acceptance failed: {values} exceeds 62.0 ms") +PY + +for model in qwen3-1.7b-4bit llama-3.2-1b-instruct-4bit; do + python3 "$WT/scripts/engine_bench_rounds.py" \ + --bin "$BENCH_BIN" \ + --model "$MODELS/$model" \ + --prompt-tokens 8192 \ + --rounds 5 \ + --hostgate \ + --arm 'c2048=--path server --prefill-chunk 2048' \ + --arm 'policy=--path server' \ + --out "$OUT/ttft-$model.jsonl" \ + | tee "$OUT/ttft-$model.txt" +done + +python3 "$SCRIPT_DIR/analyze.py" "$OUT" + +echo "acceptance data: $OUT" diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/server-admit-policy-r1.log b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/server-admit-policy-r1.log new file mode 100644 index 000000000..fe3868f15 --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/server-admit-policy-r1.log @@ -0,0 +1,41 @@ +2026-10-11T04:22:00.430580Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins +2026-10-11T04:22:00.431113Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0 +2026-10-11T04:22:00.433592Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=8 kv_unified=false n_parallel=8 prefill_chunk_size=2048 prefill_chunk_size_contended=512 max_kv_size=None +2026-10-11T04:22:00.433756Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA) +2026-10-11T04:22:00.433759Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set) +2026-10-11T04:22:00.433761Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB +2026-10-11T04:22:00.775044Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"] +2026-10-11T04:22:00.777890Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256 +2026-10-11T04:22:01.075106Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model... +2026-10-11T04:22:01.075128Z INFO mlxcel::server::startup: Warming up model... +2026-10-11T04:22:01.392702Z INFO mlxcel::server::model_provider::model_worker: Model llama-3.2-1b-instruct-4bit loaded in 0.318s (resident after load: 0.00 GB) worker_model_id=llama-3.2-1b-instruct-4bit load_seconds=0.317558454 active_bytes=128 peak_bytes=128 cache_bytes=0 limit_bytes=124128081510 +2026-10-11T04:22:01.392921Z INFO mlxcel::server::model_provider::model_worker: EOS tokens from config: [128009] +2026-10-11T04:22:01.392955Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=8, max_queue_depth=32, prefill_chunk_size=2048, prefill_chunk_size_contended=512, max_batch_prefill=4, decode_storage=auto) +2026-10-11T04:22:01.393154Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 1559064 blocks (16 layers, 32-token blocks) +2026-10-11T04:22:01.393163Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 2048 blocks per layer (fused decode serves a layer only while its rows fit one slab) +2026-10-11T04:22:01.394798Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=16 kv_cache_mode_total_layers=16 +2026-10-11T04:22:01.404800Z INFO prefill{seq_id=seq-0 prompt_len=2 cached=0 start=0}: mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved +2026-10-11T04:22:02.596538Z INFO prefill{seq_id=seq-0 prompt_len=2 cached=0 start=0}: mlxcel::server::batch::scheduler::prefill: prompt-cache: request completed during prefill: cached=0/2 prompt tokens, total 1200ms prompt_tokens=2 cached_tokens=0 generation_time_ms=1200 +2026-10-11T04:22:02.596591Z INFO mlxcel::server::startup: Warmup complete +2026-10-11T04:22:02.600135Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18080 +2026-10-11T04:22:02.600152Z INFO mlxcel::server::startup: Detected 1 GPU(s) +2026-10-11T04:22:02.600157Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin) +2026-10-11T04:22:02.600159Z INFO mlxcel::server::startup: Endpoints: +2026-10-11T04:22:02.600161Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions +2026-10-11T04:22:02.600171Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions +2026-10-11T04:22:02.600173Z INFO mlxcel::server::startup: GET /v1/models - List models +2026-10-11T04:22:02.600174Z INFO mlxcel::server::startup: POST /completion - llama-server native completion +2026-10-11T04:22:02.600175Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text +2026-10-11T04:22:02.600177Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens +2026-10-11T04:22:02.600178Z INFO mlxcel::server::startup: GET /props - Server properties +2026-10-11T04:22:02.600180Z INFO mlxcel::server::startup: GET /slots - Slot status +2026-10-11T04:22:02.600181Z INFO mlxcel::server::startup: GET /health - Health check +2026-10-11T04:22:03.172867Z INFO mlxcel::server::chat_request: prompt cache on + preserve_thinking unset: defaulting preserve_thinking=true for prefix stability (override via chat_template_kwargs.preserve_thinking=false) session=__mlxcel_anon__ +2026-10-11T04:22:03.497330Z INFO decode_step{batch_size=4}: mlxcel_core::cache::paged_batch_decode: paged decode v2: gather: 656 visible KV tokens across 4 request(s) is below the 2048-token dispatch floor +2026-10-11T04:22:07.191291Z INFO decode_step{batch_size=4}: mlxcel_core::cache::paged_batch_decode: paged decode v2: fused v2 launch (batch 4, 2048 visible KV tokens, 64 chunks, merge on) +Compiling CUDA kernels (first run on this host; cached for later runs)... +2026-10-11T04:22:11.903850Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/6933 prompt tokens, total 3083ms prompt_tokens=6933 cached_tokens=0 generation_time_ms=3083 +2026-10-11T04:22:44.577021Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41400ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41400 +2026-10-11T04:22:44.577262Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41400ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41400 +2026-10-11T04:22:44.577496Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41401ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41401 +2026-10-11T04:22:44.577707Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41401ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41401 diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/server-admit-policy-r2.log b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/server-admit-policy-r2.log new file mode 100644 index 000000000..efd9fac5c --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/server-admit-policy-r2.log @@ -0,0 +1,41 @@ +2026-10-11T04:23:45.719436Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins +2026-10-11T04:23:45.719482Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0 +2026-10-11T04:23:45.719638Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=8 kv_unified=false n_parallel=8 prefill_chunk_size=2048 prefill_chunk_size_contended=512 max_kv_size=None +2026-10-11T04:23:45.719706Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA) +2026-10-11T04:23:45.719708Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set) +2026-10-11T04:23:45.719710Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB +2026-10-11T04:23:45.990907Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"] +2026-10-11T04:23:45.991466Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256 +2026-10-11T04:23:46.260971Z INFO mlxcel::server::startup: Warming up model... +2026-10-11T04:23:46.261418Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model... +2026-10-11T04:23:46.585012Z INFO mlxcel::server::model_provider::model_worker: Model llama-3.2-1b-instruct-4bit loaded in 0.324s (resident after load: 0.00 GB) worker_model_id=llama-3.2-1b-instruct-4bit load_seconds=0.323568745 active_bytes=128 peak_bytes=128 cache_bytes=0 limit_bytes=124128081510 +2026-10-11T04:23:46.585238Z INFO mlxcel::server::model_provider::model_worker: EOS tokens from config: [128009] +2026-10-11T04:23:46.585254Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=8, max_queue_depth=32, prefill_chunk_size=2048, prefill_chunk_size_contended=512, max_batch_prefill=4, decode_storage=auto) +2026-10-11T04:23:46.585424Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 1559064 blocks (16 layers, 32-token blocks) +2026-10-11T04:23:46.585434Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 2048 blocks per layer (fused decode serves a layer only while its rows fit one slab) +2026-10-11T04:23:46.585451Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=16 kv_cache_mode_total_layers=16 +2026-10-11T04:23:46.586176Z INFO prefill{seq_id=seq-0 prompt_len=2 cached=0 start=0}: mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved +2026-10-11T04:23:47.070593Z INFO prefill{seq_id=seq-0 prompt_len=2 cached=0 start=0}: mlxcel::server::batch::scheduler::prefill: prompt-cache: request completed during prefill: cached=0/2 prompt tokens, total 484ms prompt_tokens=2 cached_tokens=0 generation_time_ms=484 +2026-10-11T04:23:47.070645Z INFO mlxcel::server::startup: Warmup complete +2026-10-11T04:23:47.071716Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18080 +2026-10-11T04:23:47.071732Z INFO mlxcel::server::startup: Detected 1 GPU(s) +2026-10-11T04:23:47.071736Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin) +2026-10-11T04:23:47.071738Z INFO mlxcel::server::startup: Endpoints: +2026-10-11T04:23:47.071740Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions +2026-10-11T04:23:47.071754Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions +2026-10-11T04:23:47.071755Z INFO mlxcel::server::startup: GET /v1/models - List models +2026-10-11T04:23:47.071756Z INFO mlxcel::server::startup: POST /completion - llama-server native completion +2026-10-11T04:23:47.071758Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text +2026-10-11T04:23:47.071759Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens +2026-10-11T04:23:47.071760Z INFO mlxcel::server::startup: GET /props - Server properties +2026-10-11T04:23:47.071761Z INFO mlxcel::server::startup: GET /slots - Slot status +2026-10-11T04:23:47.071763Z INFO mlxcel::server::startup: GET /health - Health check +2026-10-11T04:23:47.454645Z INFO mlxcel::server::chat_request: prompt cache on + preserve_thinking unset: defaulting preserve_thinking=true for prefix stability (override via chat_template_kwargs.preserve_thinking=false) session=__mlxcel_anon__ +2026-10-11T04:23:47.747383Z INFO decode_step{batch_size=4}: mlxcel_core::cache::paged_batch_decode: paged decode v2: gather: 656 visible KV tokens across 4 request(s) is below the 2048-token dispatch floor +2026-10-11T04:23:51.713455Z INFO decode_step{batch_size=4}: mlxcel_core::cache::paged_batch_decode: paged decode v2: fused v2 launch (batch 4, 2048 visible KV tokens, 64 chunks, merge on) +Compiling CUDA kernels (first run on this host; cached for later runs)... +2026-10-11T04:23:56.179276Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/6933 prompt tokens, total 3146ms prompt_tokens=6933 cached_tokens=0 generation_time_ms=3146 +2026-10-11T04:24:28.969452Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41401ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41401 +2026-10-11T04:24:28.969892Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41514ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41514 +2026-10-11T04:24:28.970267Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41402ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41402 +2026-10-11T04:24:28.970669Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41402ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41402 diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/server-admit-policy-r3.log b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/server-admit-policy-r3.log new file mode 100644 index 000000000..02d373a5f --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/server-admit-policy-r3.log @@ -0,0 +1,41 @@ +2026-10-11T04:25:30.129131Z WARN mlxcel::server::startup: CORS is set to allow all origins ('*') and no API key is set; this can be a security risk (cross-origin attacks). Set --api-key, or narrow --cors-origins / --allowed-origins +2026-10-11T04:25:30.129179Z INFO mlxcel::server::startup: effective KV cache mode kv_cache_mode=fp16 kv_bits=0 +2026-10-11T04:25:30.129291Z INFO mlxcel::server::startup: resolved context and batch geometry (0 = the checkpoint's own trained context) ctx_size=0 ctx_size_per_slot=0 context_slots=8 kv_unified=false n_parallel=8 prefill_chunk_size=2048 prefill_chunk_size_contended=512 max_kv_size=None +2026-10-11T04:25:30.129359Z INFO mlxcel::server::startup: Runtime device: NVIDIA GPU (CUDA) +2026-10-11T04:25:30.129361Z INFO mlxcel::server::startup: CUDA graph-cache LRU capacity: MLX_CUDA_GRAPH_CACHE_SIZE=2000 (mlxcel raises MLX's default of 400 to 2000 so long-lived, shape-diverse decode does not hit the cache-thrashing abort from issue #818, unless an operator override is set) +2026-10-11T04:25:30.129363Z INFO mlxcel::server::startup: Wired memory limit: 121.7 GB +2026-10-11T04:25:30.402697Z INFO mlxcel::server::startup: DRY sequence breakers active (b10621 semantics: breaker token data derived from the vocabulary per request) breakers=["\n", ":", "\"", "*"] +2026-10-11T04:25:30.403208Z INFO mlxcel::server::startup: Prompt-prefix cache store enabled (+ APC, snapshots) capacity_bytes=2147483648 max_entries=1024 ttl_seconds=3600 snapshot_capacity_bytes=536870912 snapshot_max_entries=4096 snapshot_ttl_seconds=7200 min_prefix_tokens=32 apc_enabled=true apc_block_size=16 apc_hash=sha256 +2026-10-11T04:25:30.673061Z INFO mlxcel::server::startup: Warming up model... +2026-10-11T04:25:30.673332Z INFO mlxcel::server::model_provider::model_worker: Model worker thread starting, loading model... +2026-10-11T04:25:30.997245Z INFO mlxcel::server::model_provider::model_worker: Model llama-3.2-1b-instruct-4bit loaded in 0.324s (resident after load: 0.00 GB) worker_model_id=llama-3.2-1b-instruct-4bit load_seconds=0.323868789 active_bytes=128 peak_bytes=128 cache_bytes=0 limit_bytes=124128081510 +2026-10-11T04:25:30.997464Z INFO mlxcel::server::model_provider::model_worker: EOS tokens from config: [128009] +2026-10-11T04:25:30.997480Z INFO mlxcel::server::model_provider::model_worker: Starting BatchScheduler (max_batch_size=8, max_queue_depth=32, prefill_chunk_size=2048, prefill_chunk_size_contended=512, max_batch_prefill=4, decode_storage=auto) +2026-10-11T04:25:30.997640Z INFO mlxcel::server::model_provider::model_worker: Paged KV block budget: 1559064 blocks (16 layers, 32-token blocks) +2026-10-11T04:25:30.997651Z INFO mlxcel::server::model_provider::model_worker: Paged KV slab size: 2048 blocks per layer (fused decode serves a layer only while its rows fit one slab) +2026-10-11T04:25:30.997669Z INFO mlxcel::server::batch::scheduler::paged_layout: resolved KV cache mode applied to model caches kv_cache_mode_effective=fp16 kv_cache_mode_applied_layers=16 kv_cache_mode_total_layers=16 +2026-10-11T04:25:30.998394Z INFO prefill{seq_id=seq-0 prompt_len=2 cached=0 start=0}: mlxcel_core::sampling_dispatch: sampling dispatch: argmax: greedy path (temperature 0, top_k 1); no sampling kernel involved +2026-10-11T04:25:31.467524Z INFO prefill{seq_id=seq-0 prompt_len=2 cached=0 start=0}: mlxcel::server::batch::scheduler::prefill: prompt-cache: request completed during prefill: cached=0/2 prompt tokens, total 469ms prompt_tokens=2 cached_tokens=0 generation_time_ms=469 +2026-10-11T04:25:31.467572Z INFO mlxcel::server::startup: Warmup complete +2026-10-11T04:25:31.468395Z INFO mlxcel::server::startup: Starting mlxcel server on http://127.0.0.1:18080 +2026-10-11T04:25:31.468400Z INFO mlxcel::server::startup: Detected 1 GPU(s) +2026-10-11T04:25:31.468403Z INFO mlxcel::server::startup: CUDA compute capability 12.1 (sm_121); compiled for [121] (cubin) +2026-10-11T04:25:31.468405Z INFO mlxcel::server::startup: Endpoints: +2026-10-11T04:25:31.468406Z INFO mlxcel::server::startup: POST /v1/chat/completions - OpenAI chat completions +2026-10-11T04:25:31.468419Z INFO mlxcel::server::startup: POST /v1/completions - OpenAI text completions +2026-10-11T04:25:31.468420Z INFO mlxcel::server::startup: GET /v1/models - List models +2026-10-11T04:25:31.468421Z INFO mlxcel::server::startup: POST /completion - llama-server native completion +2026-10-11T04:25:31.468423Z INFO mlxcel::server::startup: POST /tokenize - Tokenize text +2026-10-11T04:25:31.468424Z INFO mlxcel::server::startup: POST /detokenize - Detokenize tokens +2026-10-11T04:25:31.468425Z INFO mlxcel::server::startup: GET /props - Server properties +2026-10-11T04:25:31.468426Z INFO mlxcel::server::startup: GET /slots - Slot status +2026-10-11T04:25:31.468428Z INFO mlxcel::server::startup: GET /health - Health check +2026-10-11T04:25:31.856178Z INFO mlxcel::server::chat_request: prompt cache on + preserve_thinking unset: defaulting preserve_thinking=true for prefix stability (override via chat_template_kwargs.preserve_thinking=false) session=__mlxcel_anon__ +2026-10-11T04:25:32.201186Z INFO decode_step{batch_size=4}: mlxcel_core::cache::paged_batch_decode: paged decode v2: gather: 656 visible KV tokens across 4 request(s) is below the 2048-token dispatch floor +2026-10-11T04:25:36.359487Z INFO decode_step{batch_size=4}: mlxcel_core::cache::paged_batch_decode: paged decode v2: fused v2 launch (batch 4, 2048 visible KV tokens, 64 chunks, merge on) +Compiling CUDA kernels (first run on this host; cached for later runs)... +2026-10-11T04:25:40.802567Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/6933 prompt tokens, total 3122ms prompt_tokens=6933 cached_tokens=0 generation_time_ms=3122 +2026-10-11T04:26:13.492260Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41635ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41635 +2026-10-11T04:26:13.492502Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41328ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41328 +2026-10-11T04:26:13.492896Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41635ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41635 +2026-10-11T04:26:13.493170Z INFO mlxcel::server::batch::scheduler::decode_tick: prompt-cache: request completed: cached=0/163 prompt tokens, total 41636ms prompt_tokens=163 cached_tokens=0 generation_time_ms=41636 diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-llama-3.2-1b-instruct-4bit.jsonl b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-llama-3.2-1b-instruct-4bit.jsonl new file mode 100644 index 000000000..d0debe093 --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-llama-3.2-1b-instruct-4bit.jsonl @@ -0,0 +1,16 @@ +{"kind": "header", "model": "/home/inureyes/Development/backend.ai/wt-auto-20261011T0224-issue-2228/models/mlx/llama-3.2-1b-instruct-4bit", "bin": "/home/inureyes/Development/backend.ai/wt-auto-20261011T0224-issue-2228/target/release/mlxcel-bench-engine", "arms": {"c2048": ["--path", "server", "--prefill-chunk", "2048"], "policy": ["--path", "server"]}, "prompt_tokens": [8192], "max_tokens": 128, "rounds": 5, "host": "spark-102", "time": 1791693709.2914042, "sdpa_deterministic": null} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 413.045263, "prefill_tok_s": 19833.177459778784, "decode_ms": 621.9840439999999, "decode_tok_s": 205.7930605049412, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048", "server_prompt_eval_ms": 412, "server_generation_ms": 622, "paged_decode_launches": 2064, "arm": "c2048", "round": 0, "started": 1791693764.727204, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 413.008809, "prefill_tok_s": 19834.92802450129, "decode_ms": 619.1212270000001, "decode_tok_s": 206.74464776508137, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "policy", "server_prompt_eval_ms": 412, "server_generation_ms": 619, "paged_decode_launches": 2064, "arm": "policy", "round": 0, "started": 1791693825.5921834, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 412.72171000000003, "prefill_tok_s": 19848.725670379685, "decode_ms": 627.567355, "decode_tok_s": 203.96217072189805, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048-null", "server_prompt_eval_ms": 412, "server_generation_ms": 627, "paged_decode_launches": 2064, "arm": "c2048-null", "round": 0, "started": 1791693886.4224012, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 415.417377, "prefill_tok_s": 19719.926159949733, "decode_ms": 620.8226999999999, "decode_tok_s": 206.1780279619286, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "policy", "server_prompt_eval_ms": 414, "server_generation_ms": 621, "paged_decode_launches": 2064, "arm": "policy", "round": 1, "started": 1791693947.2330053, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 414.680518, "prefill_tok_s": 19754.96712387149, "decode_ms": 620.590631, "decode_tok_s": 206.25512794762125, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048-null", "server_prompt_eval_ms": 414, "server_generation_ms": 620, "paged_decode_launches": 2064, "arm": "c2048-null", "round": 1, "started": 1791694008.0987415, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 416.261781, "prefill_tok_s": 19679.92348545686, "decode_ms": 641.6832380000001, "decode_tok_s": 199.47536793847183, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048", "server_prompt_eval_ms": 415, "server_generation_ms": 642, "paged_decode_launches": 2064, "arm": "c2048", "round": 1, "started": 1791694068.9750288, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 416.024281, "prefill_tok_s": 19691.158362941802, "decode_ms": 621.531952, "decode_tok_s": 205.94275095289066, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048-null", "server_prompt_eval_ms": 415, "server_generation_ms": 622, "paged_decode_launches": 2064, "arm": "c2048-null", "round": 2, "started": 1791694130.0758228, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 416.147761, "prefill_tok_s": 19685.315572321437, "decode_ms": 621.39603, "decode_tok_s": 205.98779815184852, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048", "server_prompt_eval_ms": 415, "server_generation_ms": 622, "paged_decode_launches": 2064, "arm": "c2048", "round": 2, "started": 1791694190.8663633, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 413.289136, "prefill_tok_s": 19821.474329777688, "decode_ms": 624.091996, "decode_tok_s": 205.09796764001442, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "policy", "server_prompt_eval_ms": 412, "server_generation_ms": 625, "paged_decode_launches": 2064, "arm": "policy", "round": 2, "started": 1791694251.6247394, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 416.503978, "prefill_tok_s": 19668.479612936615, "decode_ms": 621.135008, "decode_tok_s": 206.07436121198307, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048", "server_prompt_eval_ms": 416, "server_generation_ms": 621, "paged_decode_launches": 2064, "arm": "c2048", "round": 3, "started": 1791694312.4692285, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 409.813924, "prefill_tok_s": 19989.559944771423, "decode_ms": 623.385155, "decode_tok_s": 205.33052314984945, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "policy", "server_prompt_eval_ms": 409, "server_generation_ms": 623, "paged_decode_launches": 2064, "arm": "policy", "round": 3, "started": 1791694373.2444339, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 411.859882, "prefill_tok_s": 19890.259668456856, "decode_ms": 623.914382, "decode_tok_s": 205.15635428964993, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048-null", "server_prompt_eval_ms": 411, "server_generation_ms": 624, "paged_decode_launches": 2064, "arm": "c2048-null", "round": 3, "started": 1791694434.1152246, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 419.552136, "prefill_tok_s": 19525.582870587506, "decode_ms": 624.064049, "decode_tok_s": 205.10715239102004, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "policy", "server_prompt_eval_ms": 419, "server_generation_ms": 624, "paged_decode_launches": 2064, "arm": "policy", "round": 4, "started": 1791694494.9493015, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 412.060947, "prefill_tok_s": 19880.554222965467, "decode_ms": 620.2828569999999, "decode_tok_s": 206.3574682993375, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048-null", "server_prompt_eval_ms": 411, "server_generation_ms": 621, "paged_decode_launches": 2064, "arm": "c2048-null", "round": 4, "started": 1791694555.7313607, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "llama-3.2-1b-instruct-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 412.961287, "prefill_tok_s": 19837.21055189369, "decode_ms": 624.27017, "decode_tok_s": 205.0394302838465, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048", "server_prompt_eval_ms": 412, "server_generation_ms": 624, "paged_decode_launches": 2064, "arm": "c2048", "round": 4, "started": 1791694616.5659828, "gate_wait_s": 0, "ci_job_running": false} diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-llama-3.2-1b-instruct-4bit.txt b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-llama-3.2-1b-instruct-4bit.txt new file mode 100644 index 000000000..5b99b0cee --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-llama-3.2-1b-instruct-4bit.txt @@ -0,0 +1,28 @@ +round 0 c2048 server prompt 8192 TTFT 413.05 ms decode 205.79 tok/s +round 0 policy server prompt 8192 TTFT 413.01 ms decode 206.74 tok/s +round 0 c2048-null server prompt 8192 TTFT 412.72 ms decode 203.96 tok/s +round 1 policy server prompt 8192 TTFT 415.42 ms decode 206.18 tok/s +round 1 c2048-null server prompt 8192 TTFT 414.68 ms decode 206.26 tok/s +round 1 c2048 server prompt 8192 TTFT 416.26 ms decode 199.48 tok/s +round 2 c2048-null server prompt 8192 TTFT 416.02 ms decode 205.94 tok/s +round 2 c2048 server prompt 8192 TTFT 416.15 ms decode 205.99 tok/s +round 2 policy server prompt 8192 TTFT 413.29 ms decode 205.10 tok/s +round 3 c2048 server prompt 8192 TTFT 416.50 ms decode 206.07 tok/s +round 3 policy server prompt 8192 TTFT 409.81 ms decode 205.33 tok/s +round 3 c2048-null server prompt 8192 TTFT 411.86 ms decode 205.16 tok/s +round 4 policy server prompt 8192 TTFT 419.55 ms decode 205.11 tok/s +round 4 c2048-null server prompt 8192 TTFT 412.06 ms decode 206.36 tok/s +round 4 c2048 server prompt 8192 TTFT 412.96 ms decode 205.04 tok/s + +arm medians (min..max) per path and prompt length + server prompt 8192 c2048 n=5 ttft_ms 416.15 (412.96..416.50) decode_tok_s 205.79 (199.48..206.07) + server prompt 8192 policy n=5 ttft_ms 413.29 (409.81..419.55) decode_tok_s 205.33 (205.10..206.74) + server prompt 8192 c2048-null n=5 ttft_ms 412.72 (411.86..416.02) decode_tok_s 205.94 (203.96..206.36) + +paired per-round deltas, arm vs reference (positive = higher than the reference) + server prompt 8192 policy vs c2048 ttft_ms median -0.20% range -1.61%..+1.60% + server prompt 8192 policy vs c2048 decode_tok_s median +0.03% range -0.43%..+3.36% + server prompt 8192 c2048-null vs c2048 ttft_ms median -0.22% range -1.12%..-0.03% + server prompt 8192 c2048-null vs c2048 decode_tok_s median -0.02% range -0.89%..+3.40% + +the *-null rows are the noise floor: an A/B delta whose range overlaps the null range is unresolved diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-qwen3-1.7b-4bit.jsonl b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-qwen3-1.7b-4bit.jsonl new file mode 100644 index 000000000..e54443e48 --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-qwen3-1.7b-4bit.jsonl @@ -0,0 +1,16 @@ +{"kind": "header", "model": "/home/inureyes/Development/backend.ai/wt-auto-20261011T0224-issue-2228/models/mlx/qwen3-1.7b-4bit", "bin": "/home/inureyes/Development/backend.ai/wt-auto-20261011T0224-issue-2228/target/release/mlxcel-bench-engine", "arms": {"c2048": ["--path", "server", "--prefill-chunk", "2048"], "policy": ["--path", "server"]}, "prompt_tokens": [8192], "max_tokens": 128, "rounds": 5, "host": "spark-102", "time": 1791692778.8512638, "sdpa_deterministic": null} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 657.001437, "prefill_tok_s": 12468.7702928114, "decode_ms": 1208.0468609999998, "decode_tok_s": 105.95615462635601, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048", "server_prompt_eval_ms": 656, "server_generation_ms": 1208, "paged_decode_launches": 3612, "arm": "c2048", "round": 0, "started": 1791692834.3659868, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 663.413455, "prefill_tok_s": 12348.257241782954, "decode_ms": 1205.579938, "decode_tok_s": 106.17296785175932, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "policy", "server_prompt_eval_ms": 662, "server_generation_ms": 1206, "paged_decode_launches": 3612, "arm": "policy", "round": 0, "started": 1791692897.9875617, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 649.492796, "prefill_tok_s": 12612.918958380564, "decode_ms": 1203.453822, "decode_tok_s": 106.36054135195559, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048-null", "server_prompt_eval_ms": 648, "server_generation_ms": 1204, "paged_decode_launches": 3612, "arm": "c2048-null", "round": 0, "started": 1791692959.9415061, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 724.888686, "prefill_tok_s": 11301.04546837968, "decode_ms": 1198.2975959999999, "decode_tok_s": 106.81820645161338, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "policy", "server_prompt_eval_ms": 724, "server_generation_ms": 1198, "paged_decode_launches": 3612, "arm": "policy", "round": 1, "started": 1791693021.8345582, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 694.0903440000001, "prefill_tok_s": 11802.498148569546, "decode_ms": 1191.998148, "decode_tok_s": 107.38271717516126, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048-null", "server_prompt_eval_ms": 693, "server_generation_ms": 1192, "paged_decode_launches": 3612, "arm": "c2048-null", "round": 1, "started": 1791693083.8749285, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 727.136342, "prefill_tok_s": 11266.112731303974, "decode_ms": 1248.832851, "decode_tok_s": 102.49570220506635, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048", "server_prompt_eval_ms": 726, "server_generation_ms": 1249, "paged_decode_launches": 3612, "arm": "c2048", "round": 1, "started": 1791693145.7642043, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 723.553861, "prefill_tok_s": 11321.893837561873, "decode_ms": 1202.486764, "decode_tok_s": 106.44607810419109, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048-null", "server_prompt_eval_ms": 722, "server_generation_ms": 1202, "paged_decode_launches": 3612, "arm": "c2048-null", "round": 2, "started": 1791693207.8142908, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 700.703709, "prefill_tok_s": 11691.104092614414, "decode_ms": 1194.835257, "decode_tok_s": 107.12773936834039, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048", "server_prompt_eval_ms": 700, "server_generation_ms": 1195, "paged_decode_launches": 3612, "arm": "c2048", "round": 2, "started": 1791693269.7957366, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 703.819892, "prefill_tok_s": 11639.34138991343, "decode_ms": 1194.6292389999999, "decode_tok_s": 107.14621392252731, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "policy", "server_prompt_eval_ms": 703, "server_generation_ms": 1194, "paged_decode_launches": 3612, "arm": "policy", "round": 2, "started": 1791693331.6423373, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 688.763389, "prefill_tok_s": 11893.779679395822, "decode_ms": 1201.118264, "decode_tok_s": 106.56735796667563, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048", "server_prompt_eval_ms": 688, "server_generation_ms": 1201, "paged_decode_launches": 3612, "arm": "c2048", "round": 3, "started": 1791693393.6056685, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 662.626815, "prefill_tok_s": 12362.916523382773, "decode_ms": 1194.955148, "decode_tok_s": 107.11699113915195, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "policy", "server_prompt_eval_ms": 662, "server_generation_ms": 1195, "paged_decode_launches": 3612, "arm": "policy", "round": 3, "started": 1791693455.5054328, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 656.188998, "prefill_tok_s": 12484.208093961368, "decode_ms": 1195.3598969999998, "decode_tok_s": 107.08072131350748, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048-null", "server_prompt_eval_ms": 655, "server_generation_ms": 1195, "paged_decode_launches": 3612, "arm": "c2048-null", "round": 3, "started": 1791693517.360133, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 650.04821, "prefill_tok_s": 12602.14223188154, "decode_ms": 1192.051185, "decode_tok_s": 107.37793948000648, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "policy", "server_prompt_eval_ms": 649, "server_generation_ms": 1192, "paged_decode_launches": 3612, "arm": "policy", "round": 4, "started": 1791693579.1924503, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 655.735608, "prefill_tok_s": 12492.83994960359, "decode_ms": 1195.556482, "decode_tok_s": 107.06311406205901, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048-null", "server_prompt_eval_ms": 655, "server_generation_ms": 1195, "paged_decode_launches": 3612, "arm": "c2048-null", "round": 4, "started": 1791693641.0723782, "gate_wait_s": 0, "ci_job_running": false} +{"path": "server", "model": "qwen3-1.7b-4bit", "prompt_target_len": 8192, "prompt_tokens": 8192, "generated_tokens": 128, "ttft_ms": 649.952589, "prefill_tok_s": 12603.996258564024, "decode_ms": 1188.045706, "decode_tok_s": 107.73996265763196, "prefill_chunk": 2048, "decode_storage": "paged", "max_tokens": 128, "label": "c2048", "server_prompt_eval_ms": 649, "server_generation_ms": 1188, "paged_decode_launches": 3612, "arm": "c2048", "round": 4, "started": 1791693702.9438837, "gate_wait_s": 0, "ci_job_running": false} diff --git a/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-qwen3-1.7b-4bit.txt b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-qwen3-1.7b-4bit.txt new file mode 100644 index 000000000..2c46a2280 --- /dev/null +++ b/docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/ttft-qwen3-1.7b-4bit.txt @@ -0,0 +1,28 @@ +round 0 c2048 server prompt 8192 TTFT 657.00 ms decode 105.96 tok/s +round 0 policy server prompt 8192 TTFT 663.41 ms decode 106.17 tok/s +round 0 c2048-null server prompt 8192 TTFT 649.49 ms decode 106.36 tok/s +round 1 policy server prompt 8192 TTFT 724.89 ms decode 106.82 tok/s +round 1 c2048-null server prompt 8192 TTFT 694.09 ms decode 107.38 tok/s +round 1 c2048 server prompt 8192 TTFT 727.14 ms decode 102.50 tok/s +round 2 c2048-null server prompt 8192 TTFT 723.55 ms decode 106.45 tok/s +round 2 c2048 server prompt 8192 TTFT 700.70 ms decode 107.13 tok/s +round 2 policy server prompt 8192 TTFT 703.82 ms decode 107.15 tok/s +round 3 c2048 server prompt 8192 TTFT 688.76 ms decode 106.57 tok/s +round 3 policy server prompt 8192 TTFT 662.63 ms decode 107.12 tok/s +round 3 c2048-null server prompt 8192 TTFT 656.19 ms decode 107.08 tok/s +round 4 policy server prompt 8192 TTFT 650.05 ms decode 107.38 tok/s +round 4 c2048-null server prompt 8192 TTFT 655.74 ms decode 107.06 tok/s +round 4 c2048 server prompt 8192 TTFT 649.95 ms decode 107.74 tok/s + +arm medians (min..max) per path and prompt length + server prompt 8192 c2048 n=5 ttft_ms 688.76 (649.95..727.14) decode_tok_s 106.57 (102.50..107.74) + server prompt 8192 policy n=5 ttft_ms 663.41 (650.05..724.89) decode_tok_s 107.12 (106.17..107.38) + server prompt 8192 c2048-null n=5 ttft_ms 656.19 (649.49..723.55) decode_tok_s 107.06 (106.36..107.38) + +paired per-round deltas, arm vs reference (positive = higher than the reference) + server prompt 8192 policy vs c2048 ttft_ms median +0.01% range -3.79%..+0.98% + server prompt 8192 policy vs c2048 decode_tok_s median +0.20% range -0.34%..+4.22% + server prompt 8192 c2048-null vs c2048 ttft_ms median -1.14% range -4.73%..+3.26% + server prompt 8192 c2048-null vs c2048 decode_tok_s median +0.38% range -0.64%..+4.77% + +the *-null rows are the noise floor: an A/B delta whose range overlaps the null range is unresolved diff --git a/src/bin/engine_parity.rs b/src/bin/engine_parity.rs index 54b3af631..bb5e3d88f 100644 --- a/src/bin/engine_parity.rs +++ b/src/bin/engine_parity.rs @@ -341,6 +341,7 @@ fn main() -> Result<()> { // `--server-prefill-chunk` pins the chunk on every server arm. if let Some(chunk) = args.server_prefill_chunk { startup.prefill_chunk_size = chunk; + startup.prefill_chunk_size_contended = chunk; } let server = InProcessServer::start(&startup)?; for case in &cases { diff --git a/src/bin/mlx_server.rs b/src/bin/mlx_server.rs index 89aa50f5a..fcb0bbb91 100644 --- a/src/bin/mlx_server.rs +++ b/src/bin/mlx_server.rs @@ -827,13 +827,12 @@ struct ServerArgs { )] rerank_batch_size: usize, - /// Prefill chunk size in tokens (0 = disabled). Defaults to the shared - /// chunk policy: `MLXCEL_PREFILL_CHUNK`, else 2048 (ADR 0007). - #[arg( - long = "prefill-chunk-size", - default_value_t = mlxcel_core::prefill_plan::prefill_chunk_len() - )] - prefill_chunk_size: usize, + /// Prefill chunk size in tokens (0 = disabled). + /// + /// [default: 2048 when no other sequence is decoding, 512 when one is; + /// MLXCEL_PREFILL_CHUNK overrides] + #[arg(long = "prefill-chunk-size")] + prefill_chunk_size: Option, /// Decode ticks a parked chunked prefill yields before it is granted one /// (#1011). @@ -853,7 +852,7 @@ struct ServerArgs { #[arg(long = "prefill-grant-interval", value_name = "N")] prefill_grant_interval: Option, - /// Prefill batch size [llama-server alias for --prefill-chunk-size] [default: 2048] + /// Prefill batch size [llama-server alias for --prefill-chunk-size] #[arg( short = 'b', long = "batch-size", diff --git a/src/commands/serve_tests.rs b/src/commands/serve_tests.rs index 7a0d19dca..bc807009e 100644 --- a/src/commands/serve_tests.rs +++ b/src/commands/serve_tests.rs @@ -65,7 +65,7 @@ fn sample_args() -> crate::ServeArgs { embedding_request_timeout_secs: 120, reranker_model: None, rerank_batch_size: 0, - prefill_chunk_size: 512, + prefill_chunk_size: Some(512), prefill_grant_interval: None, batch_size: None, ubatch_size: None, diff --git a/src/lib/mlxcel-core/src/prefill_plan.rs b/src/lib/mlxcel-core/src/prefill_plan.rs index 3cb3148cd..d9850bdc6 100644 --- a/src/lib/mlxcel-core/src/prefill_plan.rs +++ b/src/lib/mlxcel-core/src/prefill_plan.rs @@ -45,9 +45,10 @@ //! always; embedding input only where the executor extends the embedding //! rows). A piece ending at the history boundary is never padded. //! -//! The chunk value is one policy: [`prefill_chunk_len`], `MLXCEL_PREFILL_CHUNK` -//! or [`DEFAULT_PREFILL_CHUNK`], on both front ends. The server's -//! `--prefill-chunk-size` overrides it per process. +//! The direct engine uses [`prefill_chunk_len`]. The server uses +//! [`PrefillChunkPolicy`]: the larger default while no other sequence is +//! decoding and a smaller chunk while a decode batch is live. An explicit +//! `--prefill-chunk-size` or `MLXCEL_PREFILL_CHUNK` pins both states. use std::ops::Range; @@ -59,6 +60,49 @@ use crate::utils::align_to_na_tile; /// both paths and both measured models. pub const DEFAULT_PREFILL_CHUNK: usize = 2048; +/// Prefill chunk used while at least one other sequence is decoding. +/// +/// PR #2226 measured this value at less than half the live-stream ITL p95 of +/// 2048 while admitting an 8192-token prompt. The larger default remains the +/// faster time-to-first-token choice when no decode stream needs protection. +pub const CONTENDED_PREFILL_CHUNK: usize = 512; + +/// Server prefill chunk sizes for idle and decode-contended scheduler ticks. +#[derive(Debug, Clone, Copy, PartialEq, Eq)] +pub struct PrefillChunkPolicy { + /// Chunk used when no other sequence is decoding. + pub alone: usize, + /// Chunk used while at least one other sequence is decoding. + pub contended: usize, +} + +impl PrefillChunkPolicy { + /// Resolve an explicit override or the adaptive server defaults. + #[must_use] + pub fn resolve(explicit: Option) -> Self { + match explicit { + Some(chunk) => Self { + alone: chunk, + contended: chunk, + }, + None => Self { + alone: DEFAULT_PREFILL_CHUNK, + contended: CONTENDED_PREFILL_CHUNK, + }, + } + } + + /// Select the chunk for the scheduler's current decode state. + #[must_use] + pub fn chunk(&self, others_decoding: bool) -> usize { + if others_decoding { + self.contended + } else { + self.alone + } + } +} + /// The chunk policy: `MLXCEL_PREFILL_CHUNK` (tokens, `0` forces a single-pass /// prefill), else [`DEFAULT_PREFILL_CHUNK`]. Read once per process. /// @@ -73,25 +117,30 @@ pub const DEFAULT_PREFILL_CHUNK: usize = 2048; /// default for `--prefill-chunk-size`, the Gemma 4 MTP prefill and the engine /// benchmark. pub fn prefill_chunk_len() -> usize { - static CHUNK: std::sync::OnceLock = std::sync::OnceLock::new(); + prefill_chunk_override().unwrap_or(DEFAULT_PREFILL_CHUNK) +} + +/// Explicit `MLXCEL_PREFILL_CHUNK` override, read once per process. +/// +/// `0` is a valid override that disables chunking. An unset or invalid value +/// returns `None`; invalid text emits the same warning as before. +pub fn prefill_chunk_override() -> Option { + static CHUNK: std::sync::OnceLock> = std::sync::OnceLock::new(); *CHUNK .get_or_init(|| parse_prefill_chunk(std::env::var("MLXCEL_PREFILL_CHUNK").ok().as_deref())) } -/// The chunk for an `MLXCEL_PREFILL_CHUNK` value: unset gives the default, and -/// a value that is not a non-negative integer logs one warning naming the -/// variable and the value before falling back to the default. -fn parse_prefill_chunk(value: Option<&str>) -> usize { - let Some(raw) = value else { - return DEFAULT_PREFILL_CHUNK; - }; +/// Parse an `MLXCEL_PREFILL_CHUNK` value. Unset and invalid values are not +/// overrides; invalid text logs one warning naming the variable and value. +fn parse_prefill_chunk(value: Option<&str>) -> Option { + let raw = value?; match raw.trim().parse::() { - Ok(chunk) => chunk, + Ok(chunk) => Some(chunk), Err(_) => { tracing::warn!( "MLXCEL_PREFILL_CHUNK={raw:?} is not a non-negative integer; using the default of {DEFAULT_PREFILL_CHUNK} tokens" ); - DEFAULT_PREFILL_CHUNK + None } } } @@ -257,41 +306,18 @@ impl PrefillPlan { chunk: usize, caps: PrefillCaps, ) -> Self { - let start = adopted.min(prompt_len); - // Embedding input is consumed whole and never split, so a history - // boundary does not apply to it: honoring one would cut the embedding - // rows into two forwards and advance a cursor the executor does not - // keep. - let boundary = - boundary.filter(|b| !caps.is_embedding_input() && *b > start && *b < prompt_len); + let (start, boundary, can_pad, pad_mask) = + plan_metadata(prompt_len, adopted, boundary, caps); let remaining = prompt_len - start; - let chunk = (chunk > 0 - && caps.supports_chunked_prefill - && !caps.is_embedding_input() - && remaining > chunk) - .then_some(chunk); - let can_pad = caps.can_pad(); // the model's own supports_padded_prefill() gate, see PrefillCaps::can_pad - let pad_mask = !caps.supports_maskless_padded_prefill || force_padded_prefill_array_mask(); - + let chunk = eligible_chunk(chunk, caps).filter(|chunk| remaining > *chunk); let mut pieces = Vec::new(); let mut cursor = start; if let Some(boundary) = boundary { - pieces.push(PrefillPiece { - range: cursor..boundary, - padded_len: boundary - cursor, - }); + pieces.push(piece_to_boundary(cursor, boundary)); cursor = boundary; } let step = chunk.unwrap_or(remaining.max(1)); - while cursor < prompt_len { - let end = (cursor + step).min(prompt_len); - let len = end - cursor; - pieces.push(PrefillPiece { - range: cursor..end, - padded_len: if can_pad { align_to_na_tile(len) } else { len }, - }); - cursor = end; - } + append_pieces(&mut pieces, cursor, prompt_len, step, can_pad); Self { prompt_len, @@ -303,6 +329,46 @@ impl PrefillPlan { } } + /// Rebuild a prefill plan at `cursor`, allowing the chunk to change + /// between scheduler ticks without requiring the cursor to be a split + /// point of the new partition. + /// + /// The initial call (`cursor <= adopted`) is exactly [`Self::with_prefix`]. + /// A continuation retains that plan's adopted-prefix and history-boundary + /// metadata, then partitions only `[cursor, prompt_len)`. The boundary + /// segment has already run before a continuation is parked. + #[must_use] + pub fn resume_at( + prompt_len: usize, + adopted: usize, + boundary: Option, + cursor: usize, + chunk: usize, + caps: PrefillCaps, + ) -> Self { + let (adopted, boundary, can_pad, pad_mask) = + plan_metadata(prompt_len, adopted, boundary, caps); + if cursor <= adopted { + return Self::with_prefix(prompt_len, adopted, boundary, chunk, caps); + } + + debug_assert!(cursor <= prompt_len); + debug_assert!(boundary.is_none_or(|boundary| cursor >= boundary)); + let cursor = cursor.min(prompt_len); + let chunk = eligible_chunk(chunk, caps); + let mut pieces = Vec::new(); + let step = chunk.unwrap_or((prompt_len - cursor).max(1)); + append_pieces(&mut pieces, cursor, prompt_len, step, can_pad); + Self { + prompt_len, + adopted, + boundary, + chunk, + pad_mask, + pieces, + } + } + /// Prompt length the plan covers, adopted prefix included. #[must_use] pub fn prompt_len(&self) -> usize { @@ -358,6 +424,43 @@ impl PrefillPlan { self.pieces.iter().find(|p| p.range.start == cursor) } + /// Compute only the piece that would start at the cursor in + /// Self::resume_at. + /// + /// This is the allocation-free form for capacity and reservation checks + /// that need the next piece's padding but do not execute the full suffix. + /// It follows the same initial-plan, continuation, boundary and padding + /// rules as Self::resume_at. + #[must_use] + pub fn next_piece_at( + prompt_len: usize, + adopted: usize, + boundary: Option, + cursor: usize, + chunk: usize, + caps: PrefillCaps, + ) -> Option { + let (adopted, boundary, can_pad) = plan_geometry(prompt_len, adopted, boundary, caps); + if cursor < adopted || cursor >= prompt_len { + return None; + } + if cursor == adopted { + if let Some(boundary) = boundary { + return Some(piece_to_boundary(cursor, boundary)); + } + let remaining = prompt_len - cursor; + let step = eligible_chunk(chunk, caps) + .filter(|chunk| remaining > *chunk) + .unwrap_or(remaining.max(1)); + return piece_from_cursor(cursor, prompt_len, step, can_pad); + } + + debug_assert!(boundary.is_none_or(|boundary| cursor >= boundary)); + let remaining = prompt_len - cursor; + let step = eligible_chunk(chunk, caps).unwrap_or(remaining.max(1)); + piece_from_cursor(cursor, prompt_len, step, can_pad) + } + /// Whether `piece` is the last one. #[must_use] pub fn is_terminal(&self, piece: &PrefillPiece) -> bool { @@ -413,6 +516,75 @@ impl PrefillPlan { } } +fn plan_metadata( + prompt_len: usize, + adopted: usize, + boundary: Option, + caps: PrefillCaps, +) -> (usize, Option, bool, bool) { + let (adopted, boundary, can_pad) = plan_geometry(prompt_len, adopted, boundary, caps); + let pad_mask = !caps.supports_maskless_padded_prefill || force_padded_prefill_array_mask(); + (adopted, boundary, can_pad, pad_mask) +} + +fn plan_geometry( + prompt_len: usize, + adopted: usize, + boundary: Option, + caps: PrefillCaps, +) -> (usize, Option, bool) { + let adopted = adopted.min(prompt_len); + // Embedding input is consumed whole and never split, so a history + // boundary does not apply to it: honoring one would cut the embedding + // rows into two forwards and advance a cursor the executor does not keep. + let boundary = boundary.filter(|boundary| { + !caps.is_embedding_input() && *boundary > adopted && *boundary < prompt_len + }); + let can_pad = caps.can_pad(); + (adopted, boundary, can_pad) +} + +fn eligible_chunk(chunk: usize, caps: PrefillCaps) -> Option { + (chunk > 0 && caps.supports_chunked_prefill && !caps.is_embedding_input()).then_some(chunk) +} + +fn piece_to_boundary(cursor: usize, boundary: usize) -> PrefillPiece { + PrefillPiece { + range: cursor..boundary, + padded_len: boundary - cursor, + } +} + +fn piece_from_cursor( + cursor: usize, + prompt_len: usize, + step: usize, + can_pad: bool, +) -> Option { + if cursor >= prompt_len { + return None; + } + let end = cursor.saturating_add(step).min(prompt_len); + let len = end - cursor; + Some(PrefillPiece { + range: cursor..end, + padded_len: if can_pad { align_to_na_tile(len) } else { len }, + }) +} + +fn append_pieces( + pieces: &mut Vec, + mut cursor: usize, + prompt_len: usize, + step: usize, + can_pad: bool, +) { + while let Some(piece) = piece_from_cursor(cursor, prompt_len, step, can_pad) { + cursor = piece.range.end; + pieces.push(piece); + } +} + #[cfg(test)] #[path = "prefill_plan_tests.rs"] mod tests; diff --git a/src/lib/mlxcel-core/src/prefill_plan_tests.rs b/src/lib/mlxcel-core/src/prefill_plan_tests.rs index 4f619d678..63d6afb6f 100644 --- a/src/lib/mlxcel-core/src/prefill_plan_tests.rs +++ b/src/lib/mlxcel-core/src/prefill_plan_tests.rs @@ -302,9 +302,137 @@ fn a_hit_reproduces_the_miss_only_from_a_split_point() { assert!(!PrefillPlan::with_prefix(1101, 512, None, 512, caps).reproduces(&miss)); } +#[test] +fn resume_at_matches_the_fixed_chunk_suffix_at_every_split_point() { + for chunk in [512, 2048] { + for boundary in [None, Some(777)] { + for align in [false, true] { + let caps = tokens(align); + let base = PrefillPlan::with_prefix(5000, 123, boundary, chunk, caps); + let mut cursors = base.split_points(); + cursors.push(base.prompt_len()); + for cursor in cursors { + let resumed = PrefillPlan::resume_at(5000, 123, boundary, cursor, chunk, caps); + let expected: Vec<_> = base + .pieces() + .iter() + .filter(|piece| piece.range.start >= cursor) + .cloned() + .collect(); + assert_eq!( + resumed.pieces(), + expected, + "chunk={chunk}, boundary={boundary:?}, align={align}, cursor={cursor}" + ); + assert_eq!(resumed.adopted(), base.adopted()); + assert_eq!(resumed.boundary(), base.boundary()); + assert_eq!(resumed.forwarded_len(), base.forwarded_len()); + } + } + } + } +} + +#[test] +fn resume_at_repartitions_from_the_cursor_when_the_chunk_changes() { + let plan = PrefillPlan::resume_at(8192, 0, None, 1024, 2048, tokens(false)); + assert_eq!( + ranges(&plan), + vec![1024..3072, 3072..5120, 5120..7168, 7168..8192] + ); + assert_eq!(plan.adopted(), 0); + assert_eq!(plan.forwarded_len(), 8192); +} + +#[test] +fn resumed_last_piece_still_reports_the_active_chunk() { + let plan = PrefillPlan::resume_at(1537, 0, None, 1536, 512, tokens(false)); + assert_eq!(ranges(&plan), vec![1536..1537]); + assert_eq!(plan.chunk(), Some(512)); +} + +#[test] +fn next_piece_matches_resumed_plan_across_boundaries_cursors_chunks_and_padding() { + let caps = [ + tokens(false), + tokens(true), + PrefillCaps { + supports_padded_prefill: false, + ..tokens(true) + }, + PrefillCaps { + supports_chunked_prefill: false, + ..tokens(true) + }, + tokens(true).with_input(PrefillInput::Embeddings { + executor_pads: false, + }), + tokens(true).with_input(PrefillInput::Embeddings { + executor_pads: true, + }), + ]; + for (prompt_len, adopted, boundary) in [ + (0_usize, 0_usize, None), + (70, 0, None), + (1300, 0, Some(700)), + (5000, 123, Some(777)), + ] { + let mut cursors = vec![ + adopted, + boundary.unwrap_or(adopted), + 1024, + prompt_len.saturating_sub(1), + prompt_len, + ]; + cursors.sort_unstable(); + cursors.dedup(); + for caps in caps { + for chunk in [0, 64, 512, 2048] { + for &cursor in &cursors { + if cursor < adopted + || cursor > prompt_len + || (cursor > adopted && boundary.is_some_and(|boundary| cursor < boundary)) + { + continue; + } + let plan = + PrefillPlan::resume_at(prompt_len, adopted, boundary, cursor, chunk, caps); + let expected = plan.piece_starting_at(cursor).cloned(); + assert_eq!( + PrefillPlan::next_piece_at( + prompt_len, adopted, boundary, cursor, chunk, caps, + ), + expected, + "prompt_len={prompt_len}, adopted={adopted}, boundary={boundary:?}, cursor={cursor}, chunk={chunk}, caps={caps:?}" + ); + } + } + } + } +} + +#[test] +fn chunk_policy_resolves_adaptive_defaults_and_explicit_overrides() { + assert_eq!( + PrefillChunkPolicy::resolve(None), + PrefillChunkPolicy { + alone: DEFAULT_PREFILL_CHUNK, + contended: CONTENDED_PREFILL_CHUNK, + } + ); + for chunk in [1024, 0] { + let policy = PrefillChunkPolicy::resolve(Some(chunk)); + assert_eq!(policy.alone, chunk); + assert_eq!(policy.contended, chunk); + assert_eq!(policy.chunk(false), chunk); + assert_eq!(policy.chunk(true), chunk); + } +} + #[test] fn chunk_policy_default_is_the_adr_0007_value() { assert_eq!(DEFAULT_PREFILL_CHUNK, 2048); + assert_eq!(CONTENDED_PREFILL_CHUNK, 512); // The env override is read once per process, so only the default can be // asserted here without disturbing sibling tests. assert!( @@ -313,6 +441,42 @@ fn chunk_policy_default_is_the_adr_0007_value() { ); } +#[test] +fn prefill_chunk_override_is_read_once_in_a_fresh_process() { + const CHILD: &str = "MLXCEL_PREFILL_CHUNK_TEST_CHILD"; + if std::env::var_os(CHILD).is_some() { + assert_eq!(prefill_chunk_override(), Some(1024)); + // SAFETY: this child process runs only this exact test, with one test + // thread, and no other thread reads or writes this environment key. + unsafe { std::env::set_var("MLXCEL_PREFILL_CHUNK", "512") }; + assert_eq!(prefill_chunk_override(), Some(1024)); + assert_eq!(prefill_chunk_len(), 1024); + return; + } + + let output = std::process::Command::new(std::env::current_exe().expect("current test binary")) + .args([ + "--exact", + "prefill_plan::tests::prefill_chunk_override_is_read_once_in_a_fresh_process", + "--test-threads=1", + ]) + .env(CHILD, "1") + .env("MLXCEL_PREFILL_CHUNK", "1024") + .output() + .expect("spawn isolated prefill override test"); + assert!( + output.status.success(), + "isolated test failed:\nstdout:\n{}\nstderr:\n{}", + String::from_utf8_lossy(&output.stdout), + String::from_utf8_lossy(&output.stderr) + ); + assert!( + String::from_utf8_lossy(&output.stdout).contains("1 passed"), + "isolated test did not run the exact case:\n{}", + String::from_utf8_lossy(&output.stdout) + ); +} + /// A `tracing` subscriber that records every WARN-or-worse event message. struct WarnCollector(std::sync::Arc>>); @@ -345,8 +509,8 @@ impl tracing::Subscriber for WarnCollector { } /// `MLXCEL_PREFILL_CHUNK` parses as a token count (`0` disables chunking), -/// and an unparseable value falls back to the default with exactly one -/// warning that names the variable and the value, instead of silently. +/// and an unparseable value is treated as unset with exactly one warning that +/// names the variable and the value, instead of silently. #[test] fn unparseable_chunk_env_warns_once_and_falls_back_to_the_default() { let warnings = std::sync::Arc::new(std::sync::Mutex::new(Vec::new())); @@ -360,11 +524,11 @@ fn unparseable_chunk_env_warns_once_and_falls_back_to_the_default() { parse_prefill_chunk(Some("-1")), ) }); - assert_eq!(unset, DEFAULT_PREFILL_CHUNK); - assert_eq!(zero, 0); - assert_eq!(padded, 512); - assert_eq!(bad_text, DEFAULT_PREFILL_CHUNK); - assert_eq!(negative, DEFAULT_PREFILL_CHUNK); + assert_eq!(unset, None); + assert_eq!(zero, Some(0)); + assert_eq!(padded, Some(512)); + assert_eq!(bad_text, None); + assert_eq!(negative, None); let warnings = warnings.lock().unwrap(); assert_eq!( warnings.len(), diff --git a/src/main.rs b/src/main.rs index 5689fb146..08e9cc1f2 100644 --- a/src/main.rs +++ b/src/main.rs @@ -1662,14 +1662,14 @@ pub(crate) struct ServeArgs { )] rerank_batch_size: usize, - /// Prefill chunk size in tokens (0 = disabled). Defaults to the shared - /// chunk policy: `MLXCEL_PREFILL_CHUNK`, else 2048 (ADR 0007). + /// Prefill chunk size in tokens (0 = disabled). /// /// Long prompts are broken into chunks of this size and decode steps are /// interleaved between chunks to prevent latency spikes for active - /// sequences. - #[arg(long, default_value_t = mlxcel_core::prefill_plan::prefill_chunk_len())] - prefill_chunk_size: usize, + /// sequences. [default: 2048 when no other sequence is decoding, 512 when + /// one is; MLXCEL_PREFILL_CHUNK overrides] + #[arg(long)] + prefill_chunk_size: Option, /// Decode ticks a parked chunked prefill yields before it is granted one /// (#1011). @@ -1689,7 +1689,7 @@ pub(crate) struct ServeArgs { #[arg(long, value_name = "N")] prefill_grant_interval: Option, - /// Prefill batch size [llama-server alias for --prefill-chunk-size] [default: 2048] + /// Prefill batch size [llama-server alias for --prefill-chunk-size] #[arg( short = 'b', long = "batch-size", diff --git a/src/server/batch/scheduler.rs b/src/server/batch/scheduler.rs index c7cdb6d7d..0ed15870b 100644 --- a/src/server/batch/scheduler.rs +++ b/src/server/batch/scheduler.rs @@ -255,7 +255,7 @@ impl BatchScheduler { block_size, ) .with_prefill_start_offset(prefill_start_offset) - .with_prefill_chunk_size(self.prefill_chunk_size) + .with_prefill_chunk_size(self.prefill_chunk_policy.alone) .with_prefill_boundary(prefill_boundary); Ok( crate::server::batch::speculative_slice::begin_slice_session( @@ -279,7 +279,7 @@ impl BatchScheduler { block_size, ) .with_prefill_start_offset(prefill_start_offset) - .with_prefill_chunk_size(self.prefill_chunk_size) + .with_prefill_chunk_size(self.prefill_chunk_policy.alone) .with_prefill_boundary(prefill_boundary); Ok( crate::server::batch::speculative_slice::begin_slice_session( @@ -303,7 +303,7 @@ impl BatchScheduler { block_size, ) .with_prefill_start_offset(prefill_start_offset) - .with_prefill_chunk_size(self.prefill_chunk_size) + .with_prefill_chunk_size(self.prefill_chunk_policy.alone) .with_prefill_boundary(prefill_boundary); Ok( crate::server::batch::speculative_slice::begin_slice_session( @@ -529,7 +529,7 @@ impl BatchScheduler { job.block_size, ) .with_prefill_start_offset(job.prefill_start_offset) - .with_prefill_chunk_size(self.prefill_chunk_size); + .with_prefill_chunk_size(self.prefill_chunk_policy.alone); crate::server::batch::speculative_slice::step_slice_session( adapter, &mut job, @@ -544,7 +544,7 @@ impl BatchScheduler { job.block_size, ) .with_prefill_start_offset(job.prefill_start_offset) - .with_prefill_chunk_size(self.prefill_chunk_size); + .with_prefill_chunk_size(self.prefill_chunk_policy.alone); crate::server::batch::speculative_slice::step_slice_session( adapter, &mut job, @@ -559,7 +559,7 @@ impl BatchScheduler { job.block_size, ) .with_prefill_start_offset(job.prefill_start_offset) - .with_prefill_chunk_size(self.prefill_chunk_size); + .with_prefill_chunk_size(self.prefill_chunk_policy.alone); crate::server::batch::speculative_slice::step_slice_session( adapter, &mut job, diff --git a/src/server/batch/scheduler/block_reclaim.rs b/src/server/batch/scheduler/block_reclaim.rs index 55ce2e83f..c301a6f53 100644 --- a/src/server/batch/scheduler/block_reclaim.rs +++ b/src/server/batch/scheduler/block_reclaim.rs @@ -27,6 +27,7 @@ //! forward runs. use super::*; +use mlxcel_core::prefill_plan::{PrefillCaps, PrefillInput, PrefillPlan}; /// The scheduler operations the block reclaim loop drives. pub(super) trait PagedBlockReclaimer { @@ -336,10 +337,24 @@ impl BatchScheduler { self.chunked_prefill_seq.as_ref().map_or(0, |seq| { let total = seq.prompt_tokens.len(); let remaining = total.saturating_sub(seq.prefill_offset); - let pad = self - .prefill_plan_for(seq) - .piece_starting_at(seq.prefill_offset) - .map_or(0, |piece| piece.pad_excess()); + let input = if seq.vlm_embeddings.is_some() { + PrefillInput::Embeddings { + executor_pads: false, + } + } else { + PrefillInput::Tokens + }; + let caps = PrefillCaps::for_model(self.engine.model(), should_align_prefill()) + .with_input(input); + let pad = PrefillPlan::next_piece_at( + total, + seq.prefill_start_offset, + self.history_boundary_split(seq), + seq.prefill_offset, + self.prefill_chunk_for_tick(), + caps, + ) + .map_or(0, |piece| piece.pad_excess()); self.engine .pool() .paged_blocks_to_append(seq.seq_id, remaining + pad) diff --git a/src/server/batch/scheduler/config.rs b/src/server/batch/scheduler/config.rs index c3ccf2c34..4a9e4fa5e 100644 --- a/src/server/batch/scheduler/config.rs +++ b/src/server/batch/scheduler/config.rs @@ -106,10 +106,14 @@ impl BatchScheduler { batch_metrics, batch_observability, config_eos, - prefill_chunk_size, + prefill_chunk_policy: mlxcel_core::prefill_plan::PrefillChunkPolicy::resolve(Some( + prefill_chunk_size, + )), enable_preemption, preemption_policy, chunked_prefill_seq: None, + #[cfg(test)] + prefill_forward_lengths: Vec::new(), // A one-wide decode batch has nothing to mix a prefill into, and // that is also exactly b10621's `--no-cont-batching` gate: // upstream adds pending prompts to the batch only when @@ -192,6 +196,16 @@ impl BatchScheduler { } } + /// Install the adaptive server chunk policy after the compatibility + /// constructor has derived budgets from its alone value. + pub fn with_prefill_chunk_policy( + mut self, + policy: mlxcel_core::prefill_plan::PrefillChunkPolicy, + ) -> Self { + self.prefill_chunk_policy = policy; + self + } + /// Override the #715 batched-prefill padded-token budget with the explicit /// CLI/config value (`--max-batch-prefill-tokens`). /// @@ -203,7 +217,7 @@ impl BatchScheduler { pub fn with_max_batch_prefill_tokens(mut self, configured: Option) -> Self { self.max_batch_prefill_tokens = resolve_max_batch_prefill_tokens( configured, - self.prefill_chunk_size, + self.prefill_chunk_policy.alone, self.max_batch_prefill, ); self diff --git a/src/server/batch/scheduler/mod.rs b/src/server/batch/scheduler/mod.rs index df6d4efd8..a1b6aa9cb 100644 --- a/src/server/batch/scheduler/mod.rs +++ b/src/server/batch/scheduler/mod.rs @@ -340,8 +340,8 @@ pub struct BatchScheduler { // -- Configuration -- config_eos: Vec, - /// Number of prompt tokens per prefill chunk. 0 = chunking disabled. - prefill_chunk_size: usize, + /// Prompt tokens per prefill chunk with and without live decode streams. + prefill_chunk_policy: mlxcel_core::prefill_plan::PrefillChunkPolicy, /// Whether preemptive eviction is enabled. enable_preemption: bool, /// Policy for selecting the eviction victim. @@ -352,6 +352,12 @@ pub struct BatchScheduler { /// prefill is in progress. chunked_prefill_seq: Option, + /// Logical token lengths handed to the engine by `run_prefill_piece`. + /// Test-only proof that adaptive plans are executed rather than merely + /// constructed. + #[cfg(test)] + pub(super) prefill_forward_lengths: Vec, + /// Issue #908 prototype gate, read once from `MLXCEL_MIXED_STEP` at /// construction so the tick loop never touches the environment. False (the /// default) makes [`BatchSchedulerAction::MixedStep`] unreachable and keeps diff --git a/src/server/batch/scheduler/planned_prefill.rs b/src/server/batch/scheduler/planned_prefill.rs index 908821f3c..0174e299a 100644 --- a/src/server/batch/scheduler/planned_prefill.rs +++ b/src/server/batch/scheduler/planned_prefill.rs @@ -14,11 +14,12 @@ //! The scheduler's execution of a [`PrefillPlan`] (ADR 0007, issue #2170). //! -//! The plan is a pure function of the sequence (prompt length, adopted -//! prompt-cache prefix, history boundary), the server's chunk and the model's -//! capabilities, so it is rebuilt from the sequence on every tick instead of -//! being parked next to it; `SequenceInfo::prefill_offset` is the cursor into -//! it. One tick runs one piece, so a long prompt keeps interleaving with +//! The plan is rebuilt from the sequence on every tick instead of being parked +//! next to it. Its inputs are the prompt length, adopted prompt-cache prefix, +//! history boundary, model capabilities, and the chunk selected from the live +//! scheduler state. `SequenceInfo::prefill_offset` is the resumable cursor, so +//! the chunk may change between ticks. One tick runs one piece, so a long +//! prompt keeps interleaving with //! decode (ADR 0005), with one exception that keeps the pre-plan tick shape: //! the history-boundary segment (issue #1143) is followed by the next piece in //! the same tick, because the segment is the prompt cache's snapshot point and @@ -59,7 +60,29 @@ impl BatchScheduler { /// `continue_chunked_prefill`, `chunked_prefill_reserved_blocks`, the /// prefill-role handoff. pub(super) fn prefill_plan_for(&self, seq: &SequenceInfo) -> PrefillPlan { - self.prefill_plan_for_with_chunk(seq, self.prefill_chunk_size) + let input = if seq.vlm_embeddings.is_some() { + PrefillInput::Embeddings { + executor_pads: false, + } + } else { + PrefillInput::Tokens + }; + let caps = + PrefillCaps::for_model(self.engine.model(), should_align_prefill()).with_input(input); + PrefillPlan::resume_at( + seq.prompt_tokens.len(), + seq.prefill_start_offset, + self.history_boundary_split(seq), + seq.prefill_offset, + self.prefill_chunk_for_tick(), + caps, + ) + } + + /// Chunk selected for this scheduler tick. + pub(super) fn prefill_chunk_for_tick(&self) -> usize { + self.prefill_chunk_policy + .chunk(!self.active_batch.is_empty()) } /// [`Self::prefill_plan_for`] with an explicit chunk; `0` keeps every @@ -154,6 +177,8 @@ impl BatchScheduler { "KV cache budget exhausted: no free blocks in the {total}-block KV cache budget to continue the chunked prefill" ))); } + #[cfg(test)] + self.prefill_forward_lengths.push(piece.len()); // The engine runs the piece: forward, then (for a non-terminal piece) // the forced eval that releases its transients before the next piece's // graph is built and fails just this request on an MLX throw (#822), @@ -255,11 +280,12 @@ impl BatchScheduler { /// first piece starts after it. pub(super) fn start_chunked_prefill(&mut self, mut seq: SequenceInfo) { let plan = self.prefill_plan_for(&seq); + let selected_chunk = self.prefill_chunk_for_tick(); let _span = tracing::info_span!( "chunked_prefill_start", seq_id = %seq.seq_id, prompt_len = seq.prompt_tokens.len(), - chunk_size = self.prefill_chunk_size, + chunk_size = selected_chunk, cached = seq.already_cached_tokens, start = seq.prefill_start_offset, boundary = plan.boundary(), diff --git a/src/server/batch/scheduler/speculative_finalize.rs b/src/server/batch/scheduler/speculative_finalize.rs index 0b1b83159..ee07bfa0d 100644 --- a/src/server/batch/scheduler/speculative_finalize.rs +++ b/src/server/batch/scheduler/speculative_finalize.rs @@ -225,7 +225,7 @@ impl BatchScheduler { tokenizer: &self.tokenizer, drafter_slot: &mut self.speculative_drafter_slot, dispatch: &self.speculative_dispatch, - prefill_chunk_size: self.prefill_chunk_size, + prefill_chunk_size: self.prefill_chunk_policy.alone, // Classic-step probes are a B=1 profiling concern (#736). profile_probe_rounds: 0, prefill_boundary: None, @@ -340,7 +340,7 @@ impl BatchScheduler { tokenizer: &self.tokenizer, drafter_slot: &mut self.speculative_drafter_slot, dispatch: &self.speculative_dispatch, - prefill_chunk_size: self.prefill_chunk_size, + prefill_chunk_size: self.prefill_chunk_policy.alone, // While the adaptive policy is profiling this pairing, ask // the burst for a few classic-step probe rounds so the // measured-cost estimator has a classic step time (#736); diff --git a/src/server/batch/scheduler_prompt_cache_plan_tests.rs b/src/server/batch/scheduler_prompt_cache_plan_tests.rs index 626da5bab..73635d8cc 100644 --- a/src/server/batch/scheduler_prompt_cache_plan_tests.rs +++ b/src/server/batch/scheduler_prompt_cache_plan_tests.rs @@ -24,9 +24,12 @@ //! of that follow-up, and a hit from inside a piece is reported as such by the //! plan. `docs/CONTINUOUS_BATCHING.md` records the real-checkpoint numbers. -use super::scheduler_model_owned_cache_tests::{options, prompt, test_store, tiny_gemma3}; +use super::scheduler_model_owned_cache_tests::{ + options, prompt, test_store, tiny_gemma3, tiny_gemma3_args, tiny_gemma3_weights, +}; use super::*; use crate::LoadedModel; +use crate::models::{Gemma3Model, Gemma3Wrapper}; use crate::server::config::{DecodeStorageBackend, PreemptionPolicy, PromptCacheRequestContext}; use crate::server::model_provider::{GenerateEvent, GenerationResult}; use crate::server::prompt_cache::{PromptCacheStore, key::MultimodalDigest}; @@ -41,9 +44,17 @@ use std::time::Duration; /// with an explicit prefill chunk, so the plan's chunk pieces run across /// ticks. fn plan_scheduler(store: Arc, chunk: usize) -> BatchScheduler { + plan_scheduler_with_model(tiny_gemma3(), store, chunk) +} + +fn plan_scheduler_with_model( + model: Gemma3Wrapper, + store: Arc, + chunk: usize, +) -> BatchScheduler { let (_tx, rx) = mpsc::channel(); let sched = BatchScheduler::with_config( - LoadedModel::Gemma3(tiny_gemma3()), + LoadedModel::Gemma3(model), MlxcelTokenizer::stub(), vec![7], rx, @@ -62,6 +73,123 @@ fn plan_scheduler(store: Arc, chunk: usize) -> BatchScheduler sched } +/// Scheduler with the production adaptive chunk policy. +fn adaptive_plan_scheduler() -> BatchScheduler { + let mut args = tiny_gemma3_args(); + args.max_position_embeddings = 8192; + let model = Gemma3Wrapper::new( + Gemma3Model::from_weights(&tiny_gemma3_weights(), &args).expect("tiny gemma3 loads"), + ); + plan_scheduler_with_model( + model, + test_store(), + mlxcel_core::prefill_plan::DEFAULT_PREFILL_CHUNK, + ) + .with_prefill_chunk_policy(mlxcel_core::prefill_plan::PrefillChunkPolicy::resolve(None)) +} + +/// Admit a synthetic request and return its queued sequence and response +/// stream for execution tests. +fn queued_prompt( + sched: &mut BatchScheduler, + prompt_len: usize, +) -> (SequenceInfo, mpsc::Receiver) { + let mut opts = options(); + opts.max_tokens = 1; + opts.ignore_eos = true; + let tokens: Vec = (0..prompt_len).map(|index| (index % 6) as i32).collect(); + let (tx, rx) = mpsc::channel(); + sched.enqueue_request( + "prompt".to_string(), + Some(tokens), + opts, + Vec::new(), + Vec::new(), + Vec::new(), + tx, + Arc::new(AtomicBool::new(false)), + true, + ); + ( + sched.prefill_queue.dequeue().expect("request is queued"), + rx, + ) +} + +/// Put a dequeued request back at the head and execute its first real prefill +/// forward through the scheduler and tiny Gemma test model. +fn start_recorded_prefill(sched: &mut BatchScheduler, seq: SequenceInfo) { + let seq_id = seq.seq_id; + sched + .prefill_queue + .enqueue_front(seq) + .unwrap_or_else(|_| panic!("re-queue adaptive prefill request")); + sched.execute_prefill(seq_id); +} + +/// Execute every parked continuation, then require a successful Done frame. +fn finish_recorded_prefill(sched: &mut BatchScheduler, rx: &mpsc::Receiver) { + while sched.chunked_prefill_seq.is_some() { + assert!( + sched.continue_chunked_prefill(), + "every parked continuation runs a model forward" + ); + } + loop { + match rx.recv_timeout(Duration::from_secs(5)) { + Ok(GenerateEvent::Done(_)) => break, + Ok(GenerateEvent::Error(err)) => panic!("adaptive prefill aborted: {err}"), + Ok(_) => {} + Err(err) => panic!("adaptive prefill did not complete: {err}"), + } + } +} + +#[test] +fn adaptive_chunk_policy_executes_and_resumes_real_prefill_forwards() { + // A live decode row selects 512 for every executed piece, including the + // one-token terminal continuation. + { + let mut sched = adaptive_plan_scheduler(); + let (live, _live_rx) = queued_prompt(&mut sched, 8); + sched.active_batch.add(live).expect("decode row fits"); + let (contended, rx) = queued_prompt(&mut sched, 1537); + start_recorded_prefill(&mut sched, contended); + finish_recorded_prefill(&mut sched, &rx); + assert_eq!(sched.prefill_forward_lengths, vec![512, 512, 512, 1]); + } + + // With no other row decoding, the same execution path selects 2048. + { + let mut sched = adaptive_plan_scheduler(); + let (alone, rx) = queued_prompt(&mut sched, 3000); + start_recorded_prefill(&mut sched, alone); + finish_recorded_prefill(&mut sched, &rx); + assert_eq!(sched.prefill_forward_lengths, vec![2048, 952]); + } + + // The decode row drains after the first forward. The parked request is + // repartitioned at its live 512-token cursor and completes without the old + // "no remaining tokens" abort. + { + let mut sched = adaptive_plan_scheduler(); + let (live, _live_rx) = queued_prompt(&mut sched, 8); + let live_id = live.seq_id; + sched.active_batch.add(live).expect("decode row fits"); + let (changing, rx) = queued_prompt(&mut sched, 5000); + start_recorded_prefill(&mut sched, changing); + assert_eq!(sched.prefill_forward_lengths, vec![512]); + assert!(sched.chunked_prefill_seq.is_some()); + + sched + .active_batch + .remove(live_id) + .expect("decode row exists"); + finish_recorded_prefill(&mut sched, &rx); + assert_eq!(sched.prefill_forward_lengths, vec![512, 2048, 2048, 392]); + } +} + /// A chat turn's cache context: `history` is the rendering without the /// generation prompt, which is where the plan splits the prefill. fn turn_ctx(history: Option<&[i32]>) -> PromptCacheRequestContext { diff --git a/src/server/cli_input.rs b/src/server/cli_input.rs index 3458a2573..52f570ad3 100644 --- a/src/server/cli_input.rs +++ b/src/server/cli_input.rs @@ -155,15 +155,16 @@ pub struct ServerStartupInput { pub reranker_model_path: Option, /// `--rerank-batch-size`: query/document pairs per rerank forward pass. pub rerank_batch_size: usize, - pub prefill_chunk_size: usize, + /// Explicit `--prefill-chunk-size`, or `None` to use the adaptive policy. + pub prefill_chunk_size: Option, /// #1011: `--prefill-grant-interval`, the decode ticks a parked chunked /// prefill yields before the scheduler grants it one. `None` = env override /// / shipped default; `Some(0)` = grant disabled (pre-#1011 starvation). pub prefill_grant_interval: Option, /// llama-server alias for `--prefill-chunk-size` (`--batch-size` / `-b`). /// - /// When set, maps to `prefill_chunk_size`. If both this and `prefill_chunk_size` - /// differ from the default, `prefill_chunk_size` takes precedence with a warning. + /// When set, maps to `prefill_chunk_size`. If both are present and differ, + /// `prefill_chunk_size` takes precedence with a warning. pub batch_size: Option, /// llama-server `--ubatch-size`. Accepted but ignored on Apple Silicon. pub ubatch_size: Option, @@ -812,6 +813,8 @@ impl ServerStartupInput { .resolve_kv_unified(self.parallel_auto, self.max_batch_size.is_some()); let resolution = resolve_prefill_chunk_size(self.prefill_chunk_size, self.batch_size, self.ubatch_size); + let prefill_chunk_policy = + mlxcel_core::prefill_plan::PrefillChunkPolicy::resolve(resolution.explicit); // resolve the server-wide thinking budget once, up-front. // Invalid values are logged and treated as unbounded so the server // still starts (per-request errors are surfaced as 400s at the route). @@ -1064,7 +1067,8 @@ impl ServerStartupInput { embedding_request_timeout_secs: self.embedding_request_timeout_secs, reranker_model_path: self.reranker_model_path, rerank_batch_size: self.rerank_batch_size, - prefill_chunk_size: resolution.prefill_chunk_size, + prefill_chunk_size: prefill_chunk_policy.alone, + prefill_chunk_size_contended: prefill_chunk_policy.contended, prefill_grant_interval: self.prefill_grant_interval, batch_size_conflict: resolution.batch_size_conflict, ubatch_size_provided: resolution.ubatch_size_provided, @@ -2261,8 +2265,9 @@ fn apply_optional_usize_env_fallback(value: &mut Option, key: &str, flag_ /// Result of resolving the prefill chunk size from the explicit flag and llama-server aliases. pub struct PrefillChunkResolution { - /// The effective prefill chunk size to use. - pub prefill_chunk_size: usize, + /// Explicit chunk override from the native flag, llama alias, or env. + /// `None` selects the adaptive server defaults. + pub explicit: Option, /// True when `--ubatch-size` was provided (always ignored; caller should log a notice). pub ubatch_size_provided: bool, /// True when both `--batch-size` and an explicit `--prefill-chunk-size` were supplied @@ -2275,36 +2280,24 @@ pub struct PrefillChunkResolution { /// Resolution rules: /// - `--ubatch-size` is always ignored on Apple Silicon unified memory (logged at info level). /// - `--batch-size` is an alias for `--prefill-chunk-size`. If both are provided with -/// different non-default values, `--prefill-chunk-size` takes precedence with a warning. +/// different values, `--prefill-chunk-size` takes precedence with a warning. +/// - `MLXCEL_PREFILL_CHUNK` is the final fallback and pins both scheduler states. pub fn resolve_prefill_chunk_size( - prefill_chunk_size: usize, + prefill_chunk_size: Option, batch_size: Option, ubatch_size: Option, ) -> PrefillChunkResolution { - // One chunk policy for the CLI and the server (ADR 0007, #2170). - let default_prefill_chunk_size = mlxcel_core::prefill_plan::prefill_chunk_len(); - let ubatch_size_provided = ubatch_size.is_some(); - - match batch_size { - None => PrefillChunkResolution { - prefill_chunk_size, - ubatch_size_provided, - batch_size_conflict: false, - }, - Some(bs) => { - let explicit_prefill = prefill_chunk_size != default_prefill_chunk_size; - let conflict = explicit_prefill && bs != prefill_chunk_size; - PrefillChunkResolution { - prefill_chunk_size: if explicit_prefill { - prefill_chunk_size - } else { - bs - }, - ubatch_size_provided, - batch_size_conflict: conflict, - } - } + let batch_size_conflict = prefill_chunk_size + .zip(batch_size) + .is_some_and(|(prefill, batch)| prefill != batch); + let explicit = prefill_chunk_size + .or(batch_size) + .or_else(mlxcel_core::prefill_plan::prefill_chunk_override); + PrefillChunkResolution { + explicit, + ubatch_size_provided, + batch_size_conflict, } } diff --git a/src/server/cli_input_tests.rs b/src/server/cli_input_tests.rs index 318fa3576..d16ff5550 100644 --- a/src/server/cli_input_tests.rs +++ b/src/server/cli_input_tests.rs @@ -87,7 +87,7 @@ fn sample_input() -> ServerStartupInput { embedding_request_timeout_secs: 120, reranker_model_path: None, rerank_batch_size: 0, - prefill_chunk_size: mlxcel_core::prefill_plan::prefill_chunk_len(), + prefill_chunk_size: None, prefill_grant_interval: None, batch_size: None, ubatch_size: None, @@ -447,39 +447,40 @@ fn into_startup_config_propagates_image_limits() { #[test] fn resolve_prefill_chunk_size_batch_size_alias_takes_effect() { - let default = mlxcel_core::prefill_plan::prefill_chunk_len(); - let r = resolve_prefill_chunk_size(default, Some(1024), None); - assert_eq!(r.prefill_chunk_size, 1024); + let r = resolve_prefill_chunk_size(None, Some(1024), None); + assert_eq!(r.explicit, Some(1024)); assert!(!r.batch_size_conflict); assert!(!r.ubatch_size_provided); } #[test] fn resolve_prefill_chunk_size_explicit_prefill_wins_with_conflict() { - let r = resolve_prefill_chunk_size(256, Some(1024), None); - assert_eq!(r.prefill_chunk_size, 256); + let r = resolve_prefill_chunk_size(Some(2048), Some(1024), None); + assert_eq!(r.explicit, Some(2048)); assert!(r.batch_size_conflict); } #[test] fn resolve_prefill_chunk_size_no_batch_size_returns_prefill() { - let r = resolve_prefill_chunk_size(768, None, None); - assert_eq!(r.prefill_chunk_size, 768); + let r = resolve_prefill_chunk_size(Some(768), None, None); + assert_eq!(r.explicit, Some(768)); assert!(!r.batch_size_conflict); } #[test] fn resolve_prefill_chunk_size_ubatch_sets_provided_flag() { - let default = mlxcel_core::prefill_plan::prefill_chunk_len(); - let r = resolve_prefill_chunk_size(default, None, Some(256)); + let r = resolve_prefill_chunk_size(None, None, Some(256)); assert!(r.ubatch_size_provided); - assert_eq!(r.prefill_chunk_size, default); + assert_eq!( + r.explicit, + mlxcel_core::prefill_plan::prefill_chunk_override() + ); } #[test] fn resolve_prefill_chunk_size_both_same_value_no_conflict() { - let r = resolve_prefill_chunk_size(1024, Some(1024), None); - assert_eq!(r.prefill_chunk_size, 1024); + let r = resolve_prefill_chunk_size(Some(1024), Some(1024), None); + assert_eq!(r.explicit, Some(1024)); assert!(!r.batch_size_conflict); } @@ -503,6 +504,7 @@ fn into_startup_config_resolves_batch_size_alias() { input.batch_size = Some(1024); let startup = input.into_startup_config().expect("valid startup input"); assert_eq!(startup.prefill_chunk_size, 1024); + assert_eq!(startup.prefill_chunk_size_contended, 1024); assert!(!startup.batch_size_conflict); assert!(!startup.ubatch_size_provided); } @@ -510,15 +512,87 @@ fn into_startup_config_resolves_batch_size_alias() { #[test] fn into_startup_config_detects_batch_size_conflict() { let mut input = sample_input(); - input.prefill_chunk_size = 256; + input.prefill_chunk_size = Some(2048); input.batch_size = Some(1024); input.ubatch_size = Some(64); let startup = input.into_startup_config().expect("valid startup input"); - assert_eq!(startup.prefill_chunk_size, 256); + assert_eq!(startup.prefill_chunk_size, 2048); + assert_eq!(startup.prefill_chunk_size_contended, 2048); assert!(startup.batch_size_conflict); assert!(startup.ubatch_size_provided); } +#[test] +fn into_startup_config_uses_adaptive_defaults_without_an_override() { + const CHILD: &str = "MLXCEL_CLI_ADAPTIVE_PREFILL_TEST_CHILD"; + if std::env::var_os(CHILD).is_some() { + let startup = sample_input() + .into_startup_config() + .expect("valid startup input"); + assert_eq!(startup.prefill_chunk_size, 2048); + assert_eq!(startup.prefill_chunk_size_contended, 512); + return; + } + + let output = std::process::Command::new(std::env::current_exe().expect("current test binary")) + .args([ + "--exact", + "server::cli_input::tests::into_startup_config_uses_adaptive_defaults_without_an_override", + "--test-threads=1", + ]) + .env(CHILD, "1") + .env_remove("MLXCEL_PREFILL_CHUNK") + .env_remove("LLAMA_ARG_BATCH") + .output() + .expect("spawn isolated adaptive prefill test"); + assert!( + output.status.success(), + "isolated test failed:\nstdout:\n{}\nstderr:\n{}", + String::from_utf8_lossy(&output.stdout), + String::from_utf8_lossy(&output.stderr) + ); + assert!( + String::from_utf8_lossy(&output.stdout).contains("1 passed"), + "isolated test did not run the exact case:\n{}", + String::from_utf8_lossy(&output.stdout) + ); +} + +#[test] +fn into_startup_config_uses_env_prefill_override_in_a_fresh_process() { + const CHILD: &str = "MLXCEL_CLI_PREFILL_CHUNK_TEST_CHILD"; + if std::env::var_os(CHILD).is_some() { + let startup = sample_input() + .into_startup_config() + .expect("valid startup input"); + assert_eq!(startup.prefill_chunk_size, 768); + assert_eq!(startup.prefill_chunk_size_contended, 768); + return; + } + + let output = std::process::Command::new(std::env::current_exe().expect("current test binary")) + .args([ + "--exact", + "server::cli_input::tests::into_startup_config_uses_env_prefill_override_in_a_fresh_process", + "--test-threads=1", + ]) + .env(CHILD, "1") + .env("MLXCEL_PREFILL_CHUNK", "768") + .output() + .expect("spawn isolated CLI prefill override test"); + assert!( + output.status.success(), + "isolated test failed:\nstdout:\n{}\nstderr:\n{}", + String::from_utf8_lossy(&output.stdout), + String::from_utf8_lossy(&output.stderr) + ); + assert!( + String::from_utf8_lossy(&output.stdout).contains("1 passed"), + "isolated test did not run the exact case:\n{}", + String::from_utf8_lossy(&output.stdout) + ); +} + #[test] fn into_startup_config_propagates_pp_layers() { let mut input = sample_input(); diff --git a/src/server/config.rs b/src/server/config.rs index 29e9fa0de..2ae5927f8 100644 --- a/src/server/config.rs +++ b/src/server/config.rs @@ -834,9 +834,11 @@ pub struct ServerConfig { /// takes the loaded reranker kind's own default. See /// [`DEFAULT_RERANK_BATCH_SIZE`]. pub rerank_batch_size: usize, - /// Number of tokens per prefill chunk. When 0, chunking is disabled and - /// the full prompt is prefilled in a single pass. + /// Number of tokens per prefill chunk when no other sequence is decoding. + /// When 0, chunking is disabled and the full prompt is one pass. pub prefill_chunk_size: usize, + /// Number of tokens per prefill chunk while another sequence is decoding. + pub prefill_chunk_size_contended: usize, /// #1011 prefill fairness interval (`--prefill-grant-interval`): decode /// ticks a parked chunked prefill yields before the scheduler grants it /// one, bounding the admitted request's time to first token. `None` lets @@ -1086,6 +1088,9 @@ pub struct ServerConfig { impl Default for ServerConfig { fn default() -> Self { + let prefill_chunk_policy = mlxcel_core::prefill_plan::PrefillChunkPolicy::resolve( + mlxcel_core::prefill_plan::prefill_chunk_override(), + ); Self { gcp: None, reasoning_format: crate::server::ReasoningFormat::default(), @@ -1174,7 +1179,8 @@ impl Default for ServerConfig { embedding_request_timeout_secs: DEFAULT_EMBEDDING_REQUEST_TIMEOUT_SECS, reranker_model_path: None, rerank_batch_size: DEFAULT_RERANK_BATCH_SIZE, - prefill_chunk_size: mlxcel_core::prefill_plan::prefill_chunk_len(), + prefill_chunk_size: prefill_chunk_policy.alone, + prefill_chunk_size_contended: prefill_chunk_policy.contended, // #1011: unset -> scheduler resolves the env override / default. prefill_grant_interval: None, enable_preemption: false, diff --git a/src/server/engine_probe/server_engine.rs b/src/server/engine_probe/server_engine.rs index 153bcdf80..f2c4657e2 100644 --- a/src/server/engine_probe/server_engine.rs +++ b/src/server/engine_probe/server_engine.rs @@ -58,8 +58,8 @@ pub struct ServerEngineOptions { /// records the fallback, which is what lets the probe report the storage /// it actually measured. pub decode_storage: DecodeStorageBackend, - /// `--prefill-chunk-size`; `None` keeps the server default (the shared - /// chunk policy, 2048 unless `MLXCEL_PREFILL_CHUNK` says otherwise). + /// `--prefill-chunk-size`; `None` keeps the adaptive server default (2048 + /// alone and 512 with live decode, unless `MLXCEL_PREFILL_CHUNK` pins it). pub prefill_chunk_size: Option, /// Whether the cross-request prompt cache is enabled (`--no-prompt-cache` /// when `false`). The server default is enabled. @@ -187,6 +187,7 @@ impl ServerEngine { }; if let Some(chunk) = options.prefill_chunk_size { startup.prefill_chunk_size = chunk; + startup.prefill_chunk_size_contended = chunk; } if !options.prompt_cache { startup.prompt_cache.enabled = false; diff --git a/src/server/model_provider.rs b/src/server/model_provider.rs index ae2f190b1..7a9f09aad 100644 --- a/src/server/model_provider.rs +++ b/src/server/model_provider.rs @@ -742,7 +742,10 @@ impl ModelProvider { config.lora_runtime.clone(), config.max_batch_size, config.max_queue_depth, - config.prefill_chunk_size, + mlxcel_core::prefill_plan::PrefillChunkPolicy { + alone: config.prefill_chunk_size, + contended: config.prefill_chunk_size_contended, + }, config.enable_preemption, config.preemption_policy, config.max_batch_prefill, @@ -1137,7 +1140,7 @@ impl ModelProvider { None, max_batch_size, max_queue_depth, - prefill_chunk_size, + mlxcel_core::prefill_plan::PrefillChunkPolicy::resolve(Some(prefill_chunk_size)), enable_preemption, preemption_policy, max_batch_prefill, @@ -1199,7 +1202,7 @@ impl ModelProvider { lora_runtime: Option>, max_batch_size: usize, max_queue_depth: usize, - prefill_chunk_size: usize, + prefill_chunk_policy: mlxcel_core::prefill_plan::PrefillChunkPolicy, enable_preemption: bool, preemption_policy: crate::server::config::PreemptionPolicy, max_batch_prefill: usize, @@ -1262,7 +1265,7 @@ impl ModelProvider { lora_runtime, max_batch_size, max_queue_depth, - prefill_chunk_size, + prefill_chunk_policy, enable_preemption, preemption_policy, max_batch_prefill: max_batch_prefill.max(1), @@ -1391,7 +1394,7 @@ impl ModelProvider { let sched_config = model_worker::WorkerSchedulerConfig { max_batch_size, max_queue_depth, - prefill_chunk_size: 0, + prefill_chunk_policy: mlxcel_core::prefill_plan::PrefillChunkPolicy::resolve(Some(0)), // #1011: chunking is off on this path, so no prefill can ever park // and the fairness grant is unreachable; keep the default. lora_adapters: Vec::new(), diff --git a/src/server/model_worker.rs b/src/server/model_worker.rs index 47184a65d..d875f6b50 100644 --- a/src/server/model_worker.rs +++ b/src/server/model_worker.rs @@ -59,7 +59,7 @@ pub(crate) struct WorkerSchedulerConfig { pub lora_runtime: Option>, pub max_batch_size: usize, pub max_queue_depth: usize, - pub prefill_chunk_size: usize, + pub prefill_chunk_policy: mlxcel_core::prefill_plan::PrefillChunkPolicy, /// #1011: explicit `--prefill-grant-interval` value bounding how long a /// parked chunked prefill yields to a live decode batch. `None` keeps the /// `MLXCEL_PREFILL_GRANT_INTERVAL` override or the shipped default; @@ -518,8 +518,12 @@ pub(crate) fn spawn_model_worker_with_batch_config( } } - let chunk_info = if sched_config.prefill_chunk_size > 0 { - format!(", prefill_chunk_size={}", sched_config.prefill_chunk_size) + let chunk_info = if sched_config.prefill_chunk_policy.alone > 0 { + format!( + ", prefill_chunk_size={}, prefill_chunk_size_contended={}", + sched_config.prefill_chunk_policy.alone, + sched_config.prefill_chunk_policy.contended + ) } else { String::new() }; @@ -727,12 +731,13 @@ pub(crate) fn spawn_model_worker_with_batch_config( sched_config.max_queue_depth, batch_metrics.clone(), batch_observability.clone(), - sched_config.prefill_chunk_size, + sched_config.prefill_chunk_policy.alone, sched_config.enable_preemption, sched_config.preemption_policy, sched_config.max_batch_prefill, sched_config.decode_storage_backend, ) + .with_prefill_chunk_policy(sched_config.prefill_chunk_policy) .with_vision_cache_size(sched_config.vision_cache_size) // per-batch runtime-LoRA scale application (#1439). .with_lora_runtime(sched_config.lora_runtime.clone()) @@ -1328,6 +1333,9 @@ pub(crate) fn spawn_legacy_model_worker( 1, // max_batch_prefill = 1 → sequential prefill crate::server::DecodeStorageBackend::Dense, ) + .with_prefill_chunk_policy(mlxcel_core::prefill_plan::PrefillChunkPolicy::resolve( + Some(0), + )) .with_xtc_newline_token_ids(xtc_newline_token_ids) .with_reasoning_budget(reasoning_budget, thinking_ids); scheduler.serve(); diff --git a/src/server/routes/props_tests.rs b/src/server/routes/props_tests.rs index f6030f8dc..517bda652 100644 --- a/src/server/routes/props_tests.rs +++ b/src/server/routes/props_tests.rs @@ -227,18 +227,16 @@ fn geometry_block_reports_batch_and_kv_bounds() { /// startup notice's trigger) and never changes the number (#1472). #[test] fn geometry_block_reports_the_resolved_batch_size_alias() { - // The default is the shared chunk policy (ADR 0007), not a literal: 512 - // is an explicit chunk now and would win over the alias. - let default_chunk = mlxcel_core::prefill_plan::prefill_chunk_len(); - let alias = if default_chunk == 1024 { 2048 } else { 1024 }; + let alias = 1024; let resolved = - crate::server::cli_input::resolve_prefill_chunk_size(default_chunk, Some(alias), Some(256)); - assert_eq!(resolved.prefill_chunk_size, alias); + crate::server::cli_input::resolve_prefill_chunk_size(None, Some(alias), Some(256)); + assert_eq!(resolved.explicit, Some(alias)); assert!(resolved.ubatch_size_provided); assert!(!resolved.batch_size_conflict); let block = geometry_block(&ServerConfig { - prefill_chunk_size: resolved.prefill_chunk_size, + prefill_chunk_size: alias, + prefill_chunk_size_contended: alias, max_batch_size: 4, max_kv_size: Some(4096), ..Default::default() diff --git a/src/server/runtime_settings.rs b/src/server/runtime_settings.rs index 80c0540a6..9ef9b2009 100644 --- a/src/server/runtime_settings.rs +++ b/src/server/runtime_settings.rs @@ -121,6 +121,7 @@ pub const CLASSIFIED_SERVER_CONFIG_FIELDS: &[&str] = &[ "reranker_model_path", "rerank_batch_size", "prefill_chunk_size", + "prefill_chunk_size_contended", "prefill_grant_interval", "enable_preemption", "preemption_policy", @@ -494,6 +495,7 @@ fn read_only_reason(field: &str) -> &'static str { | "embedding_request_timeout_secs" | "rerank_batch_size" | "prefill_chunk_size" + | "prefill_chunk_size_contended" | "prefill_grant_interval" | "enable_preemption" | "preemption_policy" @@ -636,6 +638,7 @@ fn read_only_value(config: &ServerConfig, field: &str) -> Value { .unwrap_or(Value::Null), "rerank_batch_size" => json!(config.rerank_batch_size), "prefill_chunk_size" => json!(config.prefill_chunk_size), + "prefill_chunk_size_contended" => json!(config.prefill_chunk_size_contended), "prefill_grant_interval" => json!(config.prefill_grant_interval), "enable_preemption" => json!(config.enable_preemption), "preemption_policy" => debug(&config.preemption_policy), @@ -716,6 +719,7 @@ fn read_only_kind(field: &str) -> KnobKind { | "embedding_request_timeout_secs" | "rerank_batch_size" | "prefill_chunk_size" + | "prefill_chunk_size_contended" | "max_batch_prefill" | "vision_cache_size" | "video_max_frames" => KnobKind::Int, diff --git a/src/server/runtime_settings_tests.rs b/src/server/runtime_settings_tests.rs index 5c16585fa..24273b23b 100644 --- a/src/server/runtime_settings_tests.rs +++ b/src/server/runtime_settings_tests.rs @@ -59,9 +59,9 @@ fn rejection_reason<'a>(result: &'a ApplyResult, name: &str) -> &'a str { } #[test] -fn server_config_schema_classifies_all_112_fields() { +fn server_config_schema_classifies_all_113_fields() { let declared = declared_server_config_fields(); - assert_eq!(declared.len(), 112, "ServerConfig field count changed"); + assert_eq!(declared.len(), 113, "ServerConfig field count changed"); assert_eq!( declared.as_slice(), CLASSIFIED_SERVER_CONFIG_FIELDS, @@ -69,7 +69,7 @@ fn server_config_schema_classifies_all_112_fields() { ); let specs = schema(&ServerConfig::default()); - assert_eq!(specs.len(), 112); + assert_eq!(specs.len(), 113); let expected_names: Vec<_> = CLASSIFIED_SERVER_CONFIG_FIELDS .iter() .copied() @@ -79,7 +79,7 @@ fn server_config_schema_classifies_all_112_fields() { assert_eq!(actual_names, expected_names); assert_eq!( actual_names.iter().copied().collect::>().len(), - 112, + 113, "every management API name must be unique" ); diff --git a/src/server/startup.rs b/src/server/startup.rs index 0c9333ef1..8572b33a6 100644 --- a/src/server/startup.rs +++ b/src/server/startup.rs @@ -276,8 +276,10 @@ pub struct ServerStartupConfig { /// `--rerank-batch-size`. Forwarded to /// [`super::config::ServerConfig::rerank_batch_size`]. pub rerank_batch_size: usize, - /// Prefill chunk size in tokens (0 = disabled). + /// Prefill chunk size when no other sequence is decoding (0 = disabled). pub prefill_chunk_size: usize, + /// Prefill chunk size while at least one other sequence is decoding. + pub prefill_chunk_size_contended: usize, /// Set when `--batch-size` and `--prefill-chunk-size` conflict; triggers a startup warning. pub batch_size_conflict: bool, /// Set when `--ubatch-size` was provided; triggers a startup info notice. @@ -637,6 +639,9 @@ pub struct ServerStartupConfig { impl Default for ServerStartupConfig { fn default() -> Self { + let prefill_chunk_policy = mlxcel_core::prefill_plan::PrefillChunkPolicy::resolve( + mlxcel_core::prefill_plan::prefill_chunk_override(), + ); Self { reasoning_format: crate::server::ReasoningFormat::default(), reasoning_alias_field: crate::server::ReasoningAliasField::default(), @@ -699,7 +704,8 @@ impl Default for ServerStartupConfig { crate::server::config::DEFAULT_EMBEDDING_REQUEST_TIMEOUT_SECS, reranker_model_path: None, rerank_batch_size: crate::server::config::DEFAULT_RERANK_BATCH_SIZE, - prefill_chunk_size: mlxcel_core::prefill_plan::prefill_chunk_len(), + prefill_chunk_size: prefill_chunk_policy.alone, + prefill_chunk_size_contended: prefill_chunk_policy.contended, batch_size_conflict: false, ubatch_size_provided: false, enable_preemption: false, @@ -1757,6 +1763,7 @@ pub(super) fn build_server_config( reranker_model_path: startup.reranker_model_path.clone(), rerank_batch_size: startup.rerank_batch_size, prefill_chunk_size: startup.prefill_chunk_size, + prefill_chunk_size_contended: startup.prefill_chunk_size_contended, // #1011: pass the explicit --prefill-grant-interval through untouched // (the scheduler resolves env / shipped default when this is None). prefill_grant_interval: startup.prefill_grant_interval, @@ -3081,6 +3088,7 @@ pub async fn start_server(mut startup: ServerStartupConfig) -> Result<()> { kv_unified = startup.kv_unified, n_parallel = startup.n_parallel, prefill_chunk_size = startup.prefill_chunk_size, + prefill_chunk_size_contended = startup.prefill_chunk_size_contended, max_kv_size = ?startup.max_kv_size, "resolved context and batch geometry (0 = the checkpoint's own trained context)" ); diff --git a/tests/vlm_wrapper_capability_delegation.rs b/tests/vlm_wrapper_capability_delegation.rs index 005b80ed9..2165aed46 100644 --- a/tests/vlm_wrapper_capability_delegation.rs +++ b/tests/vlm_wrapper_capability_delegation.rs @@ -143,6 +143,84 @@ fn field_types(src: &str) -> BTreeSet { out } +/// The brace-counted body of the first function whose header contains +/// `signature`. +fn function_body<'a>(src: &'a str, signature: &str) -> Option<&'a str> { + let header = src.find(signature)?; + let open = header + src[header..].find('{')?; + let mut depth = 0usize; + for (i, c) in src[open..].char_indices() { + match c { + '{' => depth += 1, + '}' => { + depth -= 1; + if depth == 0 { + return Some(&src[open + 1..open + i]); + } + } + _ => {} + } + } + None +} + +/// Locals that visibly derive from the padded-prefill model predicate. +/// +/// The prefill planner centralizes the predicate in `PrefillCaps::can_pad`. +/// Recognize `caps.can_pad()` only after inspecting that helper's body; a +/// same-named helper without `supports_padded_prefill` provides no evidence. +fn padded_prefill_guard_aliases(src: &str) -> Vec { + let shared_caps_guarded = function_body(src, "fn can_pad(&self) -> bool") + .is_some_and(|body| body.contains("supports_padded_prefill")); + let mut aliases = Vec::new(); + for line in src.lines() { + let Some(rest) = line.trim().strip_prefix("let ") else { + continue; + }; + if !line.contains("supports_padded_prefill") + && !(shared_caps_guarded && line.contains("caps.can_pad()")) + { + continue; + } + let name: String = rest + .chars() + .take_while(|c| c.is_alphanumeric() || *c == '_') + .collect(); + if !name.is_empty() { + aliases.push(name); + } + } + aliases +} + +fn unguarded_tile_aligned_prefills(src: &str) -> Vec<(usize, String)> { + const WINDOW: usize = 12; + let lines: Vec<&str> = src.lines().collect(); + let aliases = padded_prefill_guard_aliases(src); + let mut unguarded = Vec::new(); + for (i, line) in lines.iter().enumerate() { + let trimmed = line.trim_start(); + // The definition, its doc examples and its own unit tests are not + // prefill sites and have nothing to guard. + if !line.contains("align_to_na_tile(") + || trimmed.starts_with("//") + || trimmed.starts_with("fn ") + || trimmed.starts_with("pub fn ") + || trimmed.starts_with("assert") + { + continue; + } + let start = i.saturating_sub(WINDOW); + if lines[start..=i].iter().any(|l| { + l.contains("supports_padded_prefill") || aliases.iter().any(|a| l.contains(a.as_str())) + }) { + continue; + } + unguarded.push((i + 1, line.trim().to_string())); + } + unguarded +} + #[test] fn a_vision_wrapper_delegates_padded_prefill_when_its_backbone_refuses_it() { let mut files = Vec::new(); @@ -213,7 +291,6 @@ fn a_vision_wrapper_delegates_padded_prefill_when_its_backbone_refuses_it() { /// that is one a reader cannot see either. #[test] fn every_tile_aligned_prefill_is_guarded_on_the_model_predicate() { - const WINDOW: usize = 12; let mut files = Vec::new(); rust_sources(&repo_root().join("src"), &mut files); files.sort(); @@ -223,49 +300,12 @@ fn every_tile_aligned_prefill_is_guarded_on_the_model_predicate() { let Ok(src) = fs::read_to_string(file) else { continue; }; - let lines: Vec<&str> = src.lines().collect(); - // Locals bound to the predicate count as the guard, so a site that - // hoists it out of a loop is not reported as unguarded. - let mut aliases: Vec = Vec::new(); - for line in &lines { - let Some(rest) = line.trim().strip_prefix("let ") else { - continue; - }; - if !line.contains("supports_padded_prefill") { - continue; - } - let name: String = rest - .chars() - .take_while(|c| c.is_alphanumeric() || *c == '_') - .collect(); - if !name.is_empty() { - aliases.push(name); - } - } - for (i, line) in lines.iter().enumerate() { - let trimmed = line.trim_start(); - // The definition, its doc examples and its own unit tests are not - // prefill sites and have nothing to guard. - if !line.contains("align_to_na_tile(") - || trimmed.starts_with("//") - || trimmed.starts_with("fn ") - || trimmed.starts_with("pub fn ") - || trimmed.starts_with("assert") - { - continue; - } - let start = i.saturating_sub(WINDOW); - if lines[start..=i].iter().any(|l| { - l.contains("supports_padded_prefill") - || aliases.iter().any(|a| l.contains(a.as_str())) - }) { - continue; - } + for (line_number, line) in unguarded_tile_aligned_prefills(&src) { unguarded.push(format!( "{}:{}: {}", file.strip_prefix(repo_root()).unwrap_or(file).display(), - i + 1, - line.trim() + line_number, + line )); } } @@ -279,3 +319,42 @@ fn every_tile_aligned_prefill_is_guarded_on_the_model_predicate() { unguarded.join("\n ") ); } + +#[test] +fn shared_can_pad_alias_requires_the_model_predicate() { + let guarded = r#" +impl PrefillCaps { + fn can_pad(&self) -> bool { + self.align_prefill && self.supports_padded_prefill + } +} + +fn padded_len(caps: PrefillCaps, len: usize) -> usize { + let can_pad = caps.can_pad(); + if can_pad { align_to_na_tile(len) } else { len } +} +"#; + assert_eq!( + padded_prefill_guard_aliases(guarded), + vec!["can_pad".to_string()] + ); + assert!(unguarded_tile_aligned_prefills(guarded).is_empty()); + + let mutated = guarded.replace( + "self.align_prefill && self.supports_padded_prefill", + "self.align_prefill", + ); + assert!(padded_prefill_guard_aliases(&mutated).is_empty()); + assert_eq!(unguarded_tile_aligned_prefills(&mutated).len(), 1); + + let direct = r#" +fn padded_len(model: &Model, len: usize) -> usize { + if model.supports_padded_prefill() { + align_to_na_tile(len) + } else { + len + } +} +"#; + assert!(unguarded_tile_aligned_prefills(direct).is_empty()); +}