Skip to content

update(perf): adapt server prefill chunks to decode load - #2268

Merged
inureyes merged 10 commits into
mainfrom
update/issue-2228-adaptive-prefill-chunk
Oct 11, 2026
Merged

inureyes merged 10 commits into
mainfrom
update/issue-2228-adaptive-prefill-chunk

Conversation

@inureyes

@inureyes inureyes commented Oct 11, 2026 •

Copy link
Copy Markdown
Member

Choose the server prefill chunk from live scheduler state: 2048 tokens while no other sequence is decoding and 512 while a decode batch is active. Rebuild parked continuations from their live cursor so requests can change chunk size safely, while explicit native, llama-compatible, and environment settings continue to pin both states.

The change also adds an allocation-free next-piece calculation for repeated paged-block capacity checks, propagates the two-value policy through both server front ends, and documents the prompt-cache partition implications and measured tradeoff.

Validation

  • CUDA core prefill-plan tests: 17 passed.
  • CUDA server CLI tests: 148 passed; runtime-settings tests: 8 passed; props compatibility: passed.
  • CUDA block-reclaim tests: 20 passed.
  • Actual scheduler regression passed exact forward sequences [512, 512, 512, 1], [2048, 952], and [512, 2048, 2048, 392], and required request completion.
  • Shared-capability source audit: 3 passed, including a negative mutation fixture.
  • Workspace all-targets Clippy, cargo check --profile test-fast --features cuda --lib, formatting, and diff checks passed.
  • Fresh-release admission p95: 54.1, 54.8, and 55.2 ms; all below 62.0 ms.
  • Five-round idle-server TTFT passed on Qwen3-1.7B and Llama-3.2-1B: each adaptive-policy median stayed inside its explicit-2048 null range.
  • After PR fix: clean up paged decode and SSM precision #2270 and PR fix: contain decode readback failures and preserve sampler state #2269 merged, workspace test compilation, all-targets Clippy, formatting, all three contract suites, and make verify-test-cuda passed with zero failed CUDA summaries, including the BF16 and F32 SSM update parity tests.

Closes #2228

@inureyes inureyes added status:in-progress Currently being worked on type:performance Performance improvements priority:medium Medium priority area:core mlxcel-core: MLX FFI, primitives, KV cache, layers area:inference Generation, sampling, decoding (incl. speculative, DRY) labels Oct 11, 2026
Choose 2048-token prefill chunks when the scheduler has no active decode rows and 512-token chunks while streams are decoding, preserving the existing direct-engine and speculative alone value. Explicit native, llama-compatible, and environment overrides pin both states.

Add cursor-resumable prefill plans so a parked request can safely change chunk size between ticks, propagate both values through startup and worker configuration, and document the measured TTFT/ITL tradeoff.

Validate core planning, CLI precedence, runtime schema, props compatibility, and real scheduler forward sequences under CUDA.

Refs #2228
Use the Option question mark operator when the prefill chunk environment value is absent. This preserves the existing fallback behavior and satisfies the workspace Clippy warning policy.

Refs #2228
Compute the next prefill piece directly when paged-block capacity checks only need its tile padding. The accessor shares chunk eligibility, boundary, range, and padding helpers with full plan construction, so repeated decode and prompt-lookup checks stay constant-space.

Cover equivalence with resumable plans across initial and continuation cursors, chunk values, history boundaries, padding modes, model opt-outs, and embedding inputs.

Refs #2228
Annotate the first next-piece equivalence case so saturating cursor arithmetic resolves to usize under Clippy's all-targets build.

Refs #2228
Teach the tile-alignment source audit to recognize a PrefillCaps can_pad alias only when its brace-counted helper body still reads supports_padded_prefill. Keep the original direct-guard recognition and add a mutation fixture proving removal of the model predicate is reported.

Refs #2228
Verify directly that the guarded shared capability helper produces the can_pad alias and that removing the model predicate removes it. This prevents the positive fixture from passing through only the nearby direct-guard window.

Refs #2228
Document the passing GB10 admission and idle-server TTFT gates with exact build provenance and policy decisions.

Archive the raw records, server logs, reproducible driver and analyzer, and finalize the bilingual technical report.

Refs #2228
Name the bilingual PR reports with the required PR, slug, and date convention.

Refs #2228
Explain that the archived build commit requires the unsquashed PR lineage, add a safe fresh-build command for squash-merged or newer revisions, and make the replay driver fail with an actionable ancestry error.

Validation:
- git diff --check
- bash -n docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/run.sh

Refs #2228
Replace the obsolete sibling-failure caveat with the complete workspace, contract, and CUDA validation results after both preceding PRs merged.

Mark the bilingual technical reports complete while preserving the original measurement provenance and fresh-build replay instructions.

Refs #2228
@inureyes
inureyes force-pushed the update/issue-2228-adaptive-prefill-chunk branch from 808bccc to c379623 Compare October 11, 2026 08:16
@inureyes
inureyes marked this pull request as ready for review October 11, 2026 08:16
@inureyes inureyes added status:review Under review status:done Completed and removed status:in-progress Currently being worked on status:review Under review labels Oct 11, 2026
@inureyes
inureyes merged commit 73700aa into main Oct 11, 2026
52 of 53 checks passed
@inureyes
inureyes deleted the update/issue-2228-adaptive-prefill-chunk branch October 11, 2026 08:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:core mlxcel-core: MLX FFI, primitives, KV cache, layers area:inference Generation, sampling, decoding (incl. speculative, DRY) priority:medium Medium priority status:done Completed type:performance Performance improvements

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf: choose the prefill chunk by whether other sequences are decoding

1 participant