Repository navigation
update(perf): adapt server prefill chunks to decode load - #2268
Merged
Merged
Conversation
Choose 2048-token prefill chunks when the scheduler has no active decode rows and 512-token chunks while streams are decoding, preserving the existing direct-engine and speculative alone value. Explicit native, llama-compatible, and environment overrides pin both states. Add cursor-resumable prefill plans so a parked request can safely change chunk size between ticks, propagate both values through startup and worker configuration, and document the measured TTFT/ITL tradeoff. Validate core planning, CLI precedence, runtime schema, props compatibility, and real scheduler forward sequences under CUDA. Refs #2228
Use the Option question mark operator when the prefill chunk environment value is absent. This preserves the existing fallback behavior and satisfies the workspace Clippy warning policy. Refs #2228
Compute the next prefill piece directly when paged-block capacity checks only need its tile padding. The accessor shares chunk eligibility, boundary, range, and padding helpers with full plan construction, so repeated decode and prompt-lookup checks stay constant-space. Cover equivalence with resumable plans across initial and continuation cursors, chunk values, history boundaries, padding modes, model opt-outs, and embedding inputs. Refs #2228
Annotate the first next-piece equivalence case so saturating cursor arithmetic resolves to usize under Clippy's all-targets build. Refs #2228
Teach the tile-alignment source audit to recognize a PrefillCaps can_pad alias only when its brace-counted helper body still reads supports_padded_prefill. Keep the original direct-guard recognition and add a mutation fixture proving removal of the model predicate is reported. Refs #2228
Verify directly that the guarded shared capability helper produces the can_pad alias and that removing the model predicate removes it. This prevents the positive fixture from passing through only the nearby direct-guard window. Refs #2228
Document the passing GB10 admission and idle-server TTFT gates with exact build provenance and policy decisions. Archive the raw records, server logs, reproducible driver and analyzer, and finalize the bilingual technical report. Refs #2228
Name the bilingual PR reports with the required PR, slug, and date convention. Refs #2228
Explain that the archived build commit requires the unsquashed PR lineage, add a safe fresh-build command for squash-merged or newer revisions, and make the replay driver fail with an actionable ancestry error. Validation: - git diff --check - bash -n docs/benchmark_results/data/adaptive-prefill-chunk-gb10-2026-10-11/run.sh Refs #2228
Replace the obsolete sibling-failure caveat with the complete workspace, contract, and CUDA validation results after both preceding PRs merged. Mark the bilingual technical reports complete while preserving the original measurement provenance and fresh-build replay instructions. Refs #2228
inureyes
force-pushed
the
update/issue-2228-adaptive-prefill-chunk
branch
from
October 11, 2026 08:16
808bccc to
c379623
Compare
inureyes
marked this pull request as ready for review
October 11, 2026 08:16
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Choose the server prefill chunk from live scheduler state: 2048 tokens while no other sequence is decoding and 512 while a decode batch is active. Rebuild parked continuations from their live cursor so requests can change chunk size safely, while explicit native, llama-compatible, and environment settings continue to pin both states.
The change also adds an allocation-free next-piece calculation for repeated paged-block capacity checks, propagates the two-value policy through both server front ends, and documents the prompt-cache partition implications and measured tradeoff.
Validation
[512, 512, 512, 1],[2048, 952], and[512, 2048, 2048, 392], and required request completion.cargo check --profile test-fast --features cuda --lib, formatting, and diff checks passed.make verify-test-cudapassed with zero failed CUDA summaries, including the BF16 and F32 SSM update parity tests.Closes #2228