Skip to content

fix(perf): avoid replaying terminal B1 decode - #2271

Merged
inureyes merged 4 commits into
mainfrom
update/issue-2227-b1-decode
Oct 11, 2026
Merged

inureyes merged 4 commits into
mainfrom
update/issue-2227-b1-decode

Conversation

@inureyes

@inureyes inureyes commented Oct 11, 2026 •

Copy link
Copy Markdown
Member

For a fused-lookahead-eligible capped B=1 request, the scheduler submitted the next lookahead at the length boundary, discarded both speculative appends, and synchronously replayed a token it had already sampled. This changes the terminal B=1 path to commit that pending token through the existing finish/finalization logic, removing two complete decode forwards per affected request while leaving B>1 behavior unchanged. Closes #2227.

Profiling on GB10 showed row plumbing (0.0053%), metrics (0.0462%), provider channel send (0.0698%), and lookahead collect (0.0836%) below the issue's 0.1% threshold, so those candidates remain unchanged. A guard mutation makes the live-lookahead test fail with four forwards instead of two, confirming that the new branch accounts for the removed work.

Performance, five interleaved quiet-host rounds against pre-epic 4c44e317:

Model Prompt Server dense vs baseline CLI Baseline null arm
Qwen3-1.7B 4-bit 256 +0.57% (-0.25%..+1.54%) +0.04% (-0.41%..+0.76%)
Qwen3-1.7B 4-bit 8192 +0.32% (+0.04%..+0.63%) +0.22% (-0.13%..+0.66%)
Llama-3.2-1B 4-bit 256 +0.48% (+0.19%..+1.41%) +0.07% (-12.38%..+0.54%)
Llama-3.2-1B 4-bit 8192 -0.16% (-1.15%..+0.05%) -0.17% (-0.29%..+0.14%)

All four medians pass the -1.0% threshold. Every candidate range overlaps its null range, so the measurements establish the threshold without claiming a speedup. The unexplained Llama 256 null outlier (-12.38% in round 3) remains in the raw data and analysis without exclusion or rerun.

Validation:

  • cargo fmt --all -- --check
  • cargo clippy --workspace --all-targets --features cuda -- -D warnings
  • restored full CUDA suite: 8,930 root tests and 2,072 core tests passed; all integration binaries had zero failures
  • repository contract checks
  • four live-lookahead terminal scheduler tests, including max-token boundaries, stochastic sampling, EOS precedence, and unchanged B=2 behavior
  • controlled guard mutation: expected assertion failure, actual 4 forwards versus expected 2
  • real mlxcel run before/after generated text: byte-identical
  • Llama full engine parity: 13 identical, zero divergent, one not applicable
  • Qwen no-prompt-cache engine parity: nine identical, zero divergent, one not applicable

The full Qwen parity command still exits 1 on its prompt-cache miss/hit pair at greedy token 58 (374 versus 594) and seeded token 61 (576 versus 362). Forced synchronous decode reproduces the same positions and token IDs, while the new guard is unreachable; the no-cache pairs affected by this change are identical. The benchmark report records this existing prompt-cache limitation, which is outside this issue's scope.

Evidence: benchmark report and raw data.

@inureyes inureyes added status:review Under review type:performance Performance improvements priority:high High priority area:inference Generation, sampling, decoding (incl. speculative, DRY) area:cli Command-line interface / CLI flags labels Oct 11, 2026
Add the bilingual technical report and record the verified performance, correctness, provenance, and Qwen prompt-cache limitation for PR #2271. Remove generated Nsight capture binaries from git while retaining their hashes, sizes, local provenance, and committed text exports under the repository's binary-assets policy.

Validation:
- `sha256sum -c SHA256SUMS`
- `python3 scripts/ci/check_binary_assets.py`
- `bash scripts/ci/check_binary_assets_test.sh`
- archived acceptance analyzer: 4/4 cells passed

Refs #2227
@inureyes inureyes added status:done Completed and removed status:review Under review labels Oct 11, 2026
@inureyes
inureyes merged commit 3ae626c into main Oct 11, 2026
27 checks passed
@inureyes
inureyes deleted the update/issue-2227-b1-decode branch October 11, 2026 11:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:cli Command-line interface / CLI flags area:inference Generation, sampling, decoding (incl. speculative, DRY) priority:high High priority status:done Completed type:performance Performance improvements

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf: bring mlxcel run and chat B=1 decode within the 1.0 percent threshold

1 participant