Repository navigation
update(perf): bound DRY and sparse penalty step costs - #2272
Merged
Merged
Conversation
Replace quadratic backward matching with a linear reverse-Z pass and scatter DRY and rebuild frequency/presence updates only across touched token IDs. Preserve backend-specific dense dtype behavior and the existing incremental f32 path. Add dense-reference parity matrices, breaker-boundary and linear comparison-count checks, bf16 incremental regression coverage, and an ignored timing report. Build and tests have not run yet because the orchestrated CUDA profiling window is still active. Refs #2259
Build a reversed-token matcher once per DRY step so exact IDs and multi-token breakers cost expected O(P + N) instead of multiplying breaker configuration size by the history window. Preserve scalar breaker semantics in direct per-position parity tests and cover large exact sets plus shared-tail failure chains with separate build and scan counters. Validation: cargo fmt --all. Build and tests are pending the root validation gate. Refs #2259
Publish the quiet GB10 host matching and lazy graph-construction timings for issue #2259 with byte-preserved logs, validation output, checksums, and source, artifact, driver, and host-gate provenance. The record distinguishes these measurements from GPU end-to-end latency. Validation: sha256sum --check SHA256SUMS and git diff --cached --check. Refs #2259
Add bilingual technical reports and preserve the original two-commit implementation patch used for GB10 timing. Record the measured and rebased source identities, prove the scoped core, native, and lockfile diff is empty, and document replay from the public measurement base. Validation: - `sha256sum --check SHA256SUMS` - Replayed the archived patch from `73700aaf` and matched tree `58b9a04161eaf88048699d0930a69dd59ccdd7b3` - `git diff --cached --check` Refs #2259, #2272
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A full-history DRY step previously rescanned repeated suffixes quadratically, and checking every position against request-supplied breaker lists could multiply window length by breaker configuration size. This change computes suffix matches with reverse Z and breaker starts with a reversed-token Aho-Corasick matcher, bounding expected host work to
O(N + P)for history lengthNand configured breaker-pattern tokensP.DRY and rebuilt frequency/presence penalties now scatter only touched token IDs instead of constructing vocabulary-sized host buffers. The sparse paths preserve the pinned MLX backend semantics: CUDA bf16 mixed arithmetic remains bf16 on rebuild paths, while the pre-existing incremental frequency/presence path still explicitly promotes to f32.
Quiet GB10 host matching and lazy MLX graph-construction timing used vocabulary 4, a period-3 history, and no breakers. Reference iterations were 3/3/1 and new-path iterations were 20/20/20. Returned arrays were not evaluated, so these figures are not GPU end-to-end latency:
Validation on CUDA GB10:
cargo fmt --checkpassed.Timing evidence:
docs/benchmark_results/dry-penalty-step-cost-gb10-2026-10-11.mdand its linked raw package.Closes #2259
Refs #2229, #2256