Vectorize x86-64 RC2 mashing schedule scans with SSE2 - #465
Merged
reaperhulk merged 2 commits intoOct 1, 2026
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
RC2-CBC block processing currently scans the 64-word key schedule one scalar candidate at a time during each mashing lookup. This replaces those scans with baseline SSE2 operations that select eight schedule words per vector, while retaining the fixed traversal of every candidate.
The existing RC2 block and CBC contracts, public API, dispatch, and scratch sizes are unchanged. Each lookup temporarily uses a fixed 16-byte slot in the existing scratch buffer and restores its original contents. The proof establishes the exact schedule selection and unchanged memory, then composes it into the existing block and CBC verification. No
Spec/orTCB/changes.This follows the key-expansion vector scan in #462. OpenSSL and AWS-LC instead perform secret-indexed schedule loads in their mashing rounds; those accesses cannot be copied here. Primary upstream references: OpenSSL RC2 rounds, AWS-LC RC2.
Two alternating baseline/candidate pairs on an Intel Xeon E5-2696 v4 (Broadwell), pinned to core 2, using the normal public API with
VG_CPU_FEATURES=none, 30 samples, 0.5-second warmup and 1-second measurement per workload. Baseline is #462. Times include initialization, update, finalization, and wiping.Team builds and other benchmarks were excluded by locks. An unrelated Lean build outside those locks was present on the host; absolute times are higher than earlier samples, but the two alternating comparisons agree closely. These local results do not establish CI performance.
Validation: full
lake build(5,542 jobs), emitter generation and--checkincluding artifact audits; all required Python checks; Cargo formatting and Clippy for library/tests and benchmarks; complete Wycheproof tests both normally and withVG_CPU_FEATURES=nonethroughcpu-features-env. The rebase onto #462's current CI-fixed head leaves all RC2 implementation/proof/generated code unchanged and is checked with the RC2 artifact target. No resource limits or proof shortcuts were added.Benchmark dependency metadata now includes the generated
rc2module alongsiderc2_cbc, so RC2-only changes select these benchmarks.The ARM benchmark container also uses a fresh
RUSTUP_HOME, matching the merged CI fix in #458. Its first run failed while rustup removed a stale image component; this one-line repair is shared with #460 and changes no cryptographic code.