Vectorize x86-64 RC2 key expansion and mashing with SSE2 - #462
Merged
Merged
Conversation
reaperhulk
force-pushed
the
codex/rc2-x86-64-sse2-scan
branch
from
October 1, 2026 13:37
e017429 to
bfd21c4
Compare
* Vectorize x86-64 RC2 mashing schedule scans with SSE2 * Use a fresh rustup home for ARM benchmark containers --------- Co-authored-by: Paul Kehrer <161495+reaperhulk@users.noreply.github.com>
reaperhulk
enabled auto-merge
October 1, 2026 13:56
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
RC2-CBC previously scanned each PITABLE byte during key expansion and each schedule word during mashing with scalar instructions. This changes both x86-64 paths to baseline SSE2 scans that select eight candidates per vector. The existing public initialization and block/CBC paths use the optimized code directly.
Each 16-bit lane forms an equality mask with XOR, subtraction, and arithmetic shift, then masks and ORs the candidate value. A horizontal OR collects the selected value. Extraction temporarily uses 16 bytes at offset 64 in the existing scratch buffer, then restores those bytes exactly. The implementation is proven against the existing merged key-expansion, block, and CBC contracts. There are no Spec/TCB, public API, scratch-size, or CPU-dispatch changes; all assembly is emitted from Lean.
OpenSSL key expansion, OpenSSL rounds, and AWS-LC RC2 use secret-indexed table loads. Studying their setup and bulk paths helped isolate both costs. The vector scans preserve fixed memory addresses and control flow.
Local measurements use the normal public RC2-CBC API on a Xeon E5-2696 v4 (Broadwell), core 2,
VG_CPU_FEATURES=none, with two alternating baseline/candidate pairs, 30 samples, 0.5-second warmup and 1-second measurement. Times include initialization, update, finalization, and wiping.Key-expansion-only comparison against merged scalar #443:
Additional bulk gain, comparing key-expansion-only SSE2 against the final key-expansion plus mashing SSE2 implementation from #465:
Locks excluded team builds and other benchmarks. An unrelated Lean build outside those locks was present during the bulk comparison; absolute times are higher, but both alternating comparisons agree closely. The stages were measured separately: their absolute times should not be combined. Local results do not establish CI performance.
Validation for both stages: full Lean builds (5,537 and 5,542 jobs), emitter generation/
--checkincluding standard-axiom and compiler-override audits; all required Python checks and algorithm-table regeneration; library and benchmark fmt/clippy, includingcpu-features-env; complete Wycheproof tests normally and withVG_CPU_FEATURES=none. After the CI-only rebase, the RC2 artifact target passed again. Generated changes are confined tosrc/asm/x86_64/rc2.rs.Benchmark dependency metadata includes
rc2alongsiderc2_cbcso subsequent RC2-only changes select these benchmarks. The ARM benchmark container uses a freshRUSTUP_HOME, matching #458 and #460, to repair its stale-image toolchain installation failure.The final CI benchmark run passed all four x86-64 feature configurations. In the baseline-feature comparison, 16 KiB encryption fell from 1.04 ms to 550.88 µs (47.2% lower time), and decryption from 905.14 µs to 312.12 µs (65.5% lower). The resulting gaps to OpenSSL were 1.7× and 2.3×, respectively. All AArch64 comparisons also passed; unrelated ARM/x86 full-suite jobs exceeded their 15-minute limits with no architecture-code changes.