Skip to content

Vectorize x86-64 RC2 key expansion and mashing with SSE2 - #462

Merged
reaperhulk merged 2 commits into
mainfrom
codex/rc2-x86-64-sse2-scan
Oct 1, 2026
Merged

reaperhulk merged 2 commits into
mainfrom
codex/rc2-x86-64-sse2-scan

Conversation

@reaperhulk

@reaperhulk reaperhulk commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

RC2-CBC previously scanned each PITABLE byte during key expansion and each schedule word during mashing with scalar instructions. This changes both x86-64 paths to baseline SSE2 scans that select eight candidates per vector. The existing public initialization and block/CBC paths use the optimized code directly.

Each 16-bit lane forms an equality mask with XOR, subtraction, and arithmetic shift, then masks and ORs the candidate value. A horizontal OR collects the selected value. Extraction temporarily uses 16 bytes at offset 64 in the existing scratch buffer, then restores those bytes exactly. The implementation is proven against the existing merged key-expansion, block, and CBC contracts. There are no Spec/TCB, public API, scratch-size, or CPU-dispatch changes; all assembly is emitted from Lean.

OpenSSL key expansion, OpenSSL rounds, and AWS-LC RC2 use secret-indexed table loads. Studying their setup and bulk paths helped isolate both costs. The vector scans preserve fixed memory addresses and control flow.

Local measurements use the normal public RC2-CBC API on a Xeon E5-2696 v4 (Broadwell), core 2, VG_CPU_FEATURES=none, with two alternating baseline/candidate pairs, 30 samples, 0.5-second warmup and 1-second measurement. Times include initialization, update, finalization, and wiping.

Key-expansion-only comparison against merged scalar #443:

Operation Size Scalar runs 1 / 2 SSE2 key expansion runs 1 / 2 Lower time
Encrypt 64 B 31.15 / 30.44 µs 14.08 / 13.57 µs 54.8–55.4%
Decrypt 64 B 30.97 / 30.26 µs 13.08 / 12.97 µs 57.1–57.8%
Encrypt 1 KiB 96.70 / 97.01 µs 78.95 / 77.75 µs 18.3–19.9%
Decrypt 1 KiB 87.32 / 85.75 µs 68.53 / 69.74 µs 18.7–21.5%

Additional bulk gain, comparing key-expansion-only SSE2 against the final key-expansion plus mashing SSE2 implementation from #465:

Operation Size Baseline runs 1 / 2 Final runs 1 / 2 Lower time
Encrypt 64 B 17.464 / 17.611 µs 14.517 / 14.539 µs 16.9–17.4%
Decrypt 64 B 16.831 / 16.957 µs 13.792 / 13.771 µs 18.1–18.8%
Encrypt 1 KiB 98.663 / 99.854 µs 53.210 / 53.541 µs 46.1–46.4%
Decrypt 1 KiB 90.261 / 89.286 µs 40.269 / 40.074 µs 55.1–55.4%
Encrypt 16 KiB 1402.6 / 1408.4 µs 678.47 / 673.04 µs 51.6–52.2%
Decrypt 16 KiB 1259.3 / 1250.4 µs 463.17 / 464.23 µs 62.9–63.2%

Locks excluded team builds and other benchmarks. An unrelated Lean build outside those locks was present during the bulk comparison; absolute times are higher, but both alternating comparisons agree closely. The stages were measured separately: their absolute times should not be combined. Local results do not establish CI performance.

Validation for both stages: full Lean builds (5,537 and 5,542 jobs), emitter generation/--check including standard-axiom and compiler-override audits; all required Python checks and algorithm-table regeneration; library and benchmark fmt/clippy, including cpu-features-env; complete Wycheproof tests normally and with VG_CPU_FEATURES=none. After the CI-only rebase, the RC2 artifact target passed again. Generated changes are confined to src/asm/x86_64/rc2.rs.

Benchmark dependency metadata includes rc2 alongside rc2_cbc so subsequent RC2-only changes select these benchmarks. The ARM benchmark container uses a fresh RUSTUP_HOME, matching #458 and #460, to repair its stale-image toolchain installation failure.

The final CI benchmark run passed all four x86-64 feature configurations. In the baseline-feature comparison, 16 KiB encryption fell from 1.04 ms to 550.88 µs (47.2% lower time), and decryption from 905.14 µs to 312.12 µs (65.5% lower). The resulting gaps to OpenSSL were 1.7× and 2.3×, respectively. All AArch64 comparisons also passed; unrelated ARM/x86 full-suite jobs exceeded their 15-minute limits with no architecture-code changes.

@reaperhulk
reaperhulk force-pushed the codex/rc2-x86-64-sse2-scan branch from e017429 to bfd21c4 Compare October 1, 2026 13:37
@reaperhulk
reaperhulk added this pull request to the merge queue Oct 1, 2026
@reaperhulk
reaperhulk removed this pull request from the merge queue due to a manual request Oct 1, 2026
* Vectorize x86-64 RC2 mashing schedule scans with SSE2

* Use a fresh rustup home for ARM benchmark containers

---------

Co-authored-by: Paul Kehrer <161495+reaperhulk@users.noreply.github.com>
@reaperhulk
reaperhulk enabled auto-merge October 1, 2026 13:56
@reaperhulk
reaperhulk added this pull request to the merge queue Oct 1, 2026
@reaperhulk reaperhulk changed the title Vectorize x86-64 RC2 key-expansion scans with SSE2 Vectorize x86-64 RC2 key expansion and mashing with SSE2 Oct 1, 2026
Merged via the queue into main with commit ff98988 Oct 1, 2026
38 of 40 checks passed
@reaperhulk
reaperhulk deleted the codex/rc2-x86-64-sse2-scan branch October 1, 2026 14:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

1 participant