Spec: ML-DSA's ExpandMask four polynomials at a time (vg_mldsa_expand_mask_poly4) - #508
Merged
Merged
Conversation
The contract of a function that samples four elements of the matrix A at once, as ML-KEM's vg_mlkem_sample_ntt4 does: for each of the four 34-byte seeds, RejNTTPoly (FIPS 204 Algorithm 30) of it, reduced, or 0 if the loop does not finish within Appendix C's least bound for one of them. Its scratch is that of vg_mldsa_rej_ntt_poly (256 u64s), so callers can pass the same working space. No implementation yet: an x86-64 one, with four SHAKE128 instances in AVX2 registers, follows in its own PR. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
Four interleaved Keccak states, the second buffer of the 4-way permutation and its table of round constants already take 2368 bytes, before the squeezed output; ML-KEM's vg_mlkem_sample_ntt4 has 1024 u64s for the same. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
…_mask_poly4) vg_mldsa_expand_mask_poly4(seeds, gamma1, a, scratch) writes the four polynomials of ExpandMask of the four 66-byte seeds at seeds (seed66) to the four polynomials from a (poly4): for each, what expandMaskContract says. The four are independent, so an implementation may run four SHAKE256 instances at once, as vg_mldsa_rej_ntt_poly4 does SHAKE128. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
Base automatically changed from
claude/fervent-einstein-ukl7t7-rej4spec
to
main
October 1, 2026 22:01
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Trust change (Spec only). This PR is stacked on #492, whose
poly4it reuses. Its base is #492's branch, so the diff shows only this change. It will be retargeted tomainonce #492 merges.Why
With the AVX2 polynomial arithmetic (#491) and
ExpandAfour entries at a time (on top of #492),ExpandMaskis the largest remaining Keccak cost in ML-DSA signing on x86-64. On ML-DSA-65 it is about 16–19% of the instructions per signature. Each iteration of the signing loop callsvg_mldsa_expand_mask_polyℓ times, on independent seedsρ″ ‖ IntegerToBytes(κ + r, 2). Those calls suit four SHAKE256 instances at once in AVX2 registers, asvg_mldsa_rej_ntt_poly4runs four SHAKE128 instances.What
All of this is in
Spec/MlDsa/Poly.lean. It mirrorsrejNTT4Sig/rejNTT4Contract/rejNTT4Apifrom #492, and reusesexpandMaskContract's postcondition for each polynomial:expandMask4Sig:vg_mldsa_expand_mask_poly4(seeds: *const [u8; 264], gamma1: u32, a: *mut [u32; 1024], scratch: *mut [u64; 1024]). The 8 KiB of scratch is the same asvg_mldsa_rej_ntt_poly4's: four interleaved states, the permutation's second buffer and round-constant table, and 4 × 640 bytes of squeezed output.seed66: seedkis the 66 bytes from byte66 k. Polynomialkstarts atpoly4 a k(byte1024 k), as in Spec: ML-DSA's RejNTTPoly four times (vg_mldsa_rej_ntt_poly4) #492.expandMask4Contract:gamma1is 2¹⁷ or 2¹⁹, as forexpandMaskContract.k < 4, the polynomial atpoly4 a kistoRq (BitUnpack(H(seed66 seeds k, 32c), γ₁ − 1, γ₁))and is reduced.gamma1. This is the same asexpandMaskContract, since the seeds hold the secretρ″.expandMask4Api: the name, module, contract on every target, and documentation. It reusesctDocandscratchSafety.There are no other changes: no TCB, Impl or artifacts. The implementation and its proofs come in a later PR:
vg_mldsa_expand_mask_poly;Moving signing onto it comes in that PR too.
Checks run
lake build +VerifiedGarbage.Spec.MlDsa.Polycheck_lean_imports,check_lean_speed🤖 Generated with Claude Code
https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
Generated by Claude Code