Skip to content

ML-DSA on x86-64: sample verification's  as one run of entries (verify −8% with AVX2, −9% baseline ML-DSA-87) - #544

Merged
reaperhulk merged 9 commits into
claude/fervent-einstein-ukl7t7-em4from
claude/fervent-einstein-ukl7t7-rej4v
Oct 2, 2026
Merged

reaperhulk merged 9 commits into
claude/fervent-einstein-ukl7t7-em4from
claude/fervent-einstein-ukl7t7-rej4v

Conversation

@alex

@alex alex commented Oct 2, 2026

Copy link
Copy Markdown
Member

Stacked on #537; the diff shows only this change. This is the follow-up promised in #534. It removes the duplicated entry that made the baseline ML-DSA-87 verification 7% slower there.

What

Verification kept Â[r, s] at polynomial 20 + 8r + s, so a group of four entries for vg_mldsa_rej_ntt_poly4 could not cross a row:

  • for ℓ = 7, the second group of each row sampled entry 3 again;
  • for ℓ = 5, the last entry of each row was sampled alone.

Â[r, s] is now entry e = ℓr + s, at polynomial 20 + e (pA l r s). Verification samples the kℓ entries as key generation and signing do:

  • kℓ/4 groups of four (aGrp), where each seed's index bytes are e mod ℓ and e / ℓ;
  • then the last kℓ mod 4 entries one at a time (aOne).

As a result:

four-way calls single calls
ML-DSA-44 4 0 (as before)
ML-DSA-65 7 2 (was 6 and 6)
ML-DSA-87 14 0 (was 16 four-way calls with 8 duplicates)

The working space of vg_mldsa_rej_ntt_poly4 stays at polynomial 20 + 8k, which is after  for every ℓ, so the scratch layout is otherwise unchanged.

Proofs (untrusted)

  • StageA and CTSample now count sampled entries by number: Done ℓ e r' c' := ℓr' + c' < e, with rc_eq recovering (r', c') from e. aGrp_ok/aGrp_tr cover entries 4g … 4g + 3, and aOne_ok/aOne_tr cover entry e.
  • StageC, Correct and CTCompute only pass ℓ to pA.
  • Instrs follows the new samples.

There are no TCB or Spec changes.

Performance

Instructions per verification (callgrind: 20 verifications, minus key generation and one signature; base is #537's head):

AVX2 baseline (VG_CPU_FEATURES=none)
ML-DSA-44 unchanged unchanged
ML-DSA-65 1.061M → 0.977M (−7.9%) unchanged
ML-DSA-87 1.677M → 1.540M (−8.1%) 3.276M → 2.971M (−9.3%)

Checks run

  • lake build and the emitter
  • check_lean_imports, check_lean_speed, check_vectors, check_arch_gates, check_variants, check_mcdt
  • cargo fmt --check and cargo clippy --all-targets -D warnings
  • cargo test --release with Wycheproof, both on the default path and with VG_CPU_FEATURES=none --features cpu-features-env

🤖 Generated with Claude Code

https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr


Generated by Claude Code

claude added 2 commits October 2, 2026 00:16
Verification kept Â[r, s] at polynomial 20 + 8r + s, so a group of four
entries of vg_mldsa_rej_ntt_poly4 could not cross a row: for ℓ = 7 the
second group of each row sampled entry 3 again, and for ℓ = 5 the last
entry of each row was sampled alone. Â[r, s] is now entry e = ℓr + s at
polynomial 20 + e, and verification samples the kℓ entries as key
generation and signing do: kℓ/4 groups of four (each seed's indices are
e mod ℓ and e / ℓ), then the last kℓ mod 4 one at a time. ML-DSA-87
(56 entries) has no duplicate, and ML-DSA-65 makes 7 four-way calls and
2 single ones instead of 6 and 6.

The proofs of the samplers (StageA, CTSample) now count sampled entries
by their number, Done ℓ e r' c' := ℓr' + c' < e; the rest only pass ℓ
to pA. The working space of vg_mldsa_rej_ntt_poly4 stays at polynomial
20 + 8k, after  for every ℓ.

Instructions per verification (callgrind, 20 verifications less key
generation and one signature): AVX2 ML-DSA-65 1.061M → 0.977M (−7.9%),
ML-DSA-87 1.677M → 1.540M (−8.1%); baseline ML-DSA-87 3.276M → 2.971M
(−9.3%, undoing the duplicate's cost); ML-DSA-44 and the ML-DSA-65
baseline are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
…-einstein-ukl7t7-rej4v

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
…-einstein-ukl7t7-rej4v

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
claude added 6 commits October 2, 2026 03:31
…-einstein-ukl7t7-rej4v

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
…-einstein-ukl7t7-rej4v

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
…-einstein-ukl7t7-rej4v

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
…-einstein-ukl7t7-rej4v

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
…-einstein-ukl7t7-rej4v

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
…-einstein-ukl7t7-rej4v

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
@reaperhulk
reaperhulk merged commit 551c0ef into claude/fervent-einstein-ukl7t7-em4 Oct 2, 2026
58 checks passed
@reaperhulk
reaperhulk deleted the claude/fervent-einstein-ukl7t7-rej4v branch October 2, 2026 10:22
alex pushed a commit that referenced this pull request Oct 2, 2026
…e merge of #526

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Ddof3szoTi7HB8iCsM2MCr
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Development

Successfully merging this pull request may close these issues.

3 participants