Skip to content

[TEST] x86_64: Use Plantard reduction - #1871

Draft
mkannwischer wants to merge 2 commits into
mainfrom
plantard-reduce
Draft

[TEST] x86_64: Use Plantard reduction#1871
mkannwischer wants to merge 2 commits into
mainfrom
plantard-reduce

Conversation

@mkannwischer

@mkannwischer mkannwischer commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

Bo-Yin Yang suggested a five-instruction AVX2 sequence that multiplies a
coefficient by a compile-time constant modulo q and returns the canonical
representative, z = (a*w) mod q with 0 <= z < q:

vpmulhw  t, a, L    ; bits[31:16] of a*L        (signed)
vpmullw  u, a, H    ; bits[15:0]  of a*H
vpaddw   m, t, u    ; bits[31:16] of (a*W mod 2^32)
vpaddw   n, m, S    ; + bias s, wraps mod 2^16
vpmulhuw z, n, Q    ; floor(n*q / 2^16), n read UNSIGNED

The constants come from w: b = (-2^32 * w) mod q, W = (b * q^-1) mod 2^32
split as H*2^16 + L, bias s = 2.

The only place I was able to make use of it so far is mlk_reduce_avx2_asm,
i.e., multiplying by w = 1, replacing Barrett reduction plus conditional
subtraction: 5 instructions per vector instead of 8.

Performance

machine poly_reduce polyvec_reduce (k=3)
AMD EPYC 4th gen (c7a) 26 → 19 (-27%) 78 → 59 (-24%)
Intel Xeon 4th gen (c7i) 32 → 19 (-41%) 96 → 56 (-42%)
AMD EPYC 3rd gen (c6a) 29 → 21 (-28%) 88 → 65 (-26%)
Intel Xeon 3rd gen (c6i) 181 → 115 (-37%) 532 → 337 (-37%)

Unfortunately, we don't spend much time in reduction overall, so scheme benchmarks don't really reflect it.

Verification

The new sequence is proven against the same HOL-Light specification as before: each output coefficient is
ival(x) rem 3329.

However, we also prove the modmul sequence more generally, as
PLANTARD_MULCONST in proofs/hol_light/x86_64/proofs/mlkem_reduce_avx2_asm.ml:

let PLANTARD_MULCONST = prove
 (`!(h:int16) (l:int16) (qw:int16) (x:int16) b w.
      &0 < (&(val qw):int) /\ (&(val qw):int) <= &26214 /\
      abs b <= &(val qw) /\
      &(val qw) * (ival h * &65536 + ival l) = b + &4294967296 * w
      ==> ival(plantard_seq (h,l,word 2,qw) x) =
          (ival x * w) rem &(val qw)`,

plantard_seq is the five instructions on one 16-bit lane.

Beyond the proof: the routine was checked exhaustively against the reference
over all 2^16 inputs.

Signed-off-by: Matthias J. Kannwischer <matthias@zerorisc.com>
@mkannwischer mkannwischer added the benchmark this PR should be benchmarked in CI label Aug 19, 2026

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mac Mini (M1, 2020) benchmarks

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 12322 cycles 12322 cycles 1
ML-KEM-512 encaps 14757 cycles 14758 cycles 1.00
ML-KEM-512 decaps 19315 cycles 19315 cycles 1
ML-KEM-768 keypair 21179 cycles 21179 cycles 1
ML-KEM-768 encaps 23515 cycles 23513 cycles 1.00
ML-KEM-768 decaps 30058 cycles 30054 cycles 1.00
ML-KEM-1024 keypair 30331 cycles 30331 cycles 1
ML-KEM-1024 encaps 34109 cycles 34112 cycles 1.00
ML-KEM-1024 decaps 43668 cycles 43668 cycles 1

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ppc64le (POWER10) benchmarks

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 38790 cycles 38538 cycles 1.01
ML-KEM-512 encaps 43688 cycles 43318 cycles 1.01
ML-KEM-512 decaps 54274 cycles 53838 cycles 1.01
ML-KEM-768 keypair 67589 cycles 68035 cycles 0.99
ML-KEM-768 encaps 76425 cycles 76846 cycles 0.99
ML-KEM-768 decaps 90698 cycles 91282 cycles 0.99
ML-KEM-1024 keypair 114107 cycles 113586 cycles 1.00
ML-KEM-1024 encaps 124618 cycles 124287 cycles 1.00
ML-KEM-1024 decaps 143304 cycles 142943 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton5

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 13434 cycles 13463 cycles 1.00
ML-KEM-512 encaps 15787 cycles 15776 cycles 1.00
ML-KEM-512 decaps 21299 cycles 21331 cycles 1.00
ML-KEM-768 keypair 22875 cycles 22838 cycles 1.00
ML-KEM-768 encaps 25296 cycles 25303 cycles 1.00
ML-KEM-768 decaps 33171 cycles 33124 cycles 1.00
ML-KEM-1024 keypair 33287 cycles 33226 cycles 1.00
ML-KEM-1024 encaps 36932 cycles 37038 cycles 1.00
ML-KEM-1024 decaps 47156 cycles 47155 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A72 (Raspberry Pi 4) benchmarks

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 52536 cycles 52580 cycles 1.00
ML-KEM-512 encaps 60402 cycles 60439 cycles 1.00
ML-KEM-512 decaps 76863 cycles 76532 cycles 1.00
ML-KEM-768 keypair 87637 cycles 88666 cycles 0.99
ML-KEM-768 encaps 95783 cycles 97336 cycles 0.98
ML-KEM-768 decaps 119523 cycles 121395 cycles 0.98
ML-KEM-1024 keypair 133351 cycles 132943 cycles 1.00
ML-KEM-1024 encaps 145589 cycles 145722 cycles 1.00
ML-KEM-1024 decaps 178446 cycles 180073 cycles 0.99

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A76 (Raspberry Pi 5) benchmarks

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 28246 cycles 28243 cycles 1.00
ML-KEM-512 encaps 34127 cycles 34119 cycles 1.00
ML-KEM-512 decaps 44531 cycles 44481 cycles 1.00
ML-KEM-768 keypair 47566 cycles 47617 cycles 1.00
ML-KEM-768 encaps 53875 cycles 53856 cycles 1.00
ML-KEM-768 decaps 68530 cycles 68499 cycles 1.00
ML-KEM-1024 keypair 70152 cycles 70190 cycles 1.00
ML-KEM-1024 encaps 78501 cycles 78569 cycles 1.00
ML-KEM-1024 decaps 98288 cycles 98321 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton3

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 18638 cycles 18655 cycles 1.00
ML-KEM-512 encaps 21833 cycles 21836 cycles 1.00
ML-KEM-512 decaps 28840 cycles 28877 cycles 1.00
ML-KEM-768 keypair 31535 cycles 31504 cycles 1.00
ML-KEM-768 encaps 34749 cycles 34746 cycles 1.00
ML-KEM-768 decaps 44804 cycles 44735 cycles 1.00
ML-KEM-1024 keypair 46145 cycles 46085 cycles 1.00
ML-KEM-1024 encaps 51439 cycles 51517 cycles 1.00
ML-KEM-1024 decaps 64928 cycles 64942 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton5 (no-opt)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 27343 cycles 27362 cycles 1.00
ML-KEM-512 encaps 31715 cycles 31620 cycles 1.00
ML-KEM-512 decaps 40426 cycles 40535 cycles 1.00
ML-KEM-768 keypair 44008 cycles 44091 cycles 1.00
ML-KEM-768 encaps 50702 cycles 50594 cycles 1.00
ML-KEM-768 decaps 62140 cycles 62271 cycles 1.00
ML-KEM-1024 keypair 68272 cycles 68282 cycles 1.00
ML-KEM-1024 encaps 75988 cycles 76085 cycles 1.00
ML-KEM-1024 decaps 90663 cycles 90625 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 4th gen (c7i)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 11870 cycles 11798 cycles 1.01
ML-KEM-512 encaps 13113 cycles 13069 cycles 1.00
ML-KEM-512 decaps 16994 cycles 17171 cycles 0.99
ML-KEM-768 keypair 19400 cycles 19388 cycles 1.00
ML-KEM-768 encaps 20643 cycles 20700 cycles 1.00
ML-KEM-768 decaps 26307 cycles 26389 cycles 1.00
ML-KEM-1024 keypair 28092 cycles 28192 cycles 1.00
ML-KEM-1024 encaps 29986 cycles 30153 cycles 0.99
ML-KEM-1024 decaps 37461 cycles 37577 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton3 (no-opt)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 38772 cycles 38739 cycles 1.00
ML-KEM-512 encaps 44450 cycles 44436 cycles 1.00
ML-KEM-512 decaps 56102 cycles 56164 cycles 1.00
ML-KEM-768 keypair 62346 cycles 62264 cycles 1.00
ML-KEM-768 encaps 70577 cycles 70509 cycles 1.00
ML-KEM-768 decaps 86248 cycles 86285 cycles 1.00
ML-KEM-1024 keypair 95594 cycles 95852 cycles 1.00
ML-KEM-1024 encaps 105883 cycles 105903 cycles 1.00
ML-KEM-1024 decaps 126037 cycles 125765 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 4th gen (c7i) (no-opt)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 27569 cycles 27515 cycles 1.00
ML-KEM-512 encaps 34312 cycles 34207 cycles 1.00
ML-KEM-512 decaps 43940 cycles 43827 cycles 1.00
ML-KEM-768 keypair 44250 cycles 44287 cycles 1.00
ML-KEM-768 encaps 54979 cycles 54814 cycles 1.00
ML-KEM-768 decaps 68202 cycles 68225 cycles 1.00
ML-KEM-1024 keypair 67521 cycles 67427 cycles 1.00
ML-KEM-1024 encaps 78756 cycles 78750 cycles 1.00
ML-KEM-1024 decaps 96062 cycles 96113 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton4

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 17667 cycles 17688 cycles 1.00
ML-KEM-512 encaps 20547 cycles 20546 cycles 1.00
ML-KEM-512 decaps 26995 cycles 27033 cycles 1.00
ML-KEM-768 keypair 29870 cycles 29832 cycles 1.00
ML-KEM-768 encaps 32701 cycles 32722 cycles 1.00
ML-KEM-768 decaps 41918 cycles 41866 cycles 1.00
ML-KEM-1024 keypair 43719 cycles 43685 cycles 1.00
ML-KEM-1024 encaps 48637 cycles 48702 cycles 1.00
ML-KEM-1024 decaps 61396 cycles 61368 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 3rd gen (c6a)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 14368 cycles 14429 cycles 1.00
ML-KEM-512 encaps 15900 cycles 15946 cycles 1.00
ML-KEM-512 decaps 21291 cycles 21331 cycles 1.00
ML-KEM-768 keypair 23536 cycles 23603 cycles 1.00
ML-KEM-768 encaps 24949 cycles 25013 cycles 1.00
ML-KEM-768 decaps 32783 cycles 32948 cycles 0.99
ML-KEM-1024 keypair 33473 cycles 33493 cycles 1.00
ML-KEM-1024 encaps 35775 cycles 35840 cycles 1.00
ML-KEM-1024 decaps 46215 cycles 46267 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton4 (no-opt)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 35215 cycles 35183 cycles 1.00
ML-KEM-512 encaps 40028 cycles 39969 cycles 1.00
ML-KEM-512 decaps 50563 cycles 50602 cycles 1.00
ML-KEM-768 keypair 56692 cycles 56650 cycles 1.00
ML-KEM-768 encaps 64081 cycles 63981 cycles 1.00
ML-KEM-768 decaps 78445 cycles 78485 cycles 1.00
ML-KEM-1024 keypair 87372 cycles 87509 cycles 1.00
ML-KEM-1024 encaps 96558 cycles 96613 cycles 1.00
ML-KEM-1024 decaps 114871 cycles 114678 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 4th gen (c7a)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 12765 cycles 12822 cycles 1.00
ML-KEM-512 encaps 14180 cycles 14217 cycles 1.00
ML-KEM-512 decaps 18947 cycles 18985 cycles 1.00
ML-KEM-768 keypair 21565 cycles 21613 cycles 1.00
ML-KEM-768 encaps 22681 cycles 22677 cycles 1.00
ML-KEM-768 decaps 29636 cycles 29688 cycles 1.00
ML-KEM-1024 keypair 30391 cycles 30443 cycles 1.00
ML-KEM-1024 encaps 32556 cycles 32629 cycles 1.00
ML-KEM-1024 decaps 42035 cycles 42042 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 3rd gen (c6a) (no-opt)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 39829 cycles 39879 cycles 1.00
ML-KEM-512 encaps 48164 cycles 48246 cycles 1.00
ML-KEM-512 decaps 61635 cycles 61724 cycles 1.00
ML-KEM-768 keypair 62541 cycles 62560 cycles 1.00
ML-KEM-768 encaps 75061 cycles 74791 cycles 1.00
ML-KEM-768 decaps 92286 cycles 92376 cycles 1.00
ML-KEM-1024 keypair 95102 cycles 95217 cycles 1.00
ML-KEM-1024 encaps 110240 cycles 110082 cycles 1.00
ML-KEM-1024 decaps 132655 cycles 132637 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 4th gen (c7a) (no-opt)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 36724 cycles 36751 cycles 1.00
ML-KEM-512 encaps 42850 cycles 42805 cycles 1.00
ML-KEM-512 decaps 55524 cycles 55557 cycles 1.00
ML-KEM-768 keypair 58371 cycles 58383 cycles 1.00
ML-KEM-768 encaps 67156 cycles 67012 cycles 1.00
ML-KEM-768 decaps 83911 cycles 83974 cycles 1.00
ML-KEM-1024 keypair 88490 cycles 88506 cycles 1.00
ML-KEM-1024 encaps 98941 cycles 99034 cycles 1.00
ML-KEM-1024 decaps 120418 cycles 120490 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 3rd gen (c6i)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 17918 cycles 17947 cycles 1.00
ML-KEM-512 encaps 20085 cycles 20118 cycles 1.00
ML-KEM-512 decaps 26739 cycles 26823 cycles 1.00
ML-KEM-768 keypair 30014 cycles 30079 cycles 1.00
ML-KEM-768 encaps 31808 cycles 31884 cycles 1.00
ML-KEM-768 decaps 41437 cycles 41523 cycles 1.00
ML-KEM-1024 keypair 44431 cycles 42190 cycles 1.05
ML-KEM-1024 encaps 44998 cycles 45298 cycles 0.99
ML-KEM-1024 decaps 58004 cycles 58302 cycles 0.99

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Performance Alert ⚠️

Possible performance regression was detected for benchmark 'Intel Xeon 3rd gen (c6i)'.
Benchmark result of this commit is worse than the previous benchmark result exceeding threshold 1.03.

Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-1024 keypair 44431 cycles 42190 cycles 1.05

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 3rd gen (c6i) (no-opt)

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 46958 cycles 46968 cycles 1.00
ML-KEM-512 encaps 55677 cycles 55707 cycles 1.00
ML-KEM-512 decaps 71370 cycles 71408 cycles 1.00
ML-KEM-768 keypair 74327 cycles 74383 cycles 1.00
ML-KEM-768 encaps 86064 cycles 85946 cycles 1.00
ML-KEM-768 decaps 106957 cycles 106930 cycles 1.00
ML-KEM-1024 keypair 111389 cycles 111331 cycles 1.00
ML-KEM-1024 encaps 126115 cycles 126009 cycles 1.00
ML-KEM-1024 decaps 152125 cycles 152070 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-M55 (NUCLEO-N657X0-Q) benchmarks

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 647367 cycles 647367 cycles 1
ML-KEM-512 encaps 728024 cycles 728024 cycles 1
ML-KEM-512 decaps 928692 cycles 928692 cycles 1
ML-KEM-768 keypair 1032214 cycles 1032214 cycles 1
ML-KEM-768 encaps 1162339 cycles 1162339 cycles 1
ML-KEM-768 decaps 1433590 cycles 1433590 cycles 1
ML-KEM-1024 keypair 1594936 cycles 1594936 cycles 1
ML-KEM-1024 encaps 1747347 cycles 1747347 cycles 1
ML-KEM-1024 decaps 2094866 cycles 2094866 cycles 1

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A55 (Snapdragon 888) benchmarks

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 59815 cycles 59783 cycles 1.00
ML-KEM-512 encaps 67496 cycles 67444 cycles 1.00
ML-KEM-512 decaps 86482 cycles 86410 cycles 1.00
ML-KEM-768 keypair 97550 cycles 97468 cycles 1.00
ML-KEM-768 encaps 111114 cycles 111015 cycles 1.00
ML-KEM-768 decaps 138412 cycles 138144 cycles 1.00
ML-KEM-1024 keypair 155175 cycles 154627 cycles 1.00
ML-KEM-1024 encaps 171950 cycles 171103 cycles 1.00
ML-KEM-1024 decaps 209844 cycles 207755 cycles 1.01

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot

oqs-bot commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-KEM-1024)

⚠️ Attention Required

Proof Status Current Previous Change
mlk_indcpa_keypair_derand ⚠️ 299s 176s +70%
mlk_polyvec_basemul_acc_montgomery_cached_c ⚠️ 133s 79s +68%
Full Results (191 proofs)
Proof Status Current Previous Change
**TOTAL** 1877s 1885s -0.4%
mlk_indcpa_enc 421s 402s +5%
mlk_indcpa_keypair_derand ⚠️ 299s 176s +70%
mlk_rej_uniform_c 156s 144s +8%
mlk_polyvec_basemul_acc_montgomery_cached_c ⚠️ 133s 79s +68%
rej_uniform_native_aarch64 45s 53s -15%
mlk_poly_reduce_native 41s 40s +2%
rej_uniform_native_x86_64 39s 67s -42%
mlk_poly_rej_uniform 38s 161s -76%
poly_ntt_native 35s 43s -19%
mlk_ntt_layer 32s 38s -16%
polyvec_basemul_acc_montgomery_cached_native 32s 31s +3%
mlk_keccak_squeezeblocks_x4 27s 26s +4%
keccakf1600x4_permute_native_x4 19s 15s +27%
mlk_polyvec_add 19s 18s +6%
mlk_poly_decompress_d11_native 18s 16s +12%
mlk_indcpa_dec 12s 11s +9%
mlk_poly_decompress_d5_native 12s 18s -33%
mlk_poly_frommsg 11s 10s +10%
mlk_poly_frombytes_native 10s 10s +0%
mlk_keccak_squeeze_once 9s 10s -10%
mlk_keccak_squeezeblocks 8s 9s -11%
mlk_ntt_butterfly_block 8s 9s -11%
kem_dec 7s 5s +40%
mlk_gen_matrix_serial 7s 4s +75%
mlk_poly_rej_uniform_x4 7s 7s +0%
mlk_invntt_layer 6s 6s +0%
mlk_keccakf1600_permute_c 6s 6s +0%
mlk_poly_ntt 6s 8s -25%
mlk_poly_sub 6s 8s -25%
mlk_scalar_decompress_d10 6s 2s +200%
mlk_shake128_squeezeblocks 6s 1s +500%
poly_frombytes_native_x86_64 6s 5s +20%
mlk_keccak_absorb_once_x4 5s 5s +0%
mlk_matvec_mul 5s 5s +0%
mlk_poly_mulcache_compute_c 5s 3s +67%
poly_decompress_d11_native_x86_64 5s 5s +0%
polyvec_basemul_acc_montgomery_cached_k4_native_x86_64 5s 2s +150%
mlk_enc_getnoise_eta1_eta2 4s 4s +0%
mlk_fqmul 4s 20s -80%
mlk_gen_matrix 4s 2s +100%
mlk_keccakf1600x4_permute 4s 2s +100%
mlk_poly_add 4s 5s -20%
mlk_poly_compress_d4 4s 3s +33%
mlk_poly_frombytes_c 4s 3s +33%
mlk_poly_getnoise_eta1_4x 4s 2s +100%
mlk_poly_tomsg 4s 3s +33%
mlk_polymat_permute_bitrev_to_custom 4s 5s -20%
mlk_polyvec_compress_du 4s 3s +33%
mlk_polyvec_frombytes 4s 5s -20%
mlk_polyvec_invntt_tomont 4s 1s +300%
mlk_polyvec_ntt 4s 12s -67%
mlk_polyvec_permute_bitrev_to_custom_native 4s 6s -33%
mlk_scalar_decompress_d11 4s 2s +100%
mlk_sha3_256 4s 2s +100%
mlk_shake128x4_absorb_once 4s 2s +100%
mlk_value_barrier_i32 4s 3s +33%
ntt_native_aarch64 4s 3s +33%
ntt_native_x86_64 4s 2s +100%
poly_decompress_d4_native_x86_64 4s 2s +100%
polyvec_basemul_acc_montgomery_cached_k3_native_x86_64 4s 2s +100%
polyvec_basemul_acc_montgomery_cached_k4_native_aarch64 4s 2s +100%
intt_native_x86_64 3s 2s +50%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 3s 3s +0%
keccakf1600x4_extract_bytes_native 3s 2s +50%
kem_check_sk 3s 3s +0%
kem_enc 3s 2s +50%
kem_enc_derand 3s 3s +0%
kem_keypair_derand 3s 1s +200%
mlk_barrett_reduce 3s 2s +50%
mlk_check_pct 3s 1s +200%
mlk_ct_cmask_nonzero_u8 3s 1s +200%
mlk_ct_get_optblocker_u32 3s 3s +0%
mlk_ct_get_optblocker_u8 3s 2s +50%
mlk_ct_memcmp 3s 3s +0%
mlk_keccak_absorb_once 3s 4s -25%
mlk_keccakf1600_extract_bytes 3s 2s +50%
mlk_keccakf1600x4_xor_bytes 3s 2s +50%
mlk_poly_cbd_eta2 3s 3s +0%
mlk_poly_compress_d4_c 3s 4s -25%
mlk_poly_compress_d5_native 3s 4s -25%
mlk_poly_decompress_d11_c 3s 2s +50%
mlk_poly_decompress_d4 3s 3s +0%
mlk_poly_decompress_d4_c 3s 4s -25%
mlk_poly_decompress_dv 3s 2s +50%
mlk_poly_frombytes 3s 2s +50%
mlk_poly_getnoise_eta1122_4x 3s 3s +0%
mlk_poly_reduce_c 3s 1s +200%
mlk_poly_tobytes 3s 3s +0%
mlk_poly_tobytes_native 3s 1s +200%
mlk_poly_tomont 3s 2s +50%
mlk_poly_tomont_native 3s 4s -25%
mlk_polyvec_basemul_acc_montgomery_cached 3s 2s +50%
mlk_shake256 3s 1s +200%
poly_compress_d11_native_x86_64 3s 3s +0%
poly_decompress_d5_native_x86_64 3s 5s -40%
poly_reduce_native_aarch64 3s 2s +50%
poly_reduce_native_x86_64 3s 3s +0%
poly_tobytes_native_aarch64 3s 3s +0%
poly_tobytes_native_x86_64 3s 3s +0%
intt_native_aarch64 2s 3s -33%
keccak_f1600_x4_native_aarch64_v84a 2s 3s -33%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 2s 3s -33%
keccak_f1600_x4_native_avx2 2s 2s +0%
keccakf1600_permute_native 2s 2s +0%
kem_check_pk 2s 4s -50%
kem_keypair 2s 2s +0%
mlk_ct_cmask_neg_i16 2s 3s -33%
mlk_ct_cmask_nonzero_u16 2s 2s +0%
mlk_ct_cmov_zero 2s 3s -33%
mlk_ct_get_optblocker_i32 2s 3s -33%
mlk_ct_sel_int16 2s 4s -50%
mlk_ct_sel_uint8 2s 2s +0%
mlk_keccakf1600_permute 2s 2s +0%
mlk_keccakf1600x4_extract_bytes 2s 2s +0%
mlk_keccakf1600x4_extract_bytes_c 2s 3s -33%
mlk_keccakf1600x4_xor_bytes_c 2s 5s -60%
mlk_keypair_getnoise_eta1 2s 5s -60%
mlk_montgomery_reduce 2s 2s +0%
mlk_poly_cbd_eta1 2s 3s -33%
mlk_poly_compress_d10 2s 2s +0%
mlk_poly_compress_d10_c 2s 3s -33%
mlk_poly_compress_d10_native 2s 3s -33%
mlk_poly_compress_d11_native 2s 3s -33%
mlk_poly_compress_d4_native 2s 2s +0%
mlk_poly_compress_d5_c 2s 4s -50%
mlk_poly_decompress_d10 2s 2s +0%
mlk_poly_decompress_d10_c 2s 2s +0%
mlk_poly_decompress_d10_native 2s 2s +0%
mlk_poly_decompress_d11 2s 3s -33%
mlk_poly_decompress_d5 2s 3s -33%
mlk_poly_decompress_du 2s 1s +100%
mlk_poly_getnoise_eta1_4x_native 2s 3s -33%
mlk_poly_getnoise_eta2 2s 1s +100%
mlk_poly_invntt_tomont 2s 3s -33%
mlk_poly_invntt_tomont_c 2s 2s +0%
mlk_poly_mulcache_compute 2s 2s +0%
mlk_poly_ntt_c 2s 1s +100%
mlk_poly_reduce 2s 2s +0%
mlk_poly_tobytes_c 2s 4s -50%
mlk_polyvec_decompress_du 2s 2s +0%
mlk_polyvec_mulcache_compute 2s 4s -50%
mlk_polyvec_permute_bitrev_to_custom 2s 4s -50%
mlk_polyvec_tobytes 2s 1s +100%
mlk_polyvec_tomont 2s 2s +0%
mlk_scalar_compress_d10 2s 2s +0%
mlk_scalar_compress_d11 2s 4s -50%
mlk_scalar_compress_d4 2s 1s +100%
mlk_scalar_decompress_d5 2s 2s +0%
mlk_scalar_signed_to_unsigned_q 2s 3s -33%
mlk_shake128_absorb_once 2s 2s +0%
mlk_shake256x4 2s 4s -50%
mlk_value_barrier_u32 2s 3s -33%
nttunpack_native_x86_64 2s 2s +0%
poly_compress_d10_native_x86_64 2s 2s +0%
poly_compress_d4_native_x86_64 2s 2s +0%
poly_compress_d5_native_x86_64 2s 4s -50%
poly_getnoise_eta1122_4x_native 2s 3s -33%
poly_invntt_tomont_native 2s 3s -33%
poly_mulcache_compute_native_aarch64 2s 3s -33%
polyvec_basemul_acc_montgomery_cached_k2_native_aarch64 2s 1s +100%
polyvec_basemul_acc_montgomery_cached_k2_native_x86_64 2s 2s +0%
rej_uniform_native 2s 1s +100%
keccak_f1600_x1_native_aarch64 1s 1s +0%
keccak_f1600_x1_native_aarch64_v84a 1s 2s -50%
keccakf1600x4_xor_bytes_native 1s 2s -50%
mlk_keccakf1600_extract_bytes (big endian) 1s 2s -50%
mlk_keccakf1600_xor_bytes 1s 1s +0%
mlk_keccakf1600_xor_bytes (big endian) 1s 2s -50%
mlk_poly_compress_d11 1s 3s -67%
mlk_poly_compress_d11_c 1s 4s -75%
mlk_poly_compress_d5 1s 5s -80%
mlk_poly_compress_du 1s 2s -50%
mlk_poly_compress_dv 1s 1s +0%
mlk_poly_decompress_d4_native 1s 3s -67%
mlk_poly_decompress_d5_c 1s 2s -50%
mlk_poly_mulcache_compute_native 1s 3s -67%
mlk_poly_tomont_c 1s 3s -67%
mlk_polyvec_reduce 1s 2s -50%
mlk_rej_uniform 1s 2s -50%
mlk_scalar_compress_d1 1s 3s -67%
mlk_scalar_compress_d5 1s 4s -75%
mlk_scalar_decompress_d4 1s 3s -67%
mlk_sha3_512 1s 1s +0%
mlk_shake128x4_squeezeblocks 1s 2s -50%
mlk_value_barrier_u8 1s 1s +0%
poly_decompress_d10_native_x86_64 1s 4s -75%
poly_mulcache_compute_native_x86_64 1s 5s -80%
poly_tomont_native_aarch64 1s 3s -67%
poly_tomont_native_x86_64 1s 2s -50%
polyvec_basemul_acc_montgomery_cached_k3_native_aarch64 1s 2s -50%
sys_check_capability 1s 2s -50%

@oqs-bot

oqs-bot commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-KEM-768)

⚠️ Attention Required

Proof Status Current Previous Change
mlk_indcpa_dec ⚠️ 22s 10s +120%
mlk_indcpa_enc ⚠️ 400s 252s +59%
mlk_polyvec_basemul_acc_montgomery_cached_c ⚠️ 73s 46s +59%
Full Results (191 proofs)
Proof Status Current Previous Change
**TOTAL** 1629s 1645s -1.0%
mlk_indcpa_enc ⚠️ 400s 252s +59%
mlk_rej_uniform_c 152s 114s +33%
mlk_indcpa_keypair_derand 127s 318s -60%
mlk_polyvec_basemul_acc_montgomery_cached_c ⚠️ 73s 46s +59%
rej_uniform_native_aarch64 46s 45s +2%
poly_ntt_native 40s 39s +3%
rej_uniform_native_x86_64 39s 55s -29%
mlk_ntt_layer 37s 29s +28%
mlk_poly_reduce_native 35s 34s +3%
mlk_poly_rej_uniform 35s 129s -73%
mlk_keccak_squeezeblocks_x4 27s 24s +12%
mlk_indcpa_dec ⚠️ 22s 10s +120%
keccakf1600x4_permute_native_x4 18s 21s -14%
polyvec_basemul_acc_montgomery_cached_native 17s 14s +21%
mlk_poly_decompress_d10_native 15s 15s +0%
mlk_poly_decompress_d4_native 15s 12s +25%
mlk_polyvec_add 13s 12s +8%
mlk_keccak_squeezeblocks 12s 8s +50%
mlk_poly_frombytes_native 11s 8s +38%
mlk_poly_frommsg 11s 9s +22%
mlk_poly_ntt 11s 5s +120%
mlk_ntt_butterfly_block 9s 7s +29%
mlk_poly_rej_uniform_x4 9s 5s +80%
mlk_poly_sub 9s 10s -10%
mlk_keccak_absorb_once_x4 8s 5s +60%
poly_decompress_d10_native_x86_64 8s 5s +60%
kem_dec 7s 5s +40%
mlk_gen_matrix_serial 7s 2s +250%
mlk_keccak_squeeze_once 7s 7s +0%
mlk_poly_add 6s 7s -14%
mlk_invntt_layer 5s 4s +25%
mlk_poly_compress_d11 5s 1s +400%
mlk_poly_compress_d5 5s 2s +150%
mlk_poly_invntt_tomont_c 5s 2s +150%
mlk_poly_tomsg 5s 3s +67%
mlk_polyvec_mulcache_compute 5s 2s +150%
poly_frombytes_native_x86_64 5s 5s +0%
kem_check_pk 4s 2s +100%
mlk_fqmul 4s 16s -75%
mlk_gen_matrix 4s 2s +100%
mlk_keccak_absorb_once 4s 3s +33%
mlk_keccakf1600_permute_c 4s 6s -33%
mlk_poly_compress_dv 4s 1s +300%
mlk_poly_decompress_d5_native 4s 3s +33%
mlk_poly_frombytes_c 4s 4s +0%
mlk_poly_getnoise_eta1_4x 4s 2s +100%
mlk_poly_mulcache_compute_c 4s 2s +100%
mlk_poly_tobytes_native 4s 1s +300%
mlk_polymat_permute_bitrev_to_custom 4s 3s +33%
mlk_polyvec_permute_bitrev_to_custom_native 4s 2s +100%
mlk_rej_uniform 4s 1s +300%
mlk_scalar_compress_d4 4s 2s +100%
mlk_value_barrier_i32 4s 2s +100%
poly_decompress_d4_native_x86_64 4s 6s -33%
poly_tomont_native_x86_64 4s 1s +300%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 3s 3s +0%
keccak_f1600_x4_native_avx2 3s 2s +50%
kem_enc_derand 3s 3s +0%
kem_keypair 3s 2s +50%
kem_keypair_derand 3s 2s +50%
mlk_check_pct 3s 3s +0%
mlk_ct_cmask_neg_i16 3s 3s +0%
mlk_ct_cmov_zero 3s 3s +0%
mlk_ct_memcmp 3s 4s -25%
mlk_keccakf1600_xor_bytes 3s 3s +0%
mlk_montgomery_reduce 3s 1s +200%
mlk_poly_cbd_eta2 3s 2s +50%
mlk_poly_compress_d10_c 3s 2s +50%
mlk_poly_compress_d10_native 3s 2s +50%
mlk_poly_compress_d4 3s 3s +0%
mlk_poly_decompress_d10 3s 2s +50%
mlk_poly_decompress_d10_c 3s 1s +200%
mlk_poly_decompress_d11_native 3s 1s +200%
mlk_poly_decompress_d5 3s 2s +50%
mlk_poly_decompress_d5_c 3s 3s +0%
mlk_poly_getnoise_eta1_4x_native 3s 3s +0%
mlk_poly_mulcache_compute 3s 2s +50%
mlk_poly_mulcache_compute_native 3s 1s +200%
mlk_poly_reduce 3s 2s +50%
mlk_poly_reduce_c 3s 3s +0%
mlk_poly_tobytes 3s 3s +0%
mlk_poly_tomont 3s 1s +200%
mlk_polyvec_decompress_du 3s 4s -25%
mlk_polyvec_invntt_tomont 3s 1s +200%
mlk_polyvec_permute_bitrev_to_custom 3s 1s +200%
mlk_polyvec_reduce 3s 1s +200%
mlk_polyvec_tobytes 3s 1s +200%
mlk_scalar_compress_d11 3s 1s +200%
mlk_scalar_signed_to_unsigned_q 3s 5s -40%
mlk_sha3_256 3s 2s +50%
mlk_shake256 3s 2s +50%
mlk_shake256x4 3s 3s +0%
mlk_value_barrier_u8 3s 2s +50%
nttunpack_native_x86_64 3s 3s +0%
poly_compress_d11_native_x86_64 3s 3s +0%
poly_compress_d5_native_x86_64 3s 2s +50%
poly_decompress_d11_native_x86_64 3s 1s +200%
poly_reduce_native_aarch64 3s 2s +50%
poly_reduce_native_x86_64 3s 2s +50%
poly_tomont_native_aarch64 3s 1s +200%
polyvec_basemul_acc_montgomery_cached_k2_native_aarch64 3s 3s +0%
polyvec_basemul_acc_montgomery_cached_k4_native_x86_64 3s 2s +50%
sys_check_capability 3s 3s +0%
intt_native_aarch64 2s 5s -60%
intt_native_x86_64 2s 3s -33%
keccak_f1600_x4_native_aarch64_v84a 2s 2s +0%
keccakf1600x4_extract_bytes_native 2s 3s -33%
keccakf1600x4_xor_bytes_native 2s 2s +0%
kem_check_sk 2s 2s +0%
kem_enc 2s 2s +0%
mlk_barrett_reduce 2s 2s +0%
mlk_ct_cmask_nonzero_u16 2s 4s -50%
mlk_ct_cmask_nonzero_u8 2s 4s -50%
mlk_ct_get_optblocker_i32 2s 1s +100%
mlk_ct_get_optblocker_u32 2s 1s +100%
mlk_ct_sel_uint8 2s 1s +100%
mlk_enc_getnoise_eta1_eta2 2s 3s -33%
mlk_keccakf1600_extract_bytes 2s 4s -50%
mlk_keccakf1600_extract_bytes (big endian) 2s 2s +0%
mlk_keccakf1600_permute 2s 1s +100%
mlk_keccakf1600x4_extract_bytes_c 2s 2s +0%
mlk_keccakf1600x4_permute 2s 4s -50%
mlk_keypair_getnoise_eta1 2s 2s +0%
mlk_poly_compress_d11_c 2s 1s +100%
mlk_poly_compress_d11_native 2s 4s -50%
mlk_poly_compress_d4_c 2s 1s +100%
mlk_poly_compress_d4_native 2s 3s -33%
mlk_poly_compress_d5_c 2s 1s +100%
mlk_poly_compress_d5_native 2s 2s +0%
mlk_poly_compress_du 2s 1s +100%
mlk_poly_decompress_d11 2s 2s +0%
mlk_poly_decompress_d11_c 2s 3s -33%
mlk_poly_decompress_d4 2s 3s -33%
mlk_poly_decompress_d4_c 2s 4s -50%
mlk_poly_decompress_du 2s 1s +100%
mlk_poly_getnoise_eta1122_4x 2s 1s +100%
mlk_poly_invntt_tomont 2s 4s -50%
mlk_poly_ntt_c 2s 1s +100%
mlk_poly_tomont_c 2s 3s -33%
mlk_poly_tomont_native 2s 3s -33%
mlk_polyvec_basemul_acc_montgomery_cached 2s 4s -50%
mlk_polyvec_compress_du 2s 1s +100%
mlk_polyvec_frombytes 2s 3s -33%
mlk_polyvec_tomont 2s 1s +100%
mlk_scalar_compress_d1 2s 2s +0%
mlk_scalar_compress_d10 2s 2s +0%
mlk_scalar_compress_d5 2s 1s +100%
mlk_scalar_decompress_d10 2s 2s +0%
mlk_scalar_decompress_d4 2s 2s +0%
mlk_scalar_decompress_d5 2s 2s +0%
mlk_sha3_512 2s 3s -33%
mlk_shake128_absorb_once 2s 1s +100%
mlk_shake128x4_absorb_once 2s 3s -33%
mlk_shake128x4_squeezeblocks 2s 1s +100%
ntt_native_aarch64 2s 3s -33%
ntt_native_x86_64 2s 3s -33%
poly_compress_d10_native_x86_64 2s 5s -60%
poly_decompress_d5_native_x86_64 2s 4s -50%
poly_getnoise_eta1122_4x_native 2s 1s +100%
poly_invntt_tomont_native 2s 2s +0%
poly_mulcache_compute_native_aarch64 2s 6s -67%
poly_mulcache_compute_native_x86_64 2s 3s -33%
poly_tobytes_native_x86_64 2s 3s -33%
polyvec_basemul_acc_montgomery_cached_k2_native_x86_64 2s 4s -50%
polyvec_basemul_acc_montgomery_cached_k3_native_aarch64 2s 2s +0%
polyvec_basemul_acc_montgomery_cached_k3_native_x86_64 2s 2s +0%
polyvec_basemul_acc_montgomery_cached_k4_native_aarch64 2s 1s +100%
keccak_f1600_x1_native_aarch64 1s 1s +0%
keccak_f1600_x1_native_aarch64_v84a 1s 3s -67%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 1s 1s +0%
keccakf1600_permute_native 1s 3s -67%
mlk_ct_get_optblocker_u8 1s 2s -50%
mlk_ct_sel_int16 1s 3s -67%
mlk_keccakf1600_xor_bytes (big endian) 1s 1s +0%
mlk_keccakf1600x4_extract_bytes 1s 4s -75%
mlk_keccakf1600x4_xor_bytes 1s 2s -50%
mlk_keccakf1600x4_xor_bytes_c 1s 1s +0%
mlk_matvec_mul 1s 4s -75%
mlk_poly_cbd_eta1 1s 1s +0%
mlk_poly_compress_d10 1s 2s -50%
mlk_poly_decompress_dv 1s 2s -50%
mlk_poly_frombytes 1s 2s -50%
mlk_poly_getnoise_eta2 1s 2s -50%
mlk_poly_tobytes_c 1s 1s +0%
mlk_polyvec_ntt 1s 4s -75%
mlk_scalar_decompress_d11 1s 3s -67%
mlk_shake128_squeezeblocks 1s 3s -67%
mlk_value_barrier_u32 1s 2s -50%
poly_compress_d4_native_x86_64 1s 3s -67%
poly_tobytes_native_aarch64 1s 1s +0%
rej_uniform_native 1s 2s -50%

@oqs-bot

oqs-bot commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-KEM-512)

⚠️ Attention Required

Proof Status Current Previous Change
mlk_indcpa_enc ⚠️ 385s 255s +51%
mlk_indcpa_keypair_derand ⚠️ 406s 213s +91%
Full Results (191 proofs)
Proof Status Current Previous Change
**TOTAL** 1844s 1614s +14.3%
mlk_indcpa_keypair_derand ⚠️ 406s 213s +91%
mlk_indcpa_enc ⚠️ 385s 255s +51%
mlk_rej_uniform_c 150s 135s +11%
mlk_polyvec_basemul_acc_montgomery_cached_c 62s 55s +13%
rej_uniform_native_aarch64 43s 43s +0%
rej_uniform_native_x86_64 42s 57s -26%
poly_ntt_native 39s 45s -13%
mlk_poly_reduce_native 37s 35s +6%
mlk_poly_rej_uniform 36s 150s -76%
mlk_ntt_layer 34s 32s +6%
mlk_keccak_squeezeblocks_x4 26s 26s +0%
mlk_indcpa_dec 18s 14s +29%
keccakf1600x4_permute_native_x4 17s 16s +6%
mlk_poly_decompress_d10_native 14s 15s -7%
mlk_poly_decompress_d4_native 14s 14s +0%
mlk_polyvec_add 13s 11s +18%
mlk_poly_frommsg 10s 10s +0%
mlk_ntt_butterfly_block 9s 9s +0%
mlk_keccak_squeeze_once 8s 8s +0%
mlk_poly_rej_uniform_x4 8s 8s +0%
polyvec_basemul_acc_montgomery_cached_native 8s 8s +0%
mlk_keccak_squeezeblocks 7s 9s -22%
mlk_keccakf1600_permute_c 7s 5s +40%
mlk_poly_frombytes_native 7s 8s -12%
mlk_poly_sub 7s 6s +17%
kem_dec 6s 5s +20%
mlk_invntt_layer 6s 4s +50%
mlk_keccak_absorb_once 6s 3s +100%
mlk_keccak_absorb_once_x4 6s 5s +20%
poly_decompress_d10_native_x86_64 6s 9s -33%
mlk_fqmul 5s 17s -71%
mlk_poly_getnoise_eta1_4x_native 5s 2s +150%
mlk_poly_ntt 5s 5s +0%
mlk_polymat_permute_bitrev_to_custom 5s 2s +150%
ntt_native_aarch64 5s 2s +150%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 4s 3s +33%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 4s 3s +33%
kem_keypair_derand 4s 3s +33%
mlk_check_pct 4s 5s -20%
mlk_enc_getnoise_eta1_eta2 4s 2s +100%
mlk_keccakf1600_xor_bytes 4s 1s +300%
mlk_poly_add 4s 4s +0%
mlk_poly_compress_d4 4s 3s +33%
mlk_poly_compress_d4_c 4s 3s +33%
mlk_poly_frombytes_c 4s 4s +0%
mlk_poly_tomont 4s 1s +300%
mlk_polyvec_compress_du 4s 2s +100%
mlk_scalar_compress_d5 4s 2s +100%
mlk_scalar_decompress_d5 4s 3s +33%
mlk_shake128_absorb_once 4s 3s +33%
mlk_shake128x4_squeezeblocks 4s 1s +300%
ntt_native_x86_64 4s 2s +100%
poly_decompress_d4_native_x86_64 4s 4s +0%
poly_frombytes_native_x86_64 4s 5s -20%
poly_mulcache_compute_native_x86_64 4s 1s +300%
poly_tomont_native_aarch64 4s 2s +100%
polyvec_basemul_acc_montgomery_cached_k3_native_aarch64 4s 3s +33%
polyvec_basemul_acc_montgomery_cached_k4_native_x86_64 4s 2s +100%
intt_native_x86_64 3s 2s +50%
keccak_f1600_x1_native_aarch64_v84a 3s 2s +50%
kem_keypair 3s 4s -25%
mlk_ct_cmask_nonzero_u8 3s 1s +200%
mlk_ct_memcmp 3s 3s +0%
mlk_gen_matrix_serial 3s 3s +0%
mlk_keccakf1600_extract_bytes 3s 2s +50%
mlk_keccakf1600_permute 3s 1s +200%
mlk_keccakf1600_xor_bytes (big endian) 3s 2s +50%
mlk_keccakf1600x4_extract_bytes_c 3s 3s +0%
mlk_poly_cbd_eta1 3s 2s +50%
mlk_poly_cbd_eta2 3s 3s +0%
mlk_poly_compress_d10 3s 4s -25%
mlk_poly_compress_d10_native 3s 1s +200%
mlk_poly_compress_d11 3s 3s +0%
mlk_poly_compress_d11_c 3s 2s +50%
mlk_poly_compress_d5_native 3s 3s +0%
mlk_poly_compress_du 3s 3s +0%
mlk_poly_decompress_dv 3s 2s +50%
mlk_poly_frombytes 3s 1s +200%
mlk_poly_getnoise_eta2 3s 3s +0%
mlk_poly_invntt_tomont_c 3s 3s +0%
mlk_poly_mulcache_compute_c 3s 4s -25%
mlk_poly_mulcache_compute_native 3s 2s +50%
mlk_poly_tobytes 3s 3s +0%
mlk_poly_tobytes_native 3s 2s +50%
mlk_polyvec_basemul_acc_montgomery_cached 3s 1s +200%
mlk_polyvec_frombytes 3s 2s +50%
mlk_polyvec_invntt_tomont 3s 3s +0%
mlk_polyvec_ntt 3s 1s +200%
mlk_polyvec_permute_bitrev_to_custom 3s 2s +50%
mlk_polyvec_permute_bitrev_to_custom_native 3s 2s +50%
mlk_scalar_compress_d11 3s 2s +50%
mlk_scalar_decompress_d11 3s 2s +50%
mlk_scalar_decompress_d4 3s 3s +0%
mlk_sha3_256 3s 2s +50%
mlk_shake128x4_absorb_once 3s 2s +50%
mlk_shake256x4 3s 5s -40%
mlk_value_barrier_i32 3s 2s +50%
mlk_value_barrier_u32 3s 3s +0%
poly_compress_d5_native_x86_64 3s 2s +50%
poly_decompress_d11_native_x86_64 3s 1s +200%
poly_getnoise_eta1122_4x_native 3s 1s +200%
poly_mulcache_compute_native_aarch64 3s 3s +0%
poly_reduce_native_aarch64 3s 1s +200%
poly_reduce_native_x86_64 3s 2s +50%
poly_tomont_native_x86_64 3s 2s +50%
polyvec_basemul_acc_montgomery_cached_k2_native_aarch64 3s 3s +0%
polyvec_basemul_acc_montgomery_cached_k2_native_x86_64 3s 3s +0%
polyvec_basemul_acc_montgomery_cached_k3_native_x86_64 3s 2s +50%
rej_uniform_native 3s 1s +200%
intt_native_aarch64 2s 2s +0%
keccak_f1600_x1_native_aarch64 2s 2s +0%
keccak_f1600_x4_native_aarch64_v84a 2s 3s -33%
keccak_f1600_x4_native_avx2 2s 2s +0%
keccakf1600_permute_native 2s 1s +100%
keccakf1600x4_extract_bytes_native 2s 1s +100%
keccakf1600x4_xor_bytes_native 2s 1s +100%
kem_check_pk 2s 4s -50%
kem_enc 2s 2s +0%
mlk_barrett_reduce 2s 2s +0%
mlk_ct_cmask_neg_i16 2s 3s -33%
mlk_ct_cmask_nonzero_u16 2s 2s +0%
mlk_ct_cmov_zero 2s 3s -33%
mlk_ct_get_optblocker_i32 2s 2s +0%
mlk_ct_get_optblocker_u8 2s 2s +0%
mlk_ct_sel_int16 2s 1s +100%
mlk_ct_sel_uint8 2s 2s +0%
mlk_gen_matrix 2s 2s +0%
mlk_keccakf1600x4_extract_bytes 2s 2s +0%
mlk_keccakf1600x4_xor_bytes 2s 2s +0%
mlk_keypair_getnoise_eta1 2s 1s +100%
mlk_matvec_mul 2s 3s -33%
mlk_montgomery_reduce 2s 2s +0%
mlk_poly_compress_d10_c 2s 2s +0%
mlk_poly_compress_d5 2s 2s +0%
mlk_poly_compress_d5_c 2s 3s -33%
mlk_poly_decompress_d11 2s 2s +0%
mlk_poly_decompress_d11_c 2s 2s +0%
mlk_poly_decompress_d11_native 2s 3s -33%
mlk_poly_decompress_d4 2s 3s -33%
mlk_poly_decompress_d5 2s 1s +100%
mlk_poly_decompress_d5_c 2s 1s +100%
mlk_poly_decompress_du 2s 2s +0%
mlk_poly_getnoise_eta1_4x 2s 5s -60%
mlk_poly_mulcache_compute 2s 2s +0%
mlk_poly_ntt_c 2s 2s +0%
mlk_poly_reduce_c 2s 3s -33%
mlk_poly_tomont_c 2s 2s +0%
mlk_poly_tomont_native 2s 2s +0%
mlk_polyvec_mulcache_compute 2s 5s -60%
mlk_polyvec_tomont 2s 1s +100%
mlk_rej_uniform 2s 5s -60%
mlk_scalar_compress_d1 2s 2s +0%
mlk_scalar_compress_d10 2s 2s +0%
mlk_scalar_compress_d4 2s 1s +100%
mlk_scalar_signed_to_unsigned_q 2s 1s +100%
mlk_shake128_squeezeblocks 2s 3s -33%
mlk_value_barrier_u8 2s 4s -50%
nttunpack_native_x86_64 2s 2s +0%
poly_compress_d11_native_x86_64 2s 2s +0%
poly_compress_d4_native_x86_64 2s 1s +100%
poly_invntt_tomont_native 2s 2s +0%
poly_tobytes_native_aarch64 2s 2s +0%
poly_tobytes_native_x86_64 2s 1s +100%
polyvec_basemul_acc_montgomery_cached_k4_native_aarch64 2s 3s -33%
kem_check_sk 1s 2s -50%
kem_enc_derand 1s 3s -67%
mlk_ct_get_optblocker_u32 1s 3s -67%
mlk_keccakf1600_extract_bytes (big endian) 1s 5s -80%
mlk_keccakf1600x4_permute 1s 4s -75%
mlk_keccakf1600x4_xor_bytes_c 1s 2s -50%
mlk_poly_compress_d11_native 1s 3s -67%
mlk_poly_compress_d4_native 1s 2s -50%
mlk_poly_compress_dv 1s 3s -67%
mlk_poly_decompress_d10 1s 2s -50%
mlk_poly_decompress_d10_c 1s 4s -75%
mlk_poly_decompress_d4_c 1s 7s -86%
mlk_poly_decompress_d5_native 1s 2s -50%
mlk_poly_getnoise_eta1122_4x 1s 3s -67%
mlk_poly_invntt_tomont 1s 2s -50%
mlk_poly_reduce 1s 1s +0%
mlk_poly_tobytes_c 1s 3s -67%
mlk_poly_tomsg 1s 3s -67%
mlk_polyvec_decompress_du 1s 2s -50%
mlk_polyvec_reduce 1s 2s -50%
mlk_polyvec_tobytes 1s 3s -67%
mlk_scalar_decompress_d10 1s 2s -50%
mlk_sha3_512 1s 2s -50%
mlk_shake256 1s 2s -50%
poly_compress_d10_native_x86_64 1s 3s -67%
poly_decompress_d5_native_x86_64 1s 2s -50%
sys_check_capability 1s 1s +0%

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SpacemiT K1 8 (Banana Pi F3) benchmarks

Details
Benchmark suite Current: af9abd7 Previous: 69d24e3 Ratio
ML-KEM-512 keypair 143635 cycles 143641 cycles 1.00
ML-KEM-512 encaps 151967 cycles 151977 cycles 1.00
ML-KEM-512 decaps 190396 cycles 190397 cycles 1.00
ML-KEM-768 keypair 233958 cycles 233991 cycles 1.00
ML-KEM-768 encaps 251113 cycles 251136 cycles 1.00
ML-KEM-768 decaps 306473 cycles 306055 cycles 1.00
ML-KEM-1024 keypair 365615 cycles 365553 cycles 1.00
ML-KEM-1024 encaps 388948 cycles 387979 cycles 1.00
ML-KEM-1024 decaps 459221 cycles 458363 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

Barrett reduction followed by a conditional subtraction needs 8
instructions per vector to map an int16 to [0,q). The sequence

  vpmulhw t,a,L   vpmullw u,a,H   vpaddw m,t,u
  vpaddw  x,m,S   vpmulhuw z,x,Q

needs 5.

Signed-off-by: Matthias J. Kannwischer <matthias@zerorisc.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

benchmark this PR should be benchmarked in CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants