ntt: two DRAM sweeps per encode, F192 on the F64 path, shift-XOR reduction - #289
Merged
Merged
Conversation
…ction A large encode is bound by memory bandwidth, so its cost is its number of sweeps of the codeword. On a Ryzen 9 9950X3D one read+write sweep of the 940 MB XMSS-900 codeword takes about 41 ms; the base encode cost about 3.7. - One gathered pass runs the top layers per L2-resident row group and builds the replicas straight from the message; L2-resident sub-blocks finish the rest. A transform that fits L3 sub-blocks is a single sweep. - The F192 encode is the F64 transform over three lanes per coefficient, which deletes the dedicated F192 transform, its AVX-512 and NEON kernels, and the separate replicate pass in ligero_commit_ext. - Streaming stores for the transpose tiles, the first pass's writes and sub-blocks built in scratch. - The AVX-512 butterfly reduces with shifts and vpternlogq instead of two more carry-less multiplies. Zen 5, rustc 1.98, target-cpu=native, 32 threads, interleaved with HEAD: transpose + base encode, 2^21 x 56 F64 181 ms -> 103 ms (1.76x) F192 encode 2^20 x 16, rate 1/16 75.3 ms -> 28.3 ms (2.66x) F192 encode 2^19 x 16, rate 1/128 26.3 ms -> 6.2 ms (4.2x) butterfly kernel, L2-resident, 1 thread 0.277 -> 0.165 ns/butterfly aggregate --xmss 900 --log-inv-rate 1 1.068 s -> 0.908 s (-15%) aggregate --sphincs 220 --log-inv-rate 1 1.716 s -> 1.547 s (-10%) recursion --n 2 --xmss-per-leaf 900 -r 2 0.582 s -> 0.455 s (-22%) aggregate --blobs 8 --log-inv-rate 1 1.270 s -> 1.095 s (-14%) Proof sizes are unchanged: the codewords are bit-identical. Peak memory rises 40 to 80 MB from per-thread scratch. The NEON arm is type-checked only, not measured on Apple silicon. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
@TomWambsgans Please double check for the Mac adaption, I expect again the speedup not to be identical on Mac device :) |
On Apple silicon the two-sweep plan made the base encode slower than main, by 16% for XMSS-900 and 20% for SPHINCS-220. The encode there is compute bound rather than bound by DRAM sweeps: timed per phase on the XMSS-900 codeword, a deep layer costs about 2.1 ms whether its sub-block is 458 KiB or 14.7 MiB, while a gathered layer costs about 3.3 ms, its rows scattered over more streams than the prefetcher follows. So `L2_WORDS` grows to 2^21 words there, which moves layers from the gathered pass into the deep one, and `LOG_SUBS_PER_WORKER` cuts four deep sub-blocks per worker instead of one, so the performance cores do not wait on the efficiency cores' share. The larger budget alone slowed the 2^20-row F192 encodes, which the extra sub-blocks recover. Both are gated on `all(target_arch = "aarch64", target_os = "macos")`, as in `parallel::topology`. Elsewhere they keep the PR's values, so the Zen plan is unchanged: a `compile_error!` in the other arm fails the x86 cross-check and not the host build. Codewords are bit-identical across every budget tried. M4 Max, medians of 9 passes, alternating main (2e6fe21), the PR as submitted (0b1fb53) and this: - base encode (transpose + NTT): XMSS-900 52.5, 61.0, 54.9 ms; SPHINCS-220 79.6, 96.3, 79.2 ms - F192 encodes, all levels: XMSS-900 30.5, 26.6, 27.2 ms; SPHINCS-220 58.7, 60.5, 56.8 ms - Commit + PCS open: XMSS-900 246.6, 244.8, 238.7 ms; SPHINCS-220 376.3, 387.6, 367.9 ms End to end, medians of 12 passes in the same order, near the run-to-run spread since the encode is a small share of a proof: - aggregate --xmss 900 --log-inv-rate 1: 544.0, 552.5, 548.5 ms - aggregate --sphincs 220 --log-inv-rate 1: 839.0, 850.5, 827.5 ms - recursion --n 2 --xmss-per-leaf 900 --log-inv-rate 2: 278.0, 284.5, 280.0 ms - aggregate --blobs 8 --log-inv-rate 1: 560.5, 555.5, 551.0 ms Proof sizes are unchanged and peak memory moves by under 0.1 GiB. The test comments naming the plan each shape forces now say those are the plans off Apple silicon. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Collaborator
|
nice |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
In one line
The WHIR encodes now make two passes over memory instead of three or four, so proving gets 10 to 22% faster across the four benchmarks, with bit-identical codewords.
The problem
A large NTT here is not limited by arithmetic. It is limited by memory bandwidth.
On a Ryzen 9 9950X3D, one read+write sweep of the 940 MB XMSS-900 codeword takes about 41 ms.
The base encode cost about 3.7 of those sweeps:
The F192 encodes of the deeper WHIR levels were worse off: a separate replicate pass, then three more sweeps.
The idea
Put as many layers as possible into each sweep, so each byte of the codeword travels to DRAM and back as few times as possible.
glayers pair with each other, so in scratch they form a small transform of their own, with the same twiddles.Four changes
rlayers on the zero-padded message, every block is a copy of it. So no pass fills the replicas only to read them back.nF192 lanes is exactly the F64 transform over3nlanes. This deletes the dedicated F192 transform, its AVX-512 and NEON kernels, and the replicate pass inligero_commit_ext.vpternlogq 0x96replace two more carry-less multiplies per product, since Zen 5's carry-less multiplier is the scarce unit. On its own this takes the kernel from 0.277 to 0.165 ns per butterfly.Tried and dropped: keeping all 8 rows of a radix-8 group in registers. It measured flat to slightly worse, because the kernel is bound by the multiplier and ALU, not by L1 loads.
Benchmarks
Zen 5 (Ryzen 9 9950X3D), rustc 1.98,
target-cpu=native, 32 threads.HEAD and this branch were run interleaved.
End to end,
--repeat 3, two rounds each:aggregate --xmss 900 --log-inv-rate 1aggregate --sphincs 220 --log-inv-rate 1recursion --n 2 --xmss-per-leaf 900 --log-inv-rate 2aggregate --blobs 8 --log-inv-rate 1The NTT stages alone, medians of 21 runs, three interleaved rounds, p10 to p90 within 2%:
How close to the limit, measured pass by pass:
Checks
cargo testallpasses, with and withoutZK_ALLOC_POISON=1; clippy, rustdoc and fmt are clean.x86-64-v3) and plainx86-64builds.compile_error!confirmed the NEON arm is compiled.For the reviewer
The benchmark harness behind the NTT numbers (not committed)
Append it to the end of
crates/pcs/src/whir_ntt_ext.rs, then run:cargo test --release -p pcs --lib ntt_throughput -- --ignored --nocapture🤖 Generated with Claude Code