Feat/multi merkle tree gpu resident - #951
Draft
ColoCarletti wants to merge 26 commits into
Draft
Conversation
…he batched prover
…r (hybrid resident handles)
…E download for plain tables)
Per table the batched prover still ran a redundant HOST LDE FFT (main+aux) to feed prep trees / retained_main / later phases, leaving the GPU idle. This completes the device port: - aux/main: resident device expand first; host-expand only for the small-table fallback / Retain. A per-table commit-stream sync closes an async race the old host FFT delay had hidden. - OOD (phase 4): device-recompute device-only + read the trace OOD via the GPU barycentric fast path (no host coset eval, no D2H download). Byte-identical proof (roots unchanged); all batched tests pass. e22 mainnet prove 284s -> 155s (-45%), GPU util 17% -> 44%. Budget test updated: phase 4 is now a full device expansion.
Route the batched aux build through the resident GPU build (fingerprints + term columns + running-sum accumulate) and bulk-download the result into the host aux table, and transpose the main trace to column-major on device (ResidentMain::HostRowMajor) instead of on the host. Removes the host set_aux writes, the host accumulate and the ~1.36s/epoch columns_main transpose. Byte-identical; aux_commit 3.64s -> 1.52s/epoch.
Upload the composition parts to a device handle in the batched deep_codeword so the DEEP composition takes the fully-resident arm (device parts + device inv-denoms) instead of the host build_r4_inv_denoms_cpu batch-inverse. Byte-identical; deep_fri 3.96s -> 2.22s/epoch.
The parts D2H de-interleaved the 3 ext3 slabs into per-column row-major ext3 on the host (a strided gather, ~40% of the download). Do it on device via a new interleave_ext3_slabs kernel + per-part D2H straight into the owning buffers (no host copy). Byte-identical; comp_parts 3.02s -> 2.63s/epoch.
… ext3 downloads Move the composition-parts D2H de-interleave onto the GPU (new interleave_ext3_slabs kernel + per-part D2H into the owning buffers), and reinterpret the aux-trace and DEEP-codeword ext3 downloads in place instead of the per-element u64_to_ext3_vec host copy. Byte-identical; comp_parts 3.02->2.63s, aux_commit 1.52->1.29s, deep_fri 2.22->2.0s per epoch.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.