Skip to content

Feat/multi merkle tree gpu resident - #951

Draft
ColoCarletti wants to merge 26 commits into
mainfrom
feat/multi-merkle-tree-gpu-resident
Draft

Feat/multi merkle tree gpu resident#951
ColoCarletti wants to merge 26 commits into
mainfrom
feat/multi-merkle-tree-gpu-resident

Conversation

@ColoCarletti

Copy link
Copy Markdown
Collaborator

No description provided.

Per table the batched prover still ran a redundant HOST LDE FFT (main+aux)
to feed prep trees / retained_main / later phases, leaving the GPU idle.
This completes the device port:
- aux/main: resident device expand first; host-expand only for the
  small-table fallback / Retain. A per-table commit-stream sync closes an
  async race the old host FFT delay had hidden.
- OOD (phase 4): device-recompute device-only + read the trace OOD via the
  GPU barycentric fast path (no host coset eval, no D2H download).

Byte-identical proof (roots unchanged); all batched tests pass. e22 mainnet
prove 284s -> 155s (-45%), GPU util 17% -> 44%. Budget test updated: phase 4
is now a full device expansion.
Route the batched aux build through the resident GPU build (fingerprints +
term columns + running-sum accumulate) and bulk-download the result into the
host aux table, and transpose the main trace to column-major on device
(ResidentMain::HostRowMajor) instead of on the host. Removes the host
set_aux writes, the host accumulate and the ~1.36s/epoch columns_main
transpose. Byte-identical; aux_commit 3.64s -> 1.52s/epoch.
Upload the composition parts to a device handle in the batched deep_codeword
so the DEEP composition takes the fully-resident arm (device parts + device
inv-denoms) instead of the host build_r4_inv_denoms_cpu batch-inverse.
Byte-identical; deep_fri 3.96s -> 2.22s/epoch.
The parts D2H de-interleaved the 3 ext3 slabs into per-column row-major ext3
on the host (a strided gather, ~40% of the download). Do it on device via a
new interleave_ext3_slabs kernel + per-part D2H straight into the owning
buffers (no host copy). Byte-identical; comp_parts 3.02s -> 2.63s/epoch.
… ext3 downloads

Move the composition-parts D2H de-interleave onto the GPU (new
interleave_ext3_slabs kernel + per-part D2H into the owning buffers), and
reinterpret the aux-trace and DEEP-codeword ext3 downloads in place instead of
the per-element u64_to_ext3_vec host copy. Byte-identical; comp_parts
3.02->2.63s, aux_commit 1.52->1.29s, deep_fri 2.22->2.0s per epoch.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant