Skip to content

STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress - #1013

Draft
MauroToscano wants to merge 1312 commits into
mainfrom
noepoch/stark
Draft

MauroToscano wants to merge 1312 commits into
mainfrom
noepoch/stark

Conversation

@MauroToscano

@MauroToscano MauroToscano commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

One STARK proof per block, with no epochs (the prove-and-retire / VADCOP shape). This is a second prover next to #1009's epoch-based one. #1009 stays the reference and this branch does not touch it. The branch starts from #1009's head cc411aa2c, so the diff against main includes #1009. Compare against cc411aa2c to see only this work.

Whole block: 31.78 s (FAST 456 at the head 08ecc4310, mean of 3 default arms), against the epoch tree's 60.15 s (#1009; FAST 389's same-binary reference, not re-run since). Base 21.28 s, recursion 10.25 s (level 0 5.18 · interior 4.88).

Head edddc6873 (10-03): the trace spill store is on by default as auto. A block that fits spills nothing. A block above the host's limit spills the committed traces that would not fit, instead of running out of memory. LAMBDA_VM_BLOCK_SPILL=off keeps every trace. This is prover-only: digest ac406fc8… is unchanged. BIG 117 at the median block: the default spilled 0 B, with phase-A end −0.97 s against off; forced to a 70 GiB target it spilled 30.1 GiB, cost +1.43 s and verified. The same head adds the spill writer's step timings and splits the AIRs out of BLOCK MEM's unnamed heap; both are diagnostics only. See Disk spill: auto by default.

Head a93018974 (10-03): the block tree's sibling proofs now share the card through one VRAM gate, on by default (LAMBDA_VM_SHARED_VRAM_GATE=0 restores the exclusive card permit). This is scheduling only: digest ac406fc8… and top 93aa7097… are unchanged. Against 564dd02bb (FAST 871): 1× recursion −0.68 s, whole −0.48 s. Median recursion −4.72 s (BIG 471, measured before the arming drained the device; the confirmation at this head is queued). See Recursion: sibling proofs share the card.

Head 564dd02bb (10-03): adds LogUp k4 as an opt-in (LAMBDA_VM_ZF_LOGUP=k4; default pair = the same bytes; under k4, median phase B −9.56 s and 1× whole −1.04 s).

Head 8f2ce57d5 (10-03): three block-tree changes, all with the same proof bytes and program ids (FAST 673 merge gate green: 1× top 93aa7097…, pinned ids unchanged; median top 6d85454d…, BIG 586).

  • Late leaf-program emission by default (NOEPOCH_TREE_EMIT_LATE=auto): when leaf 0 puts the other leaves' programs at ≥ 8 GiB, they are emitted in phase B onto pages the retired traces freed; smaller trees emit at the shape as before. Median, spill off (BIG 585): base-phase VmRSS −10.37 GiB, whole-run peak 98.99 → 86.49 GiB, recursion −0.31 s, whole −2.15 s; inert with spill on. NOEPOCH_TREE_EMIT_LATE=off restores the old behaviour, and A/Bs against it must set that.
  • The block verifier streams its derivation (production verify_block_tree): each program and its artifacts are dropped once its child is derived. Verifier alone in a fresh process at the median (BIG 587): peak 32.33 → 22.89 GiB, derive −1.48 s. LAMBDA_VM_BLOCK_DERIVE_HOLD=level restores the old hold.
  • Compact leaf programs on the host: the instruction vector is held at its length (−21 % of a program's held bytes, ids unchanged).

Note on drivers: the whole-block numbers here come from the test harness's tree driver (block_tree_pipeline, cfg(test)). The library has the base prover (block::prove_block), the per-node program emitters (BlockTreePlan::leaf_program / node_program) and the production verifier (verify_block_tree), but no production driver that chains them into a whole-block proof yet; that is a landing item.

Head dacbd5b4e (10-03): KECCAK_RND, ECDAS and KECCAK compose on a bounded-slot interpreter by default (LAMBDA_VM_GPU_INTERP_SI=0 restores the slot-file interpreter). Prover-only: digest ac406fc8… unchanged, FAST 784's landing gate green (1× default − off −0.05 s).

Constraint composition on a bounded-slot interpreter (prover-only; proof bytes unchanged). The three programs too large for the compiled kernels (KECCAK_RND, ECDAS, KECCAK) now compose on a bounded-slot GPU interpreter. The slot-file interpreter it replaces (from OpenVM v1) keeps every value of a row in a per-thread file in global memory: 4,350 words for KECCAK_RND, 2.1 GiB per composition, not counted by the VRAM gate. The new lowering evaluates each constraint's cone on demand in constraint-index order, reads trace cells as operands, and fits a row in 16–128 words of shared memory or a local array, recomputing past the budget; each program gets its measured best shape. On one binary, the three programs compose in 0.58–0.63× the time. At the median block (25475471, 3 + 3) the deciding figure is card work −0.43 s (t −1.9), with the slot-file scratch 80 GiB a run → 0 and the phase-B device peak −1.9 GiB (t −3.4); base −1.03 s is not significant (t −1.0); at 1× base −0.06 s. The digest is unchanged (ac406fc8…); LAMBDA_VM_GPU_INTERP_SI=0 restores the slot-file interpreter. Running all 39 programs on it is slower (1.25–2.1× the compiled kernels, +0.16 s at 1×), so the compiled kernels stay for the other 36.

Head d740eb5d5 (10-03): the block tree emits each node's program as soon as its children land, instead of level by level (NOEPOCH_TREE_NODE_EMIT=level restores the old order). This is scheduling only: the same top program id in all 14 arms.

  • Median recursion −7.32 s (BIG 483: 105.97 → 98.65 s).
  • 1× net −0.15 s (FAST 668).
  • FAST 669 gates are green.
  • Two more knobs ship default off: a leaf-emission window (NOEPOCH_TREE_EMIT_WINDOW; −12.9 GiB median base peak for +6.2 s recursion) and a cap on concurrent derivation builds (LAMBDA_VM_BLOCK_DERIVE_BUILDS).

Head 86e71de77 (10-03). Three landings since f805464f6:

  • aa726ce2b: block fan-in 4 (Mauro approved 10-02; see (f)).
  • 5e4176961: KECCAK and ECSM chunked, and tree partition rule v2, a format change. See the format section below.
  • 7ef261724 + 86e71de77: producer and memory work, all prover-only with identical proof bytes:
    • the compact derived LT ops;
    • smaller LOAD/CPU op records;
    • kept subtrees capped at 8,192 elements per table;
    • the walker levers (hand-out without copy, 2,048-cycle walk batches, in-place routing, sized window lists);
    • eight generator threads instead of six (LAMBDA_VM_BLOCK_GENERATORS=6 restores six);
    • the trace spill store with the S2 policy, off by default (LAMBDA_VM_BLOCK_SPILL=always turns it on).

The median block 25475471 on BIG (job 112, Zen 3 host):

  • Proves and verifies end to end: whole block 303.59 s (base 197.39 + recursion 105.46), against 318.25 s on BIG 480.
  • Whole-run VmRSS 99.07 GiB.
  • 1× digest ac406fc8… unchanged.

Disk spill (BIG 111, base only):

  • It holds the median at 30.0–30.3 GiB against 80–82 GiB without spill, reading back 56.5 GiB over 738 slots with 0 mismatches.
  • It costs +10 s on the base, mostly phase A, because one O_DIRECT writer manages about 1.06 GiB/s on that disk.
  • That is why it is off by default for now.

The walker levers with eight generators end the median's phase A 3.45 s sooner (BIG 468) and cost +0.14 s on the 1× base (FAST 470).

Head f805464f6: two more producer and memory changes, both prover-only with identical proof bytes:

  • 6cb6cd557, smaller per-step records: the CPU op shrinks from 128 to 96 B (arg2, the branch decision and the ECALL kind are derived) and MEMW_A ops become 48-byte aligned rows. FAST 612 at 1×: base −0.38 s, digest equal. BIG 424 at the median: phase-A end −6.0 s.
  • f805464f6: LT ops kept as segments rather than concatenated, and the packed KECCAK_RND / LT finish builds back on by default. BIG 107 picked them over capping wide KECCAK_RND chunks: at the median they save about 14 GiB and end phase A 2.6 s sooner.

BIG 108 at this head: 1× verified at 15.1–16.2 GiB with digest ac406fc8…. The median block verifies at 81.8 / 82.4 GiB with base 212.4 / 216.0 s. Phase B scatters run to run by about 2.7 s (sd) at the median.

Head 330057f2a: the packed KECCAK_RND / LT finish builds are now off by default (LAMBDA_VM_BLOCK_PACKED_BUILD=1 turns them on), and the memory knobs that had no effect are removed. Two one-binary A/Bs at the median (BIG 104, 105): packed builds save 13–15 GiB of host peak but cost phase B ≈ 3–6 s. BIG 105 at this head: 1× verified at 18.32 GiB, digest ac406fc8…; median block 96.2 GiB, base 218–224 s on BIG.

Head 161b82415: six generator threads by default (LAMBDA_VM_BLOCK_GENERATORS=0 keeps the committers generating). With the faster walk the three committers, which also generated every chunk, fell behind: at the median block 32.8 GiB of op lists waited and phase A ended 19 s after the finish. Generators build and pack each chunk; the committers only commit. KECCAK_RND and LT in the finish are built packed a block at a time (LAMBDA_VM_BLOCK_PACKED_BUILD=0 builds them wide). BIG 103/104 on the median block 25475471: base 231.8 → 222–226 s, host peak 96.1 → 82.7–83.0 GiB; 1× base flat; digest ac406fc8… unchanged. Within one binary, packed builds cost phase B ≈ 6 s at the median against wide builds and save 14.75 GiB; see the head above for the current default.

Head 3be24abab: narrow storage on by default (merged from a2cd24207). Main-trace cells are stored at about 2 B each instead of 8 and widened on the card; LAMBDA_VM_BLOCK_NARROW=0 keeps them wide. BIG 100 (A/B on one binary, the 128 GiB box): proof digests packed = wide = ac406fc8…; 1× host peak 31.2 → 17.7–19.0 GiB; 1× base −1.29 s on that box. A median mainnet block (25475471, 9.78× this one, 941 sub-proofs) proves and verifies within 128 GiB: host peak 106.87 GiB, base 259 s on BIG's Zen 3 host (BIG 100; that binary predates the E1+E4 producer changes below). BIG 102 re-gated the merged head: 1× verified, digest ac406fc8…, peak 19.29 GiB.

Head 2c440b8dc: base 19.61 s (FAST 602, mean of 2). It adds two producer changes to 19fe5e9e5, whose memory changes are in Memory: a lean walk (the walker builds only what the tables need) and the executor's guest memory in 64 KiB pages instead of a hash map of words. A B B A on one binary: base 20.38 → 19.61 s (Δ −0.77, no overlap); proof bytes equal under LAMBDA_VM_FIXED_TRACE_HASH=1 + LAMBDA_VM_DETERMINISTIC_GRIND (digest ac406fc8… in both arms); host peak flat. At a median-size block (25475471, harness, windows only) the producer's windows phase falls from 42.4 to 25.2 s and the executor from 108 to 11 ns per cycle. LAMBDA_VM_WALK_LEAN=0 and LAMBDA_VM_EXEC_MEMORY=words restore the old paths. The whole tree has not been re-run at this head.

Review: SOUND-FOR-LANDING at host parity (§F2); open: (d) host attestation posture; parked: (e) the declared-height bound (Mauro 09-25).

⛔ Not mergeable. Open before this can land:

  • (a) A race in the GPU prover: root-cause the R2 H-corruption race so the global serialize lock can become a targeted drain #929 class. CLOSED. It was not GPU prover: root-cause the R2 H-corruption race so the global serialize lock can become a targeted drain #929. The kept-top-levels recompute returned with its trace snapshot still queued, and the LogUp aux build read it on another stream with no wait, so some proofs carried stale aux rows (7 of 47 whole-block proofs were refused, every one caught by the verifier). The fix e80674f6d makes the resident aux build wait on the main snapshot's ready event: 52/52 whole-block proofs verified with it (FAST 432 40/40, FAST 433 12/12), cost −0.05 s, so kept top levels' −4.89 s stands. Note: I-R2RACE.md. FAST 392 at this head: 2 of 2 arms green on the first attempt.
  • (b) Review M1: the tree's programs are not derived on the verifier side. CLOSED (follow-up review R-NOEPOCH-S3.md §F1: "M1 is closed structurally"; its minors are done at a1a3a2c22, gate FAST 395, which this PR takes with its next fast-forward).
    • Per-table shapes now come from the AIR, the trace length and the options (TableChallengeShape::derive, TableVerifyShape::derive).
    • lfm::block_plan::BlockTreePlan::derive runs verify_proof_parts' pre-checks on the claimed shape and builds the host verifier's AIR set. It fixes the partition and the carrier, and derives every leaf, node and top program with no proof.
    • verify_block_tree(elf, shape, output, top) takes no options: it derives under the block presets, so a caller cannot hand it fewer queries (F1-m-a). Both presets stamp the process's ZfFormat, so the verifier's environment enters the derived id; a mismatch refuses, which costs completeness, not soundness.
    • The plan's DECODE and ELF data-page roots are the host's recompute from the ELF (ElfConstants), never the process-wide record the prover's device wrote (F1-m-b; m4 is now fully closed).
    • The partition's cost model is versioned (PARTITION_COST_MODEL = 1, pinned by a test).
    • The epoch tree keeps the gap. The derivation functions are reusable for it.
  • (c) Review minors m1–m5. CLOSED (§F1 closes m1, m2 and m3; m4 is closed by F1-m-b; m5 by the pad-byte negative, FAST 393 and 395).
    • The output halves' canonicity is pinned by the carrier leaf's COMMIT-bus target (emit_output_bytes), carried to every leaf by the nodes' out binding, and checked exactly by top_claims. The block front absorbs only each half's live bytes and does not refuse a non-canonical half; a laptop test shows both of these behaviours.
  • (d) Host attestation posture only. The ELF-derived roots ("A") are program text in the leaves, not supplied-roots cells bound by the attestation.
  • (f) For Mauro: block fan-in 4, a tree-format change. It is −1.0 s on the recursion (FAST 399; fan-in 3 is +1.15 s against 4, FAST 450). It changes every node program and the top. BLOCK_FAN_IN is a verifier-side constant for the block tree only; STARK recursion on GPU (RPX) with ZisK-style proof formats: block 25368371 in 60.18 s #1009's FAN_IN = 2 is untouched. Approved by Mauro 10-02 and landed at aa726ce2b (FAST 661: tree negatives refused, derived top = proved top).
  • (e) Parked (Mauro 09-25): declared-height bound / k-bit PoW before γ, the same as the host verifier. The review's F1-M1: check_shape caps declared trace lengths only at the field's two-adicity, so a prover-chosen height can push a free-height table's DEEP-batching phase below 128 proven bits. This is the host verifier's own exposure, parked on 09-25. The block verifier is at parity with the host.

Review: thoughts/zf/gap2/fix2/R-NOEPOCH-S3.md. The first round read SOUND-FOR-MEASUREMENT (0 blockers, 1 major, 5 minors). §F1 closed M1 and m1–m5. §F2 reads SOUND-FOR-LANDING at host parity on a1a3a2c22, and it confirms the pad-byte correction: the refusal is emit_output_bytes in the carrier leaf, and out is not redundant.

The whole block, epoch tree vs no-epoch tree (FAST 389, one binary at e4ff3f8fb, means of 2 arms)

Block 25368371. Both trees run on one test binary, and its md5 is checked after every arm. The epoch arm is #1009's record harness, unchanged.

no-epoch (this PR) epoch (#1009) Δ
base 24.91 s 26.9 s
through the last interior node 39.01 s 60.15 s −21.14 s
top node (closes the bus) / block-artifact root, on a line of its own 1.67 s 1.70 s (+ 0.3 s verify)
complete 40.67 s 61.85 s −21.18 s
  • The epoch WHOLE of 60.15 s reproduces STARK recursion on GPU (RPX) with ZisK-style proof formats: block 25368371 in 60.18 s #1009's 60.18 s, so this binary's epoch tree is STARK recursion on GPU (RPX) with ZisK-style proof formats: block 25368371 in 60.18 s #1009's. If the root replaced the top interior node (option A; an estimate, not measured), the epoch would complete at about 60.35 s, a Δ of −19.68 s. Pre-registered row: no-epoch whole − epoch WHOLE = −19.48 s, band [−24.5, −17.5], IN.
  • The no-epoch recursion is 15.51 s (pre-registered band [12.5, 16.5], IN): harvest 2.03 s, then level 0 with 8 leaves at 6.25 s, then the interior at 7.22 s (3.41 + 2.15 + 1.67).
  • Each harness verifies on the host what it reads. That verify is not work a production driver does, so both trees overlap it with proving: STARK recursion on GPU (RPX) with ZisK-style proof formats: block 25368371 in 60.18 s #1009's per-epoch verifies run beside its wraps, and the block's 6.8 s verify of the base runs beside level 0. The join waited 0.00 s.
  • Levers measured in that run: harvest verify beside level 0 −3.21 s (EFFECTIVE); level-0 siblings 8 −0.13 s (no effect).
  • The recursion: one leaf per instance list, each leaf replaying the statement and Phase A over all 137 roots; nodes that check one state digest, one claim and summed bus shares; a top node that asserts the bus closes. Since a0954fd1b the lists are the plan's: D-NOEPOCH §12.2's rule over the closed-form costs, a pure function of the block's shape. On this block the rule gives other lists than §12.2's, which came from legmodel.py's costs, but the leaf tiers are the same. FAST 392 at cda620461 (means of 2): level 0 6.21 s, interior 7.19 s, harvest 2.03 s, whole 40.38 s, which is 389's numbers. The verifier's own derivation of the top program takes 8.35 s; that is outside the whole, and the derived top equals the proved top. A tree over another partition, a skipped instance or a duplicated instance is refused at the final check (fixture gate).

The plan's gate and the harvest lever (block 25368371, means of 2 arms, 389's posture)

FAST 389 B2 (§12.2's lists) FAST 392 cda620461 (the plan) FAST 394 a684f4f3f (+ ELF constants beside the base) FAST 395 a1a3a2c22 (+ F1 follow-ups) FAST 396 e7406d6b5 (+ rebuilt windowed builder) FAST 397 6b86432a2 (+ verifier derive)
base 24.91 s 24.70 s 25.02 s 25.05 s 21.77 s 21.86 s
harvest 2.03 s 2.03 s 0.19 s 0.19 s 0.18 s 0.19 s
level 0 (8 leaves) 6.25 s 6.21 s 6.24 s 6.23 s 6.31 s 6.25 s
interior 7.22 s 7.19 s 7.21 s 7.20 s 7.22 s 7.23 s
whole 40.67 s 40.38 s 38.89 s 38.91 s 35.73 s 35.78 s
verifier derivation (verify_block_tree, no proof read, outside the whole) n/a 8.35 s 8.28 s 10.5 s (the data-page roots are now host recomputes) 10.55 s 8.59 s cold, 5.50 s warm (ELF constants cached)
  • Harvest lever (FAST 393, A/B on one binary): EFFECTIVE, −1.80 s on the whole. The plan's ELF-only constants are computed on 4 threads while the base proves, and the harvest joins them. The harvest's plan step fell from 2.02 to 0.16 s, and the base moved +0.02 s. At 395 the constants, including the data pages, take 10.9 s beside a 25 s base, and the harvest waits 0.00 s for them.
  • FAST 396's e7406d6b5 and FAST 397's 6b86432a2 are both on this branch; 6b86432a2 is the head.
  • Rebuilt windowed builder (FAST 396, A/B on one binary): EFFECTIVE, base 21.55 s. The WHIR lane's builder (walker and accumulator split, segmented windows, BITWISE counted per window), driven the way the WHIR block drives it: a walker thread only walks, the producer routes, the committers generate each chunk before its Round-1 commit. Interleaved means: new 21.60 s, the previous schedule over the same builder 23.97 s, serial phase A 25.71 s. The 10 default runs average 21.55 s, against 24.52–25.17 s at FAST 395. The whole tree is 35.73 s, against 38.91 s. The builder came in as 7 cherry-picks (-x) of noepoch/windowed-builder 6e99c74..1d24a2c, not a merge of that branch (90 WHIR commits, 6 conflicting files), plus our parallel gating, which compiles the recursion ELFs (5fe8316, now byte-identical to the builder branch's bcccdef).
  • Verifier derive (FAST 397): cold 10.55 → 8.59 s, warm 5.50 s. verify_block_tree_with(elf, &ElfConstants, …) lets a consumer compute the per-ELF constants once (3.05 s cold), and each node level is now emitted and built in parallel (level 1: 1.31 s wall for 5.0 s of work). Split at 397: constants 3.05 · plan 0.16 · leaves 2.17 · L1 1.31 · L2 0.93 · L3 0.72 · check 0.24.
  • Recursion levers (on noepoch/stark-s3 at 3992a56f9, gated by FAST 399; A/B on one binary each):
    • Child proofs verified beside the timed path: −0.62 s, now the default. Every verify is still joined, and a refusal fails the run before anything is reported.
    • Block fan-in 4: −1.02 s, see (f).
    • The tree's programs derived beside the base with host-built artifacts: a regression, because host builds are about 100× slower than the card's. With everything ready, the mechanism is −3.8 s.
    • The pipeline instead (FAST 451: −2.56 s on the recursion, whole block 32.30 s), the default since f6445c8b6 (FAST 452, gated: −2.48 s against NOEPOCH_TREE_AHEAD=0, P8 10/10). The leaf programs are emitted beside the base. Leaf and node artifacts are built on the card during level 0, and each node program is emitted as soon as its children's artifacts exist. Base +0.01 s; every child proof verified.
    • Builder first at the card permit (FAST 453, c9c29cb5b, NOEPOCH_BUILDER_FIRST, off): recursion −0.60 s in its A/B, but not by the mechanism pre-registered. It cannot pre-empt a leaf's multi_prove, so the level-1 artifacts are still late; what it did was remove level 0's slow mode (≈ −0.4 s pooled over 451–453). Not the default.
    • Node programs emitted on the builder's own pool (FAST 455, 7da2a45d0; the default since 08ecc4310, gated by FAST 456 and 457): recursion −0.29 s, level 0 −0.40 s. Level 0 was bimodal (5.2 or 5.8 s). In the slow runs, one leaf's multi_prove held the card for 1.2 s at about a third busy, starting at the instant the builder began emitting the level-1 programs on the global rayon pool (FAST 454's card trace). With a 4-thread host-only pool of its own, the stalled hold is gone in every run. NOEPOCH_EMIT_POOL=0 restores the global pool.
    • The shared windowed builder's parallel window concatenation (i-noepoch-w's 9d15e7b, cherry-picked as e029731f7; FAST 456): base −0.57 s (P8, 5 against 5 on one binary, no overlap), replicating WHIR's FAST 417 (−0.52 s). LAMBDA_VM_BUILDER_CONCAT=serial restores the old copy.
    • Not landed (recorded on the lane's branches): programs handed out before their artifacts with per-node emission (FAST 457, NOEPOCH_TREE_EAGER): recursion −0.06 s, because level 1 is limited by the card, not by waiting. The card trace (FAST 454) puts the idle inside the recursion's card holds at 1.2–1.8 s, 13–19 %.
    • Whole block: 31.78 s at the default (FAST 456); 32.55 s at f6445c8b6 (FAST 452); 35.03 s with NOEPOCH_TREE_AHEAD=0. Fan-in 4 was measured on the inline path (33.82–33.95 s there).
  • Whole-block verifies: FAST 452, 455, 456 and 457 each ran 10 of 10 P8 base proofs through production's block verifier, all verified (456 alternating the parallel and serial concatenation), and every child proof of every tree run was verified. FAST 396 and 397 each ran 10 of 10 P8 base proofs at their shas through production's block verifier, all verified (396 also verified its 10 A/B arm proofs). FAST 395 ran 10 of 10 P8 base proofs at a1a3a2c22 through production's block verifier, all verified (bases 24.52–25.17 s), so the race fix (a) is shown on the plan's code.
  • Gates are green at every sha, every count exact: the host gates, 25 epoch-reader tests, and the fixture leaves, tree, verifier and pad-byte tests.
  • The verifier's derivation is a consumer cost, paid once per (ELF, block shape) rather than once per ELF: the top program depends on the block's table counts, page ranges and trace lengths. The per-ELF part (ElfConstants) and the per-(AIR, length) shapes can be reused across blocks.

Memory: toward typical blocks (head 19fe5e9e5)

A median mainnet block (25475471) is 9.78× this one, and the host peak grows about linearly with the block. The head carries three changes for that; none changes a proof byte.

change gate 1× base Δ 1× host peak larger block
Each streamed chunk's op lists are freed once the chunk is committed (drop_streamed_ops, default on; LAMBDA_VM_BLOCK_DROP_OPS=0 keeps them) FAST 502, EFFECTIVE −1.07 s 41.9 → 33.0 GiB 1.20× verifies at 39.7 GiB
Kept top levels 3 → 6 (BLOCK_RECOMMIT_TOP_LEVELS) FAST 503, EFFECTIVE −0.05 s 32.8 → 31.7 GiB 1.57× proves at 47.2 GiB
One trace-hash state per process for the six hash-ordered tables (LAMBDA_VM_FIXED_TRACE_HASH=1) FAST 504, EFFECTIVE — 31.2 / 31.6 GiB —
  • Freeing the op lists builds the whole-run build's tables, table for table: 138 tables, 55 streamed chunks, 0 differ (2 of 2 runs). Proof size is unchanged (82,562,392 B).
  • With LAMBDA_VM_FIXED_TRACE_HASH=1 and LAMBDA_VM_DETERMINISTIC_GRIND, two processes produce the same block proof (digest ac406fc8…); a random-key control differs. Byte-identity gates on the block use this.
  • Next, on its own branch until its default-flip gate: narrow column storage (about 2 B per main cell instead of 8). FAST 508: 1× peak 19.1 GiB at +0.28 s base; a 4.13× block proves and verifies at 55.5 GiB.

What it is

  • The whole block is proven as one monolithic VmProof. Every table is cut into instances of 2^21 rows, and KECCAK_RND into 2^16-row instances. There is no L2G table and no global proof; memory uses the monolithic PAGE argument.
  • ResidencyMode::RecomputeLdeDevice: after Round 1 only each instance's root is kept. Each instance's fused task commits its trace on the device again, and the prover refuses the proof unless the new root equals the absorbed one, so the proof bytes are unchanged. Retain and RecomputeLde behave as before.
  • prover::block::prove_block / verify_block. The block verifier is the only one that accepts a chunked KECCAK_RND (AcceleratorShape::KeccakRndChunked). Every other verifier keeps KECCAK_RND to one table.
  • Knobs, all off by default: LAMBDA_VM_GATE_PACKING (VRAM-gate packing admission), LAMBDA_VM_TABLE_TIMELINE (per-table timeline), LAMBDA_VM_RECOMMIT_TOP_LEVELS (kept top levels).
  • Phase A is streamed by default. The executor runs in 2^21-cycle windows that feed the shared WindowedTraceBuilder, and every full chunk of CPU, MEMW_R, MEMW_A, MEMW, LOAD, LT, SHIFT and STORE is committed as soon as it exists. Host tests show it equals the serial build.
  • Kept top levels (k = 6 since 0fe8e4f2f; k = 3 before) are prove_block's default. The ELF data pages' preprocessed roots are computed on the device during execution.

Base, block 25368371 on FAST (A/B on one binary; the epoch base is #1009's prove_continuation on the same binary)

step base Δ
epoch base (#1009 reference) 26.05 s
S1 block proof, first measurement 38.59 s
+ packing admission (k = 4) 37.17 s −0.64 (no effect)
+ k = 8 35.6 s (exploratory)
+ KECCAK_RND chunked at 2^16 32.40 s −3.23
+ ELF data-page roots on the device during execution 30.88 s −1.43
+ streamed phase A (windowed execution, CPU instances committed per window) 27.98 s −2.76
+ kept top levels instead of the second hash (k = 3) 23.14 s −4.89
+ shared windowed builder streaming 8 tables 24.89 s +1.75 against the row above (another run); −0.84 against the serial producer on one binary, where −2.0 or better was pre-registered
+ the rebuilt builder, its walk on a thread of its own, the committers generating (e7406d6b5, FAST 396) 21.55 s −2.37 against the previous schedule over the same builder on one binary; −4.11 against the serial producer

At e7406d6b5 the block base is about 4.5 s below the epoch base (21.55 against 26.05). The first shared builder cut Round 1's span from 6.3 to 2.4 s, but its window builds sat on the executor's path (execute 1.35 → 6.9 s). With the rebuilt builder and the walk on its own thread, execute is 3.35 s and phase A 7.13 s, and the prove is 14.4 s against the serial build's 18.4 s. Nsight on FAST shows the fused phase is 98.6 % card-busy, so the block is card-bound. Most of that card time was the second hash, which kept top levels remove: phase B recomputes the LDE alone, and the openings rebuild each queried 8-leaf subtree and check it against the kept node. The architecture's main gain is in the recursion (above).

S0 census: main 3.301 G elements, aux 1.036 G.

Format change: ECDAS chunked; KECCAK_RND and ECDAS heights capped

Landed at 23b3c8173 (FAST 531 gates GREEN; FAST 533 A/B on 25512221: base −0.17 s, no effect on time, as expected; ECDAS cells −25 %, host peak −0.40 GiB).
ECDAS is one row per double/add step (≈ 382 per ECSM call, ≈ 4.2 calls per transaction), so a median mainnet block
makes ≈ 420 k rows: one table at 2^19 proves 127.91 bits at DEEP batching on a ≈ 22.5 GiB device set, and a p90 block's
2^20 (126.91 bits, 44.7 GiB) no longer fits a 32 GiB card. The block now cuts ECDAS into instances of at most 2^17 rows;
a scalar multiplication may straddle two instances, its steps chaining only through the Ecdas bus, keyed by the call's
timestamp and the step's (round, op). AcceleratorShape::KeccakRndChunked is renamed BlockChunked and lifts the
one-table bound for ECDAS as for KECCAK_RND (count bounded by the sub-proof cross-check). New verifier constants, checked
in verify_block and in the tree's check_shape: every KECCAK_RND instance ≤ 2^16 rows (129.43 bits) and every ECDAS
instance ≤ 2^17 (129.91 bits). This closes G1 (REV-JUDGE item 13, R-NOEPOCH-S3 F1-M1) for KECCAK_RND and ECDAS
only
; every other table's height is still bounded by two-adicity alone, and the rest of G1's per-type list (ECSM,
KECCAK, the CPU family, the fixed tables) stays parked. LAMBDA_VM_BLOCK_KECCAK_RND_LOG2 takes 5..=16 (no off).
A block whose ECDAS fits one 2^17 table (25368371: 2^16) builds the same ECDAS table as before (one chunk of every step is the table generate_optional built; by construction, not byte-compared on a block); 25512221 (2^18 today, 128.91 bits, 0.04
under the minimum of record) now proves 2^17 + 2^16. The epoch and recursion verifiers (Single) are unchanged.

Format change: KECCAK and ECSM chunked; tree partition rule v2 (5e4176961)

Landed with the any-block target. Approved by the lead; Mauro to confirm. It is listed for the cryptography review as S-6 and S-9.

The caps. KECCAK is cut into instances of at most 2^18 rows (129.213 bits) and ECSM into instances of at most 2^17 (129.488 bits), as ECDAS is.

  • A row of either is one whole call (a permutation, a scalar multiplication), so a cut falls between calls.
  • A call reaches its rounds, steps and memory only through buses keyed by its timestamp.
  • The block verifier and the tree check every instance against these caps.
  • LAMBDA_VM_BLOCK_KECCAK_LOG2 (2..=18) and LAMBDA_VM_BLOCK_ECSM_LOG2 (2..=17) lower the caps, for tests.

Partition rule v2 (PARTITION_COST_MODEL = 2). Rule v1 pinned every chunk of a table to one leaf. Now the first instance keeps its seeded leaf and later chunks of KECCAK, ECSM and ECDAS fill by load. Without this, a p99 block's ≈ 15 ECDAS chunks overflowed the leaf cap and the plan was refused. The verifier derives the partition by the rule from the shape alone.

Evidence:

  • FAST 723: 1× digest and tree top unchanged; a forced 3 KECCAK + 4 ECSM chunks verified host and tree; 5 mutations caught.
  • BIG 520: the median base and tree verified, derived top = proved top.
  • Every block measured so far keeps one KECCAK and one ECSM table, and its proof bytes.

Known cost (deferred): a table just over a power of two is cut after padding, so for example 2^20 + 1 KECCAK calls make 8 instances, about half of them padding. This matters only on keccak-heavy blocks.

LogUp: four interactions per aux column (opt-in, LAMBDA_VM_ZF_LOGUP=k4)

Default pair = today's bytes. LAMBDA_VM_ZF_LOGUP=k4 lets each base table commit four bus interactions per LogUp aux column instead of two, where that commits fewer extension columns: groups of four have degree 5, which blowup 4 admits, at the price of four composition parts instead of two. The rule is per table and verifier-side (⌈N/k⌉ aux columns + parts, ties keep pairs): KECCAK_RND 516 + 2 → 258 + 4, ECSM 290 → 145, ECDAS 194 → 97, CPU 10 + 2 → 5 + 4, MEMW_A 10 → 5; LT, STORE, MEMW_R, LOAD, PAGE and the small tables keep pairs. The LFM chips keep pairs; #1014 is untouched.

  • Code: ProofFormat.logup (stark), the group/accumulator emitters for k ≥ 3 (one body for the prover folder, the verifier folder and the IR capture), the host and device aux builds grouped by arity, the four-part composition split on the card (radix-2 twice) and its host mirror, compiled kernels for the k4 table programs, the knob at block_base_options, and one new verifier refusal (parts > blowup).
  • Bytes: with the knob unset or =pair the 1× digest is ac406fc8 and every program id, kernel key and golden is unchanged (the small-tree id pin and the 44-program golden pass on the landing merge).
  • Measured (one binary each): median block 25475471 on BIG: phase B −9.56 s (t −27, 3 + 3), host peak −1.06 GiB, proof −4.6 %; about a third is stage work (aux commit −11.5 worker-s against composition commit +7.1) and two thirds packing under the 24,000 MB VRAM gate (smaller per-table device sets let three large tables run where two did). 1× block on FAST: whole −1.04 s, base −0.97 s (t −15), almost all packing.
  • Correctness: the card's split equals the host mirror (2^8…2^21) and a device-only k4 proof equals the host proof byte for byte; k4 verifies on the host and through the block tree; six tampered-proof negatives refused.
  • Security: CRYPTO-REVIEW S-11. LogUp soundness is unchanged (it counts interactions × rows); every FRI phase stays ≥ 128.946 bits (batching rises as L falls); DEEP 161.18 bits per table at degree 5.
  • Landed at 564dd02 (landing gate FAST 863: pair bytes and tree top unchanged, k4 verifies, six negatives refused).
  • Opt-in until Mauro rules on the default.

Recursion: sibling proofs share the card through one VRAM gate (on by default)

Before: the recursion's card permit was a mutex, so one proof at a time ran inside multi_prove. The card then sat idle while the holder ran its host stages: uploads, absorbs, queries. At the median block that was 24.6 s of card idle inside the holds (BIG 469).

Now: every multi_prove in the tree admits its tables through one process-wide byte gate, and the artifact commit takes its bytes from the same gate. Sibling proofs overlap wherever their bytes fit. Three parts keep the gate's account matching the card:

  • Resident bytes stay counted. A Retain prove keeps each table's main LDE, trace snapshot and tree on the card from its Round-1 commit until its fused task ends. Those bytes stay in the gate the whole time: the Round-1 task carries them past its own permit, and the fused task takes them over and is admitted only for the rest of its set.
  • Claims make the carry deadlock-free. Each prove first claims room for its residents plus its largest table's set. The claim settles after Round 1 and shrinks as tables finish, and claims are admitted only while their sum fits the budget. A prove whose claim alone exceeds the budget runs alone.
  • The pool posture and the calibration.
    • The device pool releases freed blocks at each sync only while the gate is armed; the base keeps its retained pool.
    • Arming drains the device, trims the pool, and sets the budget to the card's free memory minus 3.5 GiB, capped at the configured 23.44 GiB.
    • At most three proofs are inside multi_prove at once.

LAMBDA_VM_SHARED_VRAM_GATE=0 restores the exclusive permit. LAMBDA_VM_SHARED_GATE_TRACE=1 prints the gate's account (SGATE lines). The gate acts only while the tree arms it for concurrent proofs, so the base and every single-proof path are unchanged.

Bytes: unchanged, since the gate only reorders admission. The digest is ac406fc8…, and every A/B arm has the same top program id.

Measured:

  • Landing gate, FAST 871 (two binaries, a93018974 against 564dd02bb, 4 + 4):
    • recursion −0.68 s (t −9.6); whole −0.48 s (t −2.9); base +0.09 s (t +0.8);
    • VRAM peak 22.3–26.4 GiB against 24.1–27.3; level-0 budget 23.44 GiB in every gate-on run.
  • One binary, gate on against off, FAST 479 (8 + 8): recursion −0.59 s (t −17.2); whole −0.51 s (t −6.9); base −0.02 s.
  • Median block 25475471, BIG 471 (3 + 3, measured before the arming drained the device):
    • recursion −4.72 s (t −7.9); whole −4.29 s; level 0 −3.42 s; interior −1.30 s;
    • VRAM 25.0–26.1 GiB against 26.5–26.7; host peak at level 0 ≤ 97.9 GiB.
    • It missed the pre-registered −5 s. Without the drain, the level-0 budget (16.9–19.5 GiB) sometimes fit only one ≈ 9.7 GiB leaf claim at a time; the drain lifts it to the 23.44 GiB cap.
    • The median confirmation at this head is queued (BIG 472).
  • Interior node claims (23–24 GiB) run one at a time at any budget up to the cap. Pairing them needs tighter device-set estimates, which over-count by a median of 2–14 GiB.

Readout note: FAST 871's "claims in force" row read OUT because it counted from the optional trace, which that gate ran without. From each claim's own log line, every gate-on run peaked at 3 claims and the off runs had none.

Tests:

  • the gate's exact account at every fused start, with one driver;
  • the claim arithmetic;
  • two carrying proves that deadlock without claims and finish with them (each test fails under its mutation);
  • the permit tests pin the exclusive card where they test it.

Disk spill: auto by default

What it does. Once Round 1 has committed an instance, its packed main trace can go to a spill file that phase B reads back ahead of its walks. The words that come back are the words that went out, so no proof byte depends on the policy.

  • The file is unnamed (O_TMPFILE) and refuses a tmpfs directory. It uses O_DIRECT where a probe write allows it.
  • One writer thread works behind a 2 GiB queue. The writer takes each slot's digest, and every read checks it.
  • A store that fails to write keeps its traces resident.
  • LAMBDA_VM_BLOCK_SPILL = auto (default) | off | always | <GiB> (a resident budget for committed packed traces).

How auto decides (prover/src/block.rs, spill_target_bytes / spill_decision). A committed instance is spilled when the host's bytes, plus the reserve, plus the instance's own bytes would pass the target.

  • The host's bytes are the larger of the process's VmHWM and its cgroup's memory.current. The latter counts the page cache.
  • The reserve is 0.18 GiB per G committed cells plus 6 GiB, for the finish's transients and phase B's bump.
  • The target comes from the machine, in this order:
    1. LAMBDA_VM_BLOCK_SPILL_TARGET_GIB, if set;
    2. else the smaller of the cgroup's limit and MemTotal, less 10 GiB. The limit is v2 memory.max, or v1 memory.limit_in_bytes, read at the process's cgroup path and then at the hierarchy root, which is what a container without a cgroup namespace sees. Since 278e6a8c6: FAST 509 read 47.5 GiB on FAST's v1 limit, where the v2-only rule read 49.9;
    3. else no spill (no /proc).
  • On BIG that is 120.69 − 10 = 110.69 GiB. A 57.5 GiB host gets ≈ 47.5.
  • While a store is open, the stream's queues have byte budgets (ops 4 GiB, generated chunks 2 GiB). They keep work from piling up behind a slow writer. Under auto they apply even when nothing spills. At the median that cost nothing: generated chunks waited 39–61 generator-s and ops 2.0–2.5 s, and phase-A end was −0.97 s against off.

Measured, median block 25475471 on BIG (3 runs per arm, one binary per job):

job arm spilled phase-A end vs no spill host peak
BIG 117 auto (default) 0 B in every run −0.97 s 81.9–82.8 GiB (off 80.8–82.7)
BIG 117 auto, 70 GiB target 30.1 GiB, 427–429 slots +1.43 s 60.0–62.4 GiB
BIG 116 always 56.5 GiB, 738 slots +7.07 s 31.8–32.5 GiB
BIG 115 always 56.5 GiB +6.03 s 31.6–32.7 GiB
BIG 111 always (first measurement) 56.5 GiB +6.8 s (base +10 s) 30.0–30.3 GiB
  • Every spilled slot was read back with 0 mismatches.
  • The 1× digest ac406fc8… holds under the default (BIG 117) and under always (BIG 114–116).
  • Phase B moved within ±1 s.

A p75 block proves end to end under the default: 25440371 (39.2 M gas, 447.6 M cycles, 14.67× the bench block) on BIG: base 285.5 s at 99.7 GiB with 29.3 GiB spilled; whole block 421.0 s (level 0 71.6 s, interior 61.7 s) at 108.0 GiB, the peak being phase B's end with the late-emitted leaf programs. With the spill off the base still fit, at 116.1 GiB, by squeezing the page cache from 12 to 4 GiB (BIG 118).

What the cost follows. While the writer runs, the generators slow by ≈ 16 %. The writer's two passes over every spilled byte (digest, then the aligned copy) and the frees of the spilled buffers both contribute; BIG 115 and 116 could not split them further. At the median, phase A is bound by the generators during the walk, so the walk waits on them.

  • The 70 GiB target started spilling at t ≈ 33 s. It handed ≈ 13–15 GiB to the writer before the walk ended and the rest after, for +1.43 s.
  • always hands ≈ 39 GiB to the writer during the walk, for +6–7 s.
  • Rule of thumb (inferred): ≈ 0.1–0.18 s of phase A per GiB spilled during the walk, ≈ 0 for what is spilled after it.
  • The writer manages 0.85–1.06 GiB/s. Its seconds by step under always: digest 22, aligned copy 28, pwrite 14. The pwrite runs at the disk's own rate (3.7 GiB/s raw). The BLOCK SPILL line prints the split.

Measured and not landed (each has a FAILED-LEVERS row):

  • two writer threads (BIG 113: no effect);
  • queue budgets that bind only while the writer is behind (BIG 114: phase A +9.4 s, peak 37 GiB);
  • F1, trace buffers in page-aligned memory written straight from their pages with the digest fused per chunk (BIG 116). The copy was gone, but phase A moved only −0.9 s against the heap writer, at the price of an munmap per slot.

Tests:

  • a spilled trace round-trips bit for bit, O_DIRECT and buffered;
  • every trace spilled proves the resident bytes under every residency;
  • a flipped byte on disk, or a perturbed slot, is refused, on the card before any device work;
  • a failed store proves the resident bytes;
  • the read-back keeps to its window;
  • the stream under always, a zero budget and auto builds the resident traces;
  • the policy and decision tests; the writer's step seconds add up.

Level 0 opened with host work alone: every first-round wrap's prologue at
once (reconstruct, emit, arenas, and on WHIR the epoch harvest), with
nothing on the card. None of it needs the card or the global proof, only the
ELF and the epoch proofs the base finished long before.

By default the tree drivers now start a lead-in before the base. Its helpers
(two by default) wait until the base reports its epoch count, then build the
prologues of wraps 0..want (want = level 0's first pool round) from copies of
the leading epoch proofs, with the same functions the pool calls, so the
programs are the pool's own. Level 0 takes each prologue instead of building
it. A prologue no helper started is built by the pool as before, and a
panicking one is handed back. Nothing in the lead-in takes the card permit
or holds device memory of its own. It uses the base's DECODE derivations as
the base shares them: the STARK commitment, and on WHIR the root and the
prepared opening, whose derivation is a device commit.

The base reports through an EpochObserver installed for the calling thread
(with_epoch_observer): the epoch count from the producer as soon as the
final epoch is executed, each proved epoch, and the DECODE work. That holds
on both preparation schedules and both pipelines; with no observer
installed, the pipeline is unchanged. Level 0 still takes its own DECODE
derivations from the base, and LFM_TREE_REDERIVE_DECODE=1 still re-derives
them there; the lead-in's copies are only for its prologues. BaseDecode now
holds the Arc the base shares. An I4 schedule test reads epoch positions
through the slice-based epoch_chain_position.

LFM_TREE_PROLOGUES_AT_LEVEL0=1 builds the prologues at level 0's start
instead; LFM_TREE_TAIL_PROLOGUES and LFM_TREE_TAIL_HELPERS size the lead-in.
The driver prints `L0 PROLOGUES: ...` either way, and at level 0's start how
many prologues were ready.

Measured on block 25368371 (FAST, RTX 5090), ABBA palindromes at 5cbdf06,
on 169b668 behind a temporary knob, before the base-prep and DECODE-handoff
changes. WHIR: 107.60 -> 104.40 s (-3.20, A spread 1.20); the lead-in
3.19 -> 0.07 s; level 0 -3.1 s; base unchanged; device peak +592 MiB, from
the harvest's MLE evaluations running unreserved beside the base's tail.
STARK: 121.30 -> 118.65 s (-2.65); the lead-in 5.54 -> 0.22 s and level 0
-6.1 s, but the base +3.4 s, because the two helpers slow the base's epoch
proofs. Program identities unchanged. On this base the DECODE handoff
already removes part of the lead-in, so the gain here is smaller than above.

Tests. Card-free: the hand-off (order, the count gate, handing back an
unstarted, failed or context-failed prologue, waiting on one in progress,
close). Fixture scale, on both bases: the observer sees the count once and
every epoch byte for byte, and a prologue built from the leading epochs emits
the pool's own program and arenas.
…te does

Every staged_transfers test forces its path with the thread override, and
only the block-scale tree drivers read the lead-in's setting, so no card-free
test read either setting from the environment. One test each now does, and
prints the line that names it:
- staged_transfers: staging_pairs_enabled() is !LAMBDA_VM_STAGING_SHARED_SLAB,
  with its `[gpu] transfer staging: ...` line;
- the tree tests: lead_in_enabled() is !LFM_TREE_PROLOGUES_AT_LEVEL0, with an
  `L0 PROLOGUES setting: ...` line.
The gate runs each with --nocapture under the default and under the opt-out
and counts the named lines.
…nned staging, level 0's first wrap prologues built in the base's tail

I6: each row-major commit transfer is staged through a pinned pair of its
own instead of the worker's shared slab; opt-out
LAMBDA_VM_STAGING_SHARED_SLAB=1, named once on stderr ([gpu] transfer
staging: ...). I7: level 0's first wrap prologues are built by helpers in
the base's tail; opt-out LFM_TREE_PROLOGUES_AT_LEVEL0=1, sized by
LFM_TREE_TAIL_PROLOGUES / LFM_TREE_TAIL_HELPERS (L0 PROLOGUES: ...). No
proof byte moves. Measured on the lane's base (IDLE-B box2): I6 STARK
-5.80 s, WHIR -2.10 s; I7 WHIR -3.20 s, STARK -2.65 s net.
…xes K3/K4/K5

Brings the HASH lane's six commits onto candidate C2 (7d41668): the
half-warp Merkle tops (K3), the work-queue grind (K4), the limb-multiply
permutation variants (K5), each still behind its LAMBDA_VM_GAP_* knob,
their parity tests and host KATs, the serialised grind-counter tests and
the queue grid's context fix. No conflict; the next commits make the
three fixes the defaults.
…imb permutation by default

The three RPX device fixes measured EFFECTIVE on the WHIR block (job 160,
wt300-307: K3 -3.90 s, K4 -6.45 s, K5 variant 5 -10.05 s against 107.15 s,
identities identical in every arm) are now the defaults. Each keeps an
opt-out, and its default lives in one constant in `rpx_paths`, so a
pipeline that needs one off flips one line:

- LAMBDA_VM_RPX_WARP_MERKLE (WARP_MERKLE_DEFAULT): narrow levels and the
  tail on rpx_merkle_level_warp / rpx_merkle_tail_warp; =0 walks a
  thread per parent (rpx_merkle_level, rpx_merkle_tail).
- LAMBDA_VM_RPX_GRIND_QUEUE (GRIND_QUEUE_DEFAULT): rpx_grind_search_queue
  on a card-filling grid; =0 is rpx_grind_search on LAMBDA_VM_GRIND_GRID.
- LAMBDA_VM_RPX_LIMB_PERMUTE (LIMB_PERMUTE_DEFAULT): every RPX kernel from
  rpx_v5.cubin (32-bit limb multiply, square_n unrolled by four); =0
  loads rpx_v0.cubin, the 64-bit multiply.

Each variable takes 0 or 1 (anything else aborts) and each switch prints
one line on first use, "[gpu] RPX Merkle: ...", "[gpu] RPX grind: ...",
"[gpu] RPX permutation: ...", naming the path and whether it came from
the pipeline default or the variable. The queue's grid line becomes
"[gpu] RPX grind queue: grid ...". build.rs now builds rpx.cu twice
(variants 0 and 5) instead of five times.

The LAMBDA_VM_GAP_* knobs and the gap_hash module are gone. The parity
tests move to prover/tests/rpx_device_paths.rs, named for what they
compare (the per-parent walk, the stride grind), and the temporary
wording leaves the kernels, the host KATs and the docs.
The PROVE SPLIT line's "r4_grind (n/airs on device)" took its delta of
gpu_lde::gpu_grind_calls(), the keccak arm's counter. An RPX grind that
runs on the device counts in gpu_grind_calls_rpx(), so under RPX every
line read 0/airs while every table ground on the card: on the WHIR block
(job 160) the root proof printed 0/11 beside the harness's own count of
11 RPX device grinds for it.

device_grinds_now() sums both arms and report() takes its delta of that.
rpx_grind_device gains a test that reads it around one RPX device grind
(the keccak-only count reads 0 there and fails).
The result lines still carried the campaign's fix ids (K3, K4, K5) and
called the old paths "shipped", which stopped meaning anything once the
new paths became the defaults. They now say what is compared: the warp
walk against the per-parent walk, the queue grind against the stride
grind, the limb primitives and the permutation variants. Output only;
every check is unchanged and both binaries still pass.
…er half-warp, a queue grind, the limb permutation

K3: a Merkle level and the tail compress one permutation per half-warp
(opt-out LAMBDA_VM_RPX_WARP_MERKLE=0). K4: the device grind claims nonces
from a work queue (LAMBDA_VM_RPX_GRIND_QUEUE=0). K5: the whole RPX module
runs the limb-multiply permutation variant (LAMBDA_VM_RPX_LIMB_PERMUTE=0).
Each opt-out accepts 0 or 1 and names itself once on first use. Every
digest, root and grind nonce search is byte-identical to the previous
kernels (host known-answer tests and device parity). The prove split now
counts RPX device grinds. Measured on the pre-K1 base: WHIR K3 -3.90 s,
K4 -6.45 s, K5 -10.05 s; STARK K3 -1.00 s, K4 -0.50 s, K5 -7.95 s.
…n C3S

C3 = C2 (7d41668) + gap-fix/idle-b-int (bdb2d37, per-transfer pinned
staging and level 0's first wrap prologues built in the base's tail) +
gap-fix/hash-int (e041e9f, RPX Merkle tops per half-warp, a queue grind
and the limb permutation), each a signed merge.

per-table-gpu's head b773514 is C2S (a9defee + C2 + the LFM_HASH split
flip), so the merge base is C2 and only C3's two lanes come in. No
conflict: the flip and the one_row=auto default stay as they are.
The DEEP and out-of-domain denominators were inverted by a global
Montgomery scan: compute_denoms plus five scan kernels, each a full pass
over the domain with prefix and suffix scratch, although every row needs
only its own few inverses. By default now:

- compute_and_invert_denoms_ext3_dev runs one kernel,
  invert_denoms_rowwise_ext3_k{1..8}: each thread builds its row's
  denominators and inverts them in registers with one base-field
  inversion (adjugate over norm, the norms batched by Montgomery's
  trick, kernels/ext3_inv.cuh). More than 8 per row keep the scan.
- the fully resident R4 DEEP inverts its own row's 1 + K denominators
  (deep_composition_ext3_fused_m{1..4}) and needs no inverse buffer; it
  falls back to the buffered kernel above 3 points or on any
  precondition miss.
- the single-point OOD sums (the R3 composition parts: 1-2 columns, so
  1-2 blocks on the card) run on the row-chunked multi kernel, and the
  multi kernels' chunk count loses its 64 cap.

The values are the same field elements; raw limbs may differ by p, which
nothing downstream observes. Measured on block 25368371 on one RTX 5090,
ABBA behind a switch on 169b668: STARK (one_row=auto) -0.70 s whole
run against a 0.50 s A spread and -2.35 GiB device peak; WHIR kernels
-49.8 % with the wall inside the noise.

LAMBDA_VM_DEEP_INV_LEGACY=1 restores all three (the scan, the buffered
DEEP, one block per OOD column), read once per process with a banner,
as LAMBDA_VM_LDE_LEGACY does for the LDE. The legacy paths stay public
for tests/deep_inv_parity.rs, whose tests name both paths per call;
tests/deep_inv_setting.rs checks that the process setting is followed.
gpu_fused_deep_calls() counts the fused dispatch, and
cuda_path_integration asserts it follows the setting: a table that
silently fell back to the buffered kernel would still verify.
C4 = C3 (41549eb) + gap-fix/kern-int (d1dc455): the DEEP and OOD
denominators inverted row-wise, the R4 DEEP kernel inverting its own
denominators, and the single-point OOD sums row-chunked, all by default
(opt-out LAMBDA_VM_DEEP_INV_LEGACY=1). Every field element is the legacy
path's, so roots and proofs do not move.

C3S (4d8a096) is per-table-gpu's head: the merge base is C3, and only
the one K6 commit comes in. No conflict: the flip and the one_row=auto
default stay as they are.
recompute_lde_produces_byte_identical_proofs compares the bytes of two
proves of one instance, one per residency mode, at the test options'
grinding factor of 1. Under `parallel` the CPU nonce search is rayon's
find_any (crypto::grinding::generate_nonce), so the two proves can
return different valid nonces; the nonce is absorbed before the queries
are drawn, and every opening after it moves. The test fails whenever the
two searches disagree, whichever residency mode runs: on a laptop it
failed 11/20 at d1dc455 and 15/20 at 7d41668, and 20/20 passed with
LAMBDA_VM_DETERMINISTIC_GRIND=1 (the smallest nonce) or without
`parallel` (a sequential find).

The residency tests now prove at grinding factor 0, as zf_golden_tests
already does for the same reason. The residency mode acts on the main
LDE, which the grind never reads, so the comparison loses nothing it
could catch.
The residency-mode tests compared two proves byte for byte while grinding
with rayon's find_any, so any two nonce searches could disagree. They now
prove at grinding factor 0, as the golden tests do.
Every STARK wrap folded its program_id in-guest with one keccak
permutation. That permutation keeps the whole keccak family (LFM_KECCAK,
KECCAK_RND, KECCAK_RC) and, through its byte lookups, BITWISE in every
wrap, and makes every level-1 node re-verify those four sub-proofs per
child.

By default the wrap now computes the id at emission with
recursion::program_id_from_digest (still keccak, the SOUNDNESS.md 6.7
carve-out) over the values derived from the trusted ELF, publishes it
as program text in the fold's two-word layout, and binds every input
the fold consumed with an equality assert on the cell the verification
reads: the ELF digest halves the statement absorbs, the DECODE root
Phase A absorbs and the DECODE leg compares, pc_start, and the page
roots. A proof over any other value has no execution. The constants
are LFM_CONST rows, so they are in the wrap's program_id, which its
parent interns: the wraps, and every node and root above them, become
functions of the ELF, as the WHIR wraps already are. SOUNDNESS.md 6.9
states what binds each input.

LAMBDA_VM_STARK_WRAP_FOLD=1 keeps the in-guest fold. Its branch is the
previous emission unchanged, so every wrap, node and root program is
today's byte for byte. The setting is read once per process and named
on stderr; tests override it per thread.

With no keccak left the wrap's mask drops the keccak family and, under
RPX, BITWISE: no chip it instantiates sends BITWISE a lookup. On block
25368371 the census falls by 1.05 G cells (15 wraps -26.3 M each,
level 1 -783 M, levels 2-4 +127 M). The WHIR programs do not read the
setting.

Tests. Laptop: the standalone attestation at both root widths publishes
the fold's words, refuses a forged constant for each field and every
tampered cell, and leaves no BITWISE sender without its receiver.
Box: on the real epoch the default wrap publishes the fold's words,
emits no keccak, carries the masks above, and refuses a forged ELF
digest, pc_start or DECODE constant and a tampered cell; the WHIR
wrap's program is the same under both settings.
…evice halves

The artifact build (`build_artifacts_with_hasher`) now walks its commits
through `lfm::artifact_walk`: one plan (the eleven slot groups, the
LFM_BLAKE3 chunks, the LFM_HASH tail, under row-pair and, when the format
has it, one-row leaves) and one walk parameterized by a pass. `Pass::All`
is the build exactly as it ran: the same windows of `groups_in_flight`,
the same just-in-time BLAKE3 chunk materialization, the same
`commit_group_device_or_host_with` per group.

New, and unused by any caller yet: `build_artifacts_with_device_section`
(and `build_artifacts_sectioned`, which also returns the split). It walks
`Pass::Host` first — the groups the device would decline, committed by
`commit_group_host_with`, which cannot reach the card — and then enters a
caller-supplied section (a card permit) for `Pass::Device`, in the same
windows minus the host groups. The merge refuses a slot committed on both
sides or on neither. Routing is `gpu_lde::commit_reaches_device`, the
admission `admit_commit` applies (no device or below the row floor
declines; over budget still reaches the device and aborts there).

The registry drift tests pin the default walk's roots; the merge's two
refusals are unit-tested.
…t field

A WHIR round has three proof-of-work slots (folding, out-of-domain, query),
and the proof carries three nonces a round whatever the bits. A slot whose
grind has zero bits, and the last round's out-of-domain slot, is carried and
never read, so any value in it verifies. A query-only grind (P2) would leave
two such unbound fields a round.

ChainFormat gains `nonces: NonceLayout`:
- Three (the default): today's format, byte for byte.
- Spent: a round carries only the nonces its grinds spend. The in-guest
  arena has no word for an unspent nonce; ChainShape::carries is the one
  place the layout is written, and the word count, the hints and the arena
  words all read it. The host verifier refuses a nonzero value in a host
  field the layout does not carry (Error::UnspentNonce). RoundNonces keeps
  its three fields, so both layouts share one proof type and Three keeps its
  bytes.

GrindBits::query_only(bits) grinds before the query positions only. The
query count reads the query grind alone, so it does not move.

No production config uses Spent or a query-only grind yet, so every proof,
program and pin is unchanged. The legacy layout's chain programs, arenas and
proof bytes are pinned against values printed at 0428c39, and the
transcript closed form now prices only the grinds a config spends.
…t off)

STARK level 0 idles the card 12.9 s of 79.3 s (G1): 5.44 s inside
build_artifacts holds (64 % idle), 4.95 s inside multi_prove holds (30 %),
2.85 s with no hold. The interior's holds, on larger programs, are 10 % and
13 % idle; what differs at level 0 is the host load of the other five
workers (the epoch reconstructs above all). Three knobs, in
`lfm::card_schedule`, each moving work and never a committed byte:

- LAMBDA_VM_GAP_PREP_SCOPE=1: `build_artifacts_counted` holds the card
  only around the build's device commits (`build_artifacts_sectioned`);
  the groups under the device floor are committed on the host first. With
  the card trace on, each build prints its host/device split.
- LAMBDA_VM_GAP_PREP_NICE=<1..19>: host-only phases run on a rayon pool of
  their own whose threads take that nice value (Linux setpriority; libc as
  a Linux-only dependency): `lfm_prepare`'s execute and fill, the sectioned
  build's host half, and the STARK tree's wrap reconstruct, emit, census and
  harvest and the global child's harvest, emits and host verifies. The card
  holder keeps the CPU and the global pool.
- LAMBDA_VM_GAP_PREP_AHEAD=1: a level-0 wrap builds its artifacts on a
  helper thread while it executes and fills (`lfm_prove` is now
  `lfm_prepare` + `lfm_prove_prepared`; nothing before multi_prove reads the
  artifacts). Costs host memory: traces exist while the build may queue.

The permit stays a mutual exclusion under all three, and a build's device
commits run in the windows they always did. Unset, every path is the old
one; the STARK driver prints a CARD SCHEDULE line only when a knob is set.

Tests: the sectioned build equals the whole build over four programs
(split hash, three BLAKE3 chunks) and three one-row modes, and enters its
section once exactly when it has device work; on cuda, every device commit
lands inside the section. Execute and fill on the pool are cell-identical
to inline over the trace-identity cases; the pool's threads carry the nice
value (Linux). The prepared prove publishes the same words and verifies
(box-scale), refuses another hasher's artifacts, and the driver's AHEAD
helper returns the plain build's artifacts and a verifying proof
(box-scale).
Each WHIR base-chain round ground 20 bits before three challenges. Only the
query grind buys proven bits as placed: the folding grind sits before the
round's first sumcheck message, so the first folding challenge is redrawn by
varying that message at one hash a try, and the out-of-domain grind follows
the out-of-domain point. The new ZF lever `whir_grind` therefore defaults to
`query`: GrindBits::query_only(20) under NonceLayout::Spent, one grind and
one nonce word a round, 518 grinds a block instead of 1,472 at stack 27. The
query count reads the query grind alone and stays 112. The proven bits per
phase do not move: chain minimum 130.393 at stack 27, pipeline minimum
128.946.

LAMBDA_VM_ZF_WHIR_GRIND=all is the opt-out: GrindBits::uniform(20) under
NonceLayout::Three, the production config from before this commit. A test
pins it against a literal, and its chain programs against the values printed
at 0428c39. The banner gains `whir_grind=`. No univariate option reads the
lever, so no STARK proof, program or id moves.

Re-blessed: the production default chain's pins now describe the P2 chain
(12 grind permutations, 16,411 permutations, 32,590 arena words,
150,258 / 202,873 rows). Its previous pins move unchanged to the opt-out's
test. The banner strings in zf_format's tests gain the new key.
…head of it

The first prove of a process builds every domain and twiddle set its tables
need inside its prepass (0.38 s of the STARK base's first PROVE SPLIT on the
G1 traced run, with the card idle), and each device NTT twiddle size and
staging pair at its first use, between two kernels. These entry points let a
caller build the same state beforehand, off the prove:

- `Domain::from_options`: `Domain::new` reads only the options from its AIR,
  so the same domain can be built before any AIR exists; `new` delegates.
- `prover::warm_domain_and_twiddles`: fills the process-wide cache entry that
  `domain_and_twiddles` looks up, through the same function, so a warmed
  prove reads exactly what it would have built.
- `gpu_lde::prewarm_device`, over `device::prewarm_twiddles` and
  `device::prewarm_staging_pairs`: the backend, every twiddle size up to a
  bound, and a number of staging pairs, built by the functions that build
  them on demand.

Nothing calls them yet; no value a prove commits or reads changes.
…VM_OOD_COLUMNS_ON_CALLER)

A round-3 OOD table is one or two rows high, and `Table::columns` transposes
it with a rayon parallel iterator. The per-table drivers of `multi_prove` are
plain OS threads, so each call is injected into the global pool and the driver
waits for a worker, with its table's next device work unsubmitted. When the
pool is busy with other work, that microsecond transpose waits behind it: on
the STARK base's G1 traced run the host-only OOD absorb summed 5.16 s over
split 9's tables and 2.73 s over split 11's, while the level-0 lead-in was
verifying base epochs in the base's tail, against about 0.02 s in a quiet
split. Three sites read these columns per table: the round-3 absorb and the
two device DEEP dispatches (and the host DEEP loop).

`LAMBDA_VM_OOD_COLUMNS_ON_CALLER=1` builds them with `Table::columns_serial`,
the same per-column read on the calling thread. Unset or `0` keeps the pool
(today's); anything else panics. The values and their order are those of
`Table::columns`, so the transcript and the proof are the same bytes.

Tests (`tests::host_schedule_tests`): the serial transpose equals the parallel
one; the switch's parsing; and a whole multi-table proof, grinding off, is
byte-identical with the switch on and off, with a counter showing the named
arm ran. Reversing the serial column order fails the byte test. The file also
tests the warm-up entry points of the previous commit: an options-built domain
equals the AIR-built one, and a warmed cache entry is the one a lookup hits.
…AD_AHEAD)

Before the card has anything to do, the serial head commits DECODE's
precomputed columns on the host (RPX over a 2^22-row LDE, about a second on
the block ELF, before epoch 0 executes), epoch 0's builder then initialises
the device, and the first prove's prepass builds every domain and twiddle
set. On the G1 traced run that is 2.44 s from run start to the first PROVE
SPLIT, plus 0.38 s of prepass, all with the card idle.

`LAMBDA_VM_BASE_HEAD_AHEAD=1` starts two helpers beside the producer: one
initialises the device, commits DECODE there and then builds the device's
per-size twiddles and four staging pairs; the other builds the host domains
and twiddles of every trace size up to the epoch's. The epochs wait for the
commitment where they first use it (each epoch's preparation, or its prove),
and the base's observer hears it from the helper, before any epoch is
prepared. Unset or `0` keeps the serial head (today's); anything else panics.
Each helper step prints a `BASE HEAD:` stamp under `LAMBDA_VM_BASE_SPLIT`.

The device commitment is `decode::compute_precomputed_commitment_device_or_host`:
DECODE's five precomputed columns, row-major, through
`lfm::commit::commit_group_device_or_host_with`, whose device root its device
tests pin to the host's; `multi_prove` also rebuilds DECODE's precomputed tree
at the first epoch and refuses the proof if its root differs. The host arm is
the same interpolate, coset-evaluate and commit as the serial head's.

Tests: the device-or-host root equals the host root in both leaf layouts, for
a 13- and a 4,097-instruction program (the latter crosses the device's commit
floor on a device build) and from an ELF; building the group column-major
fails both. The head run ahead proves what the serial head proves (the shared
comparison, now `assert_same_proved`, also used by the prep-ahead test), and
the observer hears the serial head's commitment exactly once. The switch's
parsing.
…TREE_TAIL_THREADS)

The level-0 lead-in builds wrap prologues in the base's tail: each verifies
its base epoch, fanning out over every thread of the global rayon pool while
the base still proves its last epochs. The base's per-table drivers are plain
OS threads, so each parallel iterator they start is injected into that same
pool and waits for a worker. On the STARK base's G1 traced run the prologue
spans hold 4.83 s of the base's 10.88 s of card idle, and the base's host-only
OOD absorb, one such injection, summed 5.16 s over split 9's tables against
about 0.02 s in a quiet split.

`LFM_TREE_TAIL_THREADS=<n>` runs each prologue inside a rayon pool of n
threads of its own, so it fans out over those and never queues ahead of the
base's work in the global pool. The helper takes the lead-in's context before
entering the pool, so no pool thread waits on the base. Unset, empty or `0`
keeps the global pool (today's); anything but a count panics. The prologues
compute the same programs and arenas either way; the tree's program ids are
the check on a box run.

Tests: the switch's parsing, and work run in a pool of its own fans out over
that pool's threads and returns what it computed.
`LAMBDA_VM_GPU_DEVICE_ONLY_THRESHOLD` decides which tables keep a host copy
of their LDEs, and those copies are downloaded inside each table's commit:
on the STARK base's G1 traced run, 39.74 GB of retained-LDE downloads (8.08 s
of driver time), about 80 % of it KECCAK_RND and ECDAS tables that sit below
the 2^19-row default only by row count. A run that moves the threshold has to
say so in its log, so the resolved envelope is printed once, where it is read:
`[gpu] device-only envelope: LDE >= <rows> rows (<source>)`. No behaviour
changes.
Two checks for a device build, so a host fallback cannot pass as the device:

- the head helper's device stamp says whether the device took the warm-up
  (`BASE HEAD: device warm-up` or `... declined`), since a decline is
  silent by design (the prove then builds the same state on demand);
- `the_head_decode_root_is_committed_on_the_device` (cuda): the 4,097-
  instruction DECODE map, 8,192 rows at blowup 2 and so above the device's
  commit floor, moves the device group counter and still gives the host
  root. The counter is process-wide, so the test is run on its own.
The ZfFormat::DEFAULT doc quoted a block timing for whir_grind=query from an
earlier measurement. A measured number in a comment goes stale, so the doc
now says what the lever does and why it loses no proven bits.

The same reasoning in whir_chain's module header, GrindBits::query_only and
WhirGrind gave the out-of-domain grind's reason as "it follows the
out-of-domain point". That is half of it: the grind sits right before the
batching challenge and does guard it. Dropping it costs nothing because the
batching challenge has far more bits than the target without any grind.

Comments only.
the_inner_node_verifies_two_leaf_nodes needs FAN_IN^2 = 4 epochs and took
FIXTURE_EPOCH_LOG2 - 1. Since 8f9aef1 moved the fixture to the 48-cycle
continuation-fixture at FIXTURE_EPOCH_LOG2 = 5, that is 16-cycle epochs and
three of them, so this box-tier test has stopped at its own epoch-count
assert in setup, before proving anything, under either wrap attestation. A
pre-existing fixture drift, found by the gates at e413989.

Two below the shared constant is safe by an assertion the suite already
holds: the_fixture_guest_commits_in_an_intermediate_epoch keeps the guest
above one shared epoch and within two, so 8-cycle epochs give at least five
(six today), where 16-cycle ones give three or four.
…le v2)

Job 202's NICE arms lost 9.65 s at STARK level 0 while their mechanism
moved the right way (level-0 held time -5.4 s). The cause was the one
host-phase pool shared by the six level-0 workers: every whole host
phase was an injected job there, and a rayon thread blocked in a join
runs injected jobs before it returns to its own (rayon-core 1.13
wait_until_cold). A thread waiting inside one wrap's reconstruct ran
other wraps' whole phases nested on its stack, so the reconstructs that
started first finished last (2-3 s became up to 22 s) and the card sat
with no holder for 17 s of the level.

host_phase now installs into a pool of the CALLING thread's own, built
at its first host phase and dropped with the thread. A worker runs one
phase at a time, so its pool only ever holds that phase's jobs; a
caller that is already a rayon worker runs the phase inline. Each pool
prints one line naming its caller and how many of its threads took the
nice value, counted after every thread has started. Knob unset: inline,
as before.

Tests: two callers' phases never share a thread (with the shared-pool
design as the control, which puts both on the same threads); a caller
reuses its own pool; a rayon worker runs a phase inline; unset runs on
the caller; a new pool reports every thread.
A table outside the device-only envelope downloads its whole main and aux
LDE inside its commit, on the driver that would otherwise submit its next
table. At the 2^19 default the STARK block's base downloaded 39.74 GB that
way (8.08 s of driver time), most of it KECCAK_RND and ECDAS tables of 1,480
and 521 columns whose device paths already run: they sat outside the envelope
by row count alone. At 2^16, with the barycentric floor (trace >= 2^14 rows)
below it, the base downloads 7.27 GB and the block proves 1.40 s faster
(FAST job 206, ds850-861, two arms each: wall -1.40 s, base -1.10 s; program
ids unchanged, the root verified).

`LAMBDA_VM_GPU_DEVICE_ONLY_THRESHOLD=524288` restores the 2^19 envelope. The
default is process-wide, so the WHIR pipeline's recursion proofs (FRI STARKs
through the same gate) take it too; that pipeline was not measured.
The head helpers (DECODE committed on the device, the first prove's domain,
twiddle and staging state built beside epoch 0) become the default:
`LAMBDA_VM_BASE_HEAD_AHEAD` unset, empty or `1` runs the head ahead, and `0`
is the named opt-out that runs the serial head as before. FAST job 206
(ds850-861, two arms each against the serial head): the head 2.4 -> 1.0 s, the
first prepass 0.37 -> 0.00 s, the base -2.00 s, the block -1.45 s; program ids
unchanged, the root verified.

The switch's test now pins the new reading (ahead unless exactly `0`); the
equivalence test still proves both heads by parameter.
… by default

The level-0 lead-in's prologues put parallel work on whatever pool they run
in while the base still proves its last epochs, and the base's per-table
drivers, plain OS threads, queue their own parallel iterators behind it in the
global pool: the host-only OOD absorb summed 5.2-5.3 s over split 9's tables
in all three G1 arms. In a pool of 16 threads of their own (FAST job 206,
ds850-861, two arms each against the global pool) the base's worst split
absorb fell to 0.04 s, the base by 2.35 s and the block by 2.20 s, level 0
+0.05 s; program ids unchanged.

The pool now belongs to each `LeadIn`, sized by its caller: the STARK tree
defaults to `STARK_TAIL_THREADS` (16); the WHIR tree keeps the global pool,
its tail not having been measured in one. `LFM_TREE_TAIL_THREADS` overrides
either (`0` is the global pool, the named opt-out; a count, a pool of that
size), and the line naming the choice is printed where the lead-in starts.
The rationale no longer says the prologues' verify fans out over the pool;
what is established is that isolating them removed the stalls.

Tests: the switch over each pipeline's default; a lead-in with a pool of 3
builds its prologues on that pool's threads and one without a pool does not
(routing the prologue around the pool fails it).
The knob (pair | k3 | k4 | best, default pair) reaches the STARK block's
base tables through ZfFormat::base_proof_format at block_base_options.
proof_format, which every LFM proof carries, keeps the pair layout, so the
LFM chips do not move under any policy. An unknown value aborts; k3, k4 and
best abort too until the device can split their composition parts. The
banner prints logup=... on every setting.

Tests: under pair every production VM table and every LFM chip keeps its
widths, part count and constraint-artifact digest, against a golden list
taken at d740eb5 before arities existed; wide policies give each table
the layout the rule picks; every production table's artifact matches the
folders and the device blob under k3, k4 and best.
A degree-5 AIR (LogUp groups of four) splits its quotient into four parts.
The split is the existing radix-2 one applied twice on the LDE coset:
A/B from H(x) and H(-x), then H0..H3 from A and B at y and -y (y = x^2),
each part extended x4 with weights g^(-3k)/q.

- math-cuda: decompose_d4_ext3 kernel and decompose_d4_into_slabs (12 slabs,
  zero-padded) feeding the batched slab LDE at ratio 4; upload_comp_h for
  the parity tests.
- gpu_lde: try_decompose_extend_d4_dev, the d=2 arm's contract with m = 4
  (drained to host or device-only placeholders); the part-slab admission
  takes the part count.
- prover: decompose_and_extend_d4 (host mirror, also the host arm for four
  parts), cached d4 weights, 1/(2 g^2 w^2i) and host twiddles; the R2
  device arm, the downloaded-H fallback and the xcheck mirror take four
  parts; device_only_for admits 2 and 4 parts.
- LOGUP_K4_IMPLEMENTED = true: LAMBDA_VM_ZF_LOGUP=k4 is selectable; k3 and
  best still abort (no three-part device split).

Tests: the host split equals interpolate -> break_in_parts(4) -> evaluate
at blowup 4 for base and extension H, n = 4 .. 1024. GPU (ignored, box):
device split == host mirror for n = 2^8 .. 2^21 (drained and resident),
the k4 aux build through the device kernels, and a device-only k4 proof
equal to the host-composition proof byte for byte.
The compiled set now also covers the VM tables under
LAMBDA_VM_ZF_LOGUP=k4 at blowup 4. Nine programs are new (CPU, CPU32,
MEMW, MEMW_A, SHIFT, MUL, DVRM, HALT, COMMIT under k4); tables the rule
keeps on pairs share their existing kernel. LFM chips keep pairs under
every policy and the local-to-global tables are too narrow to move, so
neither is rebuilt. Without these kernels CPU and MEMW_A would fall to the
interpreter under k4.

The hot-table check names the k4 labels and keeps KECCAK_RND, ECDAS and
ECSM interpreted under both formats. The host parity check
(ccomp_host_check.cpp) passes on all 45 programs.
Two ignored box-tier tests for D-LOGUP's S2 gates:

- the_block_tree_verifies_logup_k4_and_refuses_it_under_pairs: a small
  block proved under k4 (some sub-proofs carry four composition parts)
  verifies on the host under the k4 options and is refused under the pair
  options; its leaves verify the k4 base in the guest, the block verifier
  derives the tree from the k4 options and accepts the proved top, and the
  same top is refused against the plan derived under pairs.
- logup_k4_vm_proof_negatives: a fib_iterative_160k block under k4 (CPU
  device-only, four parts split on the card) verifies; a tampered committed
  term cell, a moved absorbed multiplicity, the pair options, 3 or 5 part
  OODs and a tampered part OOD are each refused.
…gup/s2

dacbd5b makes the bounded-slot (SI) interpreter the default for the
programs with no compiled kernel (KECCAK_RND, ECDAS, KECCAK), so the k4
structures of those tables compose through it.

Conflict: compiled_constraints::production_programs, now pub(crate) on the
new head, keeps this branch's (opts, extras, suffix) signature;
compiled_option_sets is pub(crate) too. The budgeted (SI) tests take their
programs from compiled_option_sets, so the SI lowering, its host model and
the SI host parity dump now cover the VM tables' k4 programs as well
(KECCAK_RND, ECDAS, ECSM, KECCAK k4 included).
FAST 860 read two harness faults in the GPU tests, not device faults (the
device-only k4 proof equalled the host-composition proof byte for byte):

- d4_parts_equal_the_host_split starts at LDE 2^10, below the default
  device floor, where the split declines by design. Documented: it needs
  LAMBDA_VM_GPU_LDE_THRESHOLD=1024, as the proof test does.
- d4_parts_gpu_aux_build_groups_by_four read the host aux table, but on a
  cuda build the resident aux build leaves the columns on the card. The
  aux check now takes a reader: the host table for the term-column path
  (resident aux off), the downloaded resident buffer for the resident path,
  and both device paths are tested.
… 8 GiB floor

BIG 585 (the median, spill off, 2 + 2, one binary) read late emission EFFECTIVE:
base-phase VmRSS −10.37 GiB (86.10 against 96.48), recursion −0.31 s, whole −2.15 s,
one top id. The trigger fired at ≈ 149 s of a ≈ 193 s base, the emission ended ≈ 24 s
before the base, the harvest never waited, and 98 % of the programs' 30.6 GiB landed
on pages phase B had freed. With spill on it is inert (−1.28 GiB, +0.55 s).

NOEPOCH_TREE_EMIT_LATE now defaults to auto: leaf 0 at the shape, then the rest
late at margin 1.25 when leaf 0's bytes put the others' estimate at 8 GiB or more,
and at the shape below that. A small tree's programs barely move the peak, and its
short phase B may not free their bytes before the base returns, which would leave
the harvest waiting. A margin forces late at any size; 0 or off restores emission
at the shape. The programs, and so every id, are the same in every mode.
…to tree3/late

Late leaf emission, default auto above an 8 GiB floor (BIG 585: base-phase VmRSS -10.37 GiB,
recursion -0.31 s, whole -2.15 s at the median), merged onto #1013's current head. No conflict:
dacbd5b touches the stark prover's composition interpreter and its tests, not the block tree.
…e's sibling proofs share the card by bytes

LAMBDA_VM_SHARED_VRAM_GATE=1 (default off): every multi_prove in the process
admits its tables through one VramGate at the card's budget instead of a
fresh full-budget gate per call, and the LFM artifact commit, the one device
dispatch outside a prove, takes its device set from the same gate
(shared_vram_admit). With that cross-proof running total the recursion's
card permit stops being an exclusion: hold() takes no card, and sibling
proofs run their device phases together within the budget.

At the median block the recursion leaves the card idle 32.65 s of 98 s, 24.62 s
of it inside multi_prove holds (BIG 469): the holder's tables upload, absorb
and query while every sibling waits for the card (77 s of waits). Under the
shared gate a sibling's kernels fill those stretches. Off, a gate per call and
the mutual exclusion, as before. Tests: the shared admission holds and
releases its bytes only when on; under it the permit takes no card (a second
hold on one thread and a hold on another proceed), and the test fails with
the permit ignoring the gate.
…sha for both

Late leaf emission by default (BIG 585 EFFECTIVE) and the compact program form (a), the
instruction vector held at its length (FAST 672 GREEN: ids equal on the box, device = host;
BIG 586 GREEN: the median top unchanged), both on #1013's dacbd5b, stacked so one FAST merge
gate covers them. No conflict: the two touch different parts of the block tree tests, and only
compact touches the LFM compiler.
…treaming derivation joins the landing sha

The block verifier (verify_block_tree -> derive_top) drops each derived program once its child is
derived instead of holding a whole level: BIG 587 EFFECTIVE, the verifier alone in its own process
-9.44 GiB host peak at the median (32.33 -> 22.89), derive -1.48 s, ids unchanged; FAST 671 GREEN
at 1x. Stacked onto tree3/land (late by default + compact (a), on #1013's dacbd5b) so one FAST
merge gate covers all three. No conflict: vstream touches block_plan.rs and the harness's shape file
and verifier-only test.
…brated, capped, the pool released

FAST 473 and 474 overran the 31.36 GiB card with six or seven proofs on the
shared gate: first the memory pool's retained blocks (the base's frees,
reserved and invisible to the gate), then, with the pool released, about
4.5 GiB the gate never sees (caches, modules, unreleased frees) plus each
prove's bytes beyond its tables' estimates, times six.

Under LAMBDA_VM_SHARED_VRAM_GATE=1 now:
- the device's pool releases at each sync unless LAMBDA_VM_MEMPOOL_RELEASE_MB
  says otherwise;
- the gate is shared only while the tree's driver has armed it for sibling
  proofs (device_permit::arm), never in the base;
- arming calibrates it: the pool is trimmed and the budget becomes the card's
  free memory less a margin (2 GiB + 0.5 GiB per prove place), capped at the
  configured budget, printed once per arming;
- at most LAMBDA_VM_SHARED_GATE_PROVES (default 3) proofs are inside
  multi_prove at once; artifact builds take bytes, not places.

Tests: the calibrated budget's arithmetic; the release threshold the knob and
the shared gate choose; with the places full a fourth prove waits until one is
released while an artifact build does not (and the test fails with the cap
ignored).
… stay in the account

A Retain prove keeps each device-committed table's main LDE, trace snapshot
and tree on the card from its Round-1 commit until its fused task ends. The
shared gate dropped those bytes with the Round-1 permit, so sibling proves
were admitted against a budget the card no longer had (FAST 474).

Under the shared gate (LAMBDA_VM_SHARED_VRAM_GATE=1, armed) and Retain:
- the Round-1 task carries the bytes its commit left resident past its
  permit (CarriedBytes), sized from what the device handle kept;
- each fused task is admitted for its set less what it carries, and takes
  the carried bytes over until it ends;
- each prove first claims room on the gate (ResidentClaim): the residents'
  bound plus the largest table's fused set, settled after Round 1 to the
  carried bytes plus the largest top-up, shrinking as tables finish.
  Claims are admitted only while their sum fits the budget, so carried
  bytes cannot wedge the gate; a prove whose claim alone exceeds the
  budget carries nothing and runs as the only claim.
- calibration skips while a claim is in force; LAMBDA_VM_SHARED_GATE_TRACE=1
  prints one SGATE line per change of the account, for the box readout.

Off (the default), nothing is claimed or carried and every admission is
the one before; the proof bytes do not move (resident_carry_tests compares
them). Tests: the exact account at every fused start with one driver, the
claim arithmetic, and two carrying proves that wedge without claims and
finish with them; each fails under its mutation.
…s the data) into logup/s2

The test-only fix FAST 862 gated (S1 PARITY GREEN): the d4 split test
needs the 1024 device floor, and the k4 aux test reads the host table or
the resident buffer, both device paths. It was held out of logup/s2 until
the S2 runs (FAST 861, BIG 600) were done at 1bbc7f2.
…ch sync only while armed

The knob used to set release 0 for the whole process, so the base, which
runs before any arming, lost the retained pool its repeated allocations
reuse: 1x base +0.22 s (t +2.5, FAST 477/478). The pool is now created with
its usual threshold (the knob's, else retain-all); arming the shared gate
lowers it to 0 (unless LAMBDA_VM_MEMPOOL_RELEASE_MB set one) and disarming
puts it back.

Two readouts for the box: SGATE lines carry the pool's live bytes beside
the gate's admitted bytes, and LAMBDA_VM_SHARED_GATE_READOUT=1 splits the
card's used memory at each arming (at the trim, after draining every stream
and trimming again, pool live and reserved, outside the pool) and prints the
budget a drained context would give. The readout waits for the device, so it
is for a run off the clock. Both are off by default.
… streaming verifier derivation) into logup/s2

No conflict. Under the default pair policy every program id this head pins
is unchanged (the_compact_program_form_keeps_a_small_trees_ids passes on
the merge), as are the 44 production programs of the pair golden, the
compiled kernels and the stark goldens.
…rates

The calibration trimmed the pool and read the card's free memory while
frees were still queued on the base's streams: a queued free holds its
block until its stream reaches it, and the trim cannot hand back what the
pool has not been given. At the 1x level-0 arming that left 11.6-12.8 GiB
used after the trim, of which 0.33 GiB was live; draining first leaves
1.3-1.5 GiB (FAST 479a/b readout runs), so the budget is the configured cap
instead of 15-16 GiB. The drain runs only when the gate calibrates (nothing
admitted, no claim in force), between sibling levels.
… into sched/l2p at step 3 (73e3d64)

The landing train for lever 2: the shared VRAM gate (carried residents behind
per-prove claims, the pool releasing only while armed, the arming drained
first) on #1013's current head. No conflicts; #1013's changes do not touch
the gate, the admission or the device-set model the estimates come from.
pin_shared_vram_gate is process-wide; under cargo test's threads a pin
could hand a concurrent resident_carry_tests prove the shared gate and
move its account readings. One test-only lock orders them (nextest's
process-per-test runs were never exposed).
The block tree's sibling proofs now share the card by bytes unless
LAMBDA_VM_SHARED_VRAM_GATE=0 restores the exclusive card permit. The gate
still acts only while a caller arms it for concurrent proofs
(device_permit::arm with more than one worker), so the base and every
single-prove path are unchanged, and proof bytes do not move: the gate
changes admission order only (one top id across every A/B run, FAST
477-479). FAST 479 P8 against the exclusive card: 1x recursion -0.59 s
(t -17.2), whole -0.51 (t -6.9), base -0.02, VRAM <= 26.1 GiB.

The permit tests that check the exclusive card pin the gate off; without
the pin the falsifier test now fails, which is the flip taking effect.
A block plan absorbs the ELF's digest into every leaf, and the fixture ELF's bytes
depend on the clang that assembled it: the laptop's (digest 3b39e219…) and FAST's
(f5e120d0…) differ, so a single pin set holds on one machine only (FAST 670). The two
pinned-id tests now look up their ids by the plan's ELF digest, with the laptop's
pins (recorded at 541f4bd) and FAST's (recorded at fe1fb16 in FAST 672, where the
compact form's ids matched them on the device and the host). An ELF with no pins
refuses by name before any program is derived;
pinned_tree_ids_refuse_an_elf_they_do_not_know checks that.
MauroToscano added a commit that referenced this pull request Oct 3, 2026
…14fea1

crypto/stark/src/narrow.rs and crypto/stark/src/spill.rs are #1013's files,
byte-identical to mem/stark-f1 @ e5214fea1 (i-mem3's F1: the page-backed
Bytes, the streaming Digester, the fused writer, over the S1/S2 store). #1013
is their one source of truth: #1014 never edits them, re-imports any change
from there, and they merge as one when both PRs reach main.

Wiring only: the two modules in lib.rs (dead code allowed there, since some
parts serve #1013's prover alone), math's page-bytes feature, and libc
unconditional as the store needs it (it was behind disk-spill). #1013's prover
glue (TraceTable::spill_main) and its spill tests are not imported; #1014's
adapter and tests follow.
MauroToscano added a commit that referenced this pull request Oct 3, 2026
…ault off)

LAMBDA_VM_BLOCK_SPILL=always (BlockOptions::spill, BlockSpillPolicy) opens
#1013's spill store for the prove. Phase A hands each committed group's packed
tables to it as the group is installed, past the first two groups (phase B
reads them before a read-back could land) and never past the writers' queue:
the committer is the store's one producer and checks the room first, so it
never waits on the writers. Phase B reads the spilled tables back in group
order, two groups' bytes ahead of their uploads, restores each group and the
next one before their uploads, and lets a finished group's packed columns go.
While the spill is on, the packers put the packed bytes in pages of their own
(the card's pack and b2's finish pack), which the writers take whole.

The adapter between NarrowColumns and the store's NarrowMain moves the parts
and never copies the bytes (a test checks the address). A read that fails or
does not match its digest refuses the prove with MlError::SpillFailed, naming
the table. BLOCK SPILL reports the store's counters and the read-back.

Off by default: with the spill off nothing changes on any path (the heap, the
same allocations, no store). Tests: a spilled block proves the same proof
bytes as a held one under the deterministic grind, each slot read back once;
a byte flipped in a spilled table on disk is refused.
At the median the setup line's "unnamed" read 6.31 GiB in every run and arm
of BIG 110: deterministic, so sized by the program or its pages, not the
run. The line now reads the heap before VmAirs::new as well and names what
the AIRs and the AIR/trace pairing add. Measurement only (memlog).
…unted apart in BLOCK MEM

The two spill measurement rounds (BIG 115, 116) read the writer's time by
step; this keeps that split for whoever measures the spill next.

- SpillStats splits the writers' seconds into digest, copy (the aligned
  copy under O_DIRECT) and pwrite; the rest is the free and the locks.
  The BLOCK SPILL line prints them.
- SpilledMain::is_resident. The BLOCK MEM traces/setup lines count
  spilled traces apart, and their bytes in the heap only while a slot
  holds them; they were counted at 8 bytes a cell ("traces 228.87" at the
  median with the spill on).

Test: a store's step seconds add up to no more than its writer seconds,
the copy appears exactly under O_DIRECT, and written slots leave memory.
With LAMBDA_VM_BLOCK_SPILL unset the block now runs `auto`: a committed
packed trace is spilled only once the host's bytes (the larger of VmHWM
and the cgroup's memory.current), the reserve for what the block still
needs and the trace would pass the target (memory.max, or MemTotal, less
10 GiB). A block that fits spills nothing; a block above the target
spills what would not fit instead of running out of memory. `off` keeps
every trace, as the default did before. No proof byte depends on the
policy: the words read back are digest-checked.

Tests: unset reads as auto; the stream under auto precommits the same
instances and builds the resident traces.
MauroToscano added a commit that referenced this pull request Oct 3, 2026
#1013 landed 4cb0a07 (the writer's digest / copy / pwrite step timings on
SpillStats, SpilledMain::is_resident). crypto/stark/src/spill.rs is again
byte-identical to noepoch/stark @ edddc68; narrow.rs was already. Readouts
only: nothing #1014 calls changed.
MauroToscano added a commit that referenced this pull request Oct 3, 2026
… reserve

LAMBDA_VM_BLOCK_SPILL now reads as #1013's (prover/src/block.rs SpillPolicy …
spill_decision @ edddc68): auto (and unset) | off | always | <GiB> resident
budget. auto spills a committed table once the host's bytes (the larger of
VmHWM and the cgroup's memory.current), the reserve and the table would pass
the target (LAMBDA_VM_BLOCK_SPILL_TARGET_GIB, else the cgroup's memory.max
less 10 GiB, else MemTotal less 10 GiB). Phase A decides per table and never
past the writers' queue; groups 0 and 1 count but never spill.

The reserve is #1014's own, inferred from the median block (BIG 565): the
finish's p5 transient (25.5 GiB) and the tree's leaf programs alive through
phase B (16.3 GiB), about 1.25 GiB per total G cells or 1.65 per G cells
committed so far, plus 6 GiB for phase B's bump and the read-back window.

BlockSpill carries the policy's choice (SpillWanted) and the kept bytes and
committed cells it reads. A block that fits spills nothing; the store opens
for every policy but off. Tests: the knob and the decision (unit), a budget
the block fits in spills nothing and one of zero spills (block).
…Total

i-m4b found the gap: the target read only cgroup v2's memory.max. On a
v1 host with a limit below MemTotal (FAST: 57.53 GiB against 59.93) it
fell through to MemTotal less 10 GiB and aimed 2.4 GiB too high there.

- cgroup_memory reads v2's file under the unified hierarchy, else v1's
  under the memory controller (memory.limit_in_bytes for the limit,
  memory.usage_in_bytes for auto's host charge), each at the process's
  cgroup path and then at the hierarchy's root, which is what a container
  without a cgroup namespace sees (FAST: 12:memory:/docker/<id>, the
  limit at /sys/fs/cgroup/memory/memory.limit_in_bytes).
- The target is min(cgroup limit, MemTotal) less 10 GiB
  (spill_target_from): v1's unlimited sentinel gives MemTotal's, and a
  v2 limit above MemTotal no longer wins.

Tests: fake cgroup trees for v2 at its path, v2 `max` falling to v1, the
FAST layout, v1 at its path, a shared v1 hierarchy, and none; the target
rule with v1's sentinel, a cap at MemTotal, and neither. Dropping the v1
read or the minimum fails them.
MauroToscano added a commit that referenced this pull request Oct 3, 2026
…@ 278e6a8

The rule #1014 copies from #1013 read only cgroup v2's memory.max and
memory.current. On a v1 host (FAST) the target fell through to MemTotal
less 10 GiB, and the host's charge was not read at all. #1013 fixed both
in 278e6a8; this re-copies spill_target_bytes, spill_target_from,
cgroup_memory and host_bytes_now from that sha, with its two tests (the
cgroup files from fake trees, and the target rule), cited as before.

The only changes from #1013's text are the citations and the tests'
temporary directory, which gets a name of its own so the two copies of
the test never share one if they ever run in one process.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant