Skip to content

WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress - #1014

Draft
MauroToscano wants to merge 1306 commits into
mainfrom
noepoch/whir
Draft

MauroToscano wants to merge 1306 commits into
mainfrom
noepoch/whir

Conversation

@MauroToscano

@MauroToscano MauroToscano commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

One WHIR (multilinear) proof per block, with no epochs (the prove-and-retire / VADCOP shape). This is a second prover next to #1010's epoch-based one. #1010 stays the reference and this branch does not touch it. The branch starts from #1010's head f3d359998; compare against f3d359998 to see only this work.

Status: draft, work in progress. One binary (FAST 416, #1010's head 88b0d31 merged in, ABBA × 2): the block tree 23.20 s whole vs #1010's epoch tree 31.80 s, −8.60 s. Since then: 20.89 s whole at d9d0ac5 (FAST 424: base 16.76 + recursion 4.13), with the batched argue on by default, ops dropped after commit (M1), the lean walk and paged executor memory (E1/E4), and one prepared stack per group. e06dce5 added narrow trace storage (M4: same bytes, time-neutral, max RSS 1× 36.2 → 29.5 GiB, 4.13× now proves). 3cfffe2 added the KECCAK and ECSM split (§6: the same tables and proof bytes on every block measured so far, FAST 830). 01da99f uploads phase A's next group beside the current commit (prover-only, same bytes, base −0.60 s at 1×, FAST 832). 59e9890 lays out the streamed chunks on three threads by default (E3; prover-only, same bytes; whole block −0.48 s to 19.75 s and max RSS −1.51 GiB at 1×, FAST 838). 82f9046 re-hashes phase B's kept-top paths in parallel (prover-only, same bytes; phase B −0.96 s, whole block −0.82 s to 18.90 s at 1×, FAST 843 + 844 pooled). d169edd lets the tree's leaves execute while their artifacts are built (prover-only, same bytes; level 0 −0.24 s, whole block −0.24 s to 18.54 s at 1×, FAST 841 + 842 pooled). 70eee3e lays out the rest in waves and builds KECCAK_RND as its tables (W, lane i-m4b; same bytes; max RSS 88.9 → 79.5 GiB at 4.13×, base −0.14 s at 1×). fac261f lets the tree's first finished leaf execute during phase B (R2-ii; prover-only, same bytes; whole block −0.24 s to 18.39 s at 1×, measured before the merge with W, FAST 845 + 846 pooled). 365e3ab packs the finish's tables as they are built and commits them narrow (b2, lane i-m4b; same bytes; max RSS 78.8 → 52.1 GiB at 4.13×, 28.5 → 22.5 at 1×; base −0.36 s at 1×). 351a773 adds a device readout to the tree test (test-only), and the measurement posture moves to the card memory pool's code default, which retains freed memory: a baseline shift, not a prover change (whole block 17.91 → 17.27 s at 1×, FAST 848 + 849 pooled; every number above was taken under the old posture and stays as measured). Head 9a18ab5 makes the prover refuse a block over the group maximum as its groups close, with the verifier's own error (lane i-m4b; prover-only, same bytes, FAST 858). Every format change below but §6 (the port of #1013's S0a, gated FAST 830) is approved by Mauro (10-02); the cryptography team reviews them at the end (CRYPTO-REVIEW.md rows W-0…W-6, W-8).

What it is

  • prover::block_whir::prove_block_whir / verify_block_whir. The whole block is one MultiProof:
    • every table is cut into instances of at most 2^21 rows, and KECCAK_RND into 2^16-row instances;
    • memory is the monolithic PAGE argument, with no local-to-global bookend and no cross-epoch proof.
  • The tables are packed into groups of at most 3 stacked polynomials (2^27 each).
    • Phase A commits each group and retires it: the codewords go, and only the top of each tree stays on the host (the bottom 4 levels are dropped).
    • The roots block: every root goes into the transcript, then (z, α, β) are drawn once.
    • Phase B proves each group on its own fork S_post ‖ g: upload again, argue (LogUp-GKR), re-encode with the NTT alone, open. The dropped tree levels are rebuilt from the queried cosets and checked against the kept nodes.
    • The bus is checked once, over every table of the block.
  • Streamed phase A. The executor runs in 2^20-cycle windows feeding the shared WindowedTraceBuilder (noepoch/windowed-builder, also used by STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013). Every full chunk of CPU, MEMW_R, MEMW_A, MEMW, LOAD, LT, SHIFT and STORE is laid out and committed as soon as it exists.
  • The builder is deterministic: on block 25368371 the whole-run build and two windowed builds under host contention give identical per-table digests (FAST 411; 129 tables row for row, the six HashMap-ordered tables as row multisets).

⚠ Format changes (all approved by Mauro, 10-02, but §6, the port of #1013's S0a; for the cryptography team's end review; ledger rows W-0…W-6, W-8 in CRYPTO-REVIEW.md)

The block is one WHIR proof (BlockWhirProof, one MultiProof) over all of the block's tables.

  • Phase A commits the tables in groups as the trace is built. Each group is at most 3 stacked polynomials of 2^27, and each group is committed and retired as it closes.
  • The challenges. The transcript then absorbs the statement, the groups' roots and the derived prepared roots, and draws z, α, β.
  • Phase B proves each group on its own fork of the transcript (S_post ‖ g).

Every parameter below is a verifier-side constant. None is read from the proof.

constant value meaning
BLOCK_GROUP_POLYS 3 a group of two or more tables stacks into at most 3 polynomials of 2^27
BLOCK_MAX_GROUPS 64 at most 64 groups
BLOCK_MAX_TABLE_VARS 27 no table stated over 2^27 rows
BLOCK_MAX_KECCAK_RND 2^12 at most 4096 KECCAK_RND tables
BLOCK_KECCAK_RND_MAX_VARS 16 each at most 2^16 rows
BLOCK_MAX_ECDAS 2^12 at most 4096 ECDAS tables
BLOCK_ECDAS_MAX_VARS 17 each at most 2^17 rows
BLOCK_MAX_KECCAK 2^12 at most 4096 KECCAK tables
BLOCK_KECCAK_MAX_VARS 18 each at most 2^18 rows
BLOCK_MAX_ECSM 2^12 at most 4096 ECSM tables
BLOCK_ECSM_MAX_VARS 17 each at most 2^17 rows
ArgueFormat::BATCHED bins of 2^26 input cells change 5 below
LEAF_PERMS_CAP, BLOCK_FAN_IN 279,000 permutations; 3 the recursion tree's leaf cap and fan-in

1. The group partition is in the statement

What. The streamed prover packs tables into groups in the order their chunks complete. It writes the partition (per group, its tables' indices) into the statement. The transcript absorbs it before any root.

Verifier. The verifier checks:

  • that the partition is exact (every table once);
  • that each group of two or more tables fits within 3 stacked polynomials under the verifier's own stack cap;
  • that there are at most 64 groups.

It then rebuilds every group's stack layout itself.

Security. Proven bits are unchanged.

  • Each WHIR chain is 130.393 bits, and the proof minimum under the campaign's min-over-phases accounting is 128.946.
  • The chain heights are capped by BLOCK_MAX_TABLE_VARS (change 3), so no partition can make a chain taller than 2^27.

Proof bytes vary from run to run. Arrival order depends on thread timing, and the hash-ordered tables' row order (LT, EQ, BYTEWISE, BRANCH, MUL, DVRM) follows a hash state that is random per process. Mauro, 10-02: "the proof not having the same bytes it's fine, they never had the same bytes anyways".

  • Byte-identity tests run with LAMBDA_VM_FIXED_TRACE_HASH=1 (fixed hash keys) and LAMBDA_VM_DETERMINISTIC_GRIND=1. Under both, two processes prove the same bytes (FAST 423).
  • Fixed keys let a crafted program steer many operations into one bucket, so they are a test setting, not production's.

2. Prepared openings, one stack per group (approved by Mauro, 10-02: "it's a standard technique, go ahead")

In short. This is #1010's existing prepared DECODE opening, carried into the block format. It is extended to the dense genesis pages and stacked per group.

  • Costs: +0.14 s of base and −0.57 s of recursion (4.71 → 4.14 s).
  • The verifier recomputes every derived root from the ELF.

What. The recursion guest cannot evaluate DECODE's five 2^20 preprocessed columns, nor the dense genesis pages (millions of rows), on its own. So each group's prepared tables (DECODE and the dense genesis pages chosen by genesis_stack_plan) have their leading preprocessed columns stacked into one commitment.

  • Both sides derive this commitment from the ELF and the partition. The verifier recomputes every derived root from the ELF: verify_block_whir, and verify_block_tree's plan. The prover's own roots only shortcut the prover's side.
  • Its roots are absorbed after the groups' roots, before z.
  • It is opened on the group's fork after the group's own opening, each table's block at that table's point.

Negatives, on a guest with two dense genesis pages. Each is refused by the host verifier and by the recursion leaf, and admitted when the openings are skipped (the mutation):

  • two groups' openings swapped;
  • a block opened at another table's point.

Wrongly committed stacks are refused at the roots block: two pages swapped, a page holding another page's columns, another program's DECODE.

Cost. One stack per group costs +0.14 s of base (FAST 419; one commitment per table cost +0.49 s) and saves 0.57 s of recursion, by taking 90 k permutations out of the leaves. BlockFormat::prepared = false turns the openings off, for measurement only.

3. ECDAS split; table heights capped

What.

  • ECDAS is split into tables of at most 2^17 rows, as KECCAK_RND is at 2^16 (split_ecdas).
  • A scalar multiplication may straddle two tables. Its steps chain only through the Ecdas bus, keyed by the call's timestamp and the step's (round, op).
  • The frame (the host verifier and the tree plan) refuses:
    • a KECCAK_RND table over 2^16 rows;
    • an ECDAS table over 2^17 rows;
    • any table over 2^27 rows (REV-JUDGE G2). A 2^30 table would otherwise make a 2^30 chain at 127.39 bits.

Why. At a median block, one ECDAS table (2^19) argues on a 12 GiB tree. At p90 (2^20) it stacks into five polynomials on a 24 GiB tree.

Gates. FAST 534 and 535: the negatives, 3 mutations caught.

4. Recursion leaf: the leaf cap (G3) and the inverse share (W1)

  • G3. The tree plan refuses a group over LEAF_PERMS_CAP. With no fixed leaf count, it adds leaves while the heaviest leaf is over the cap. The block's own partition is unchanged.
  • W1, on by default. Each WHIR leaf publishes its bus share Σ p/q as p · ediv(1, q), which has no satisfying assignment for q = 0 whatever p is. The former ediv(p, q) left the share free at p = q = 0 (≈ 2^-160 under GKR soundness).
    • The leaf gains one XALU op per table, so every WHIR tree id changes. The base proof is untouched.
    • Gate: FAST 536. LFM_WHIR_SHARE_INVERSE=0 keeps the former form.

5. Batched argue, on by default since d9d0ac5 (BlockFormat::argue = ArgueFormat::BATCHED; Mauro, 10-02: "batch the constraints")

What. Each group's tables are argued together on the group's fork:

  • one lockstep LogUp-GKR ladder per bin, the tables packed first-fit-decreasing under 2^26 input cells;
  • one front-loaded constraint sumcheck per group;
  • no claim reduction for unshifted tables (every VM table).

Verifier side. The verifier derives the bins from the statement's shapes and its own cap; the variant is never read from the proof. The proof carries one BatchedArgue per group in BlockWhirProof.argues, and proof.tables is empty under this format. The recursion leaves verify the same format (lane i-batch2, N-4).

Security at the block's measured inputs (D-BATCH §3.2's formulas; re-checked at FAST 424's inputs: |T| ≤ 60 tables a group, ≤ 51 trees a bin, ≤ 2^28.87 input cells, bus messages ≤ 204 elements, N_C ≤ 413, D_max 4, n ≤ 21; L = 300):

term bits
LogUp fractional 147.22
GKR step batching 185.34
GKR sumcheck round 190.42
GKR line 186.33
zerocheck point 187.61
constraint β (N_C − 1 = 412) 183.31
claim batching 184.52
constraint rounds 190.00

Every term is above the WHIR phase, so the proof minimum is unchanged: 128.946 under the campaign accounting (130.393 for the WHIR fold under the calculator of record).

Measured.

  • FAST 348: phase-B argue −0.685 s, whole block −0.385 s (≈ −0.65 s attributable).
  • FAST 424 at the flip (d9d0ac5, A B B A, one binary): base 17.56 → 16.76 s (−0.80), whole block 21.85 → 20.89 s (−0.96), recursion −0.16 s. Every proof verified; determinism within each format 2/2.
  • Peaks: BIG 402/403 found none raised at 1.0 / 1.31 / 1.80×. BIG 395 at 1.80× with ops dropped: per-table 62.39 against batched 63.12 GiB (+0.72).

The knob. ArgueFormat::PerTable (one argument per table, the format before) stays measurable, with BLOCK_WHIR_ARGUE=per-table in the real-block tests. At the per-table format, the MultiProof, the prepared openings and the partition are digest-equal to the head before the batched merge (BIG 393).

6. KECCAK and ECSM split; their heights capped (the port of #1013's S0a; landed at 3cfffe2; approved by the lead under Mauro's any-block target, Mauro to confirm)

What.

  • KECCAK is split into tables of at most 2^18 rows and ECSM into tables of at most 2^17 (split_keccak, split_ecsm), as ECDAS is at 2^17.
  • A row of either is one whole call (a permutation, a scalar multiplication), so a cut falls between calls. A call reaches its rounds (KECCAK_RND), its steps and scalar bits (ECDAS) and its memory only through buses keyed by its timestamp.
  • The frame takes up to 2^12 tables of each and refuses a KECCAK table over 2^18 rows or an ECSM table over 2^17. Every other verifier keeps both to one table.

Why. A keccak-heavy block at the gas limit makes about 2^21 permutation calls. One KECCAK table that tall stacks into eight polynomials of 2^27, against a group budget of three. At its cap each table fits one polynomial.

Bytes. Every block measured so far has KECCAK ≤ 2^17 rows and ECSM ≤ 2^11, so it builds the same tables and the same proof (FAST 830: digest 07d1bd43… unchanged on block 25368371).

Gates. FAST 830: KECCAK and ECSM forced into 4 tables each on block 25368371 prove and verify, base and tree; the count and height negatives; mutation C caught.

One-binary A/B: the block tree vs #1010's epoch tree (FAST 416, noepoch/whir @ 454dad6)

Eight arms E B B E E B B E on the same binary (md5 checked after every arm):

whole block base last stage (E: root; B: top) whole − last stage harness verifies
E, #1010's epoch tree (n = 4) 31.80 s 22.48 s 1.05 s 30.75 s 0.92 s, in its timed path
B, the block tree (n = 4) 23.20 s 18.50 s 1.35 s 21.86 s 2.97 s, off the clock
B − E −8.60 s −3.98 s −8.89 s

Base, block 25368371 on FAST (each run against #1010's epoch base on the same binary)

run change block base
FAST 360 W1: non-streamed block proof 29.9 s
FAST 364 first streamed version 25.96 s
FAST 365 chunk generation off the walk thread 24.45 s
FAST 366 windowed executor 23.55 s
FAST 367 walker/accumulator split 22.68 s
FAST 368 walked windows kept whole, concatenated at finish 21.98 s
FAST 369 BITWISE counted per window, KECCAK_RND split by rows 20.73 s (epoch base 25.71 s, −4.98 s)
  • Phase B: 9.96 s. Host peak: 43 GiB.
  • Every block proof the windowed builder produced verified: 14 of 14 (FAST 363–369). Non-windowed: 6 of 6 (FAST 360–362).

Recursion (W3): proved and verified on block 25368371 (FAST 413)

  • Leaves are cut along groups: a group's opening covers all of its tables, so a group is the smallest unit a leaf can verify alone. The tree plan (lfm::whir_block::WhirBlockPlan) runs the host verifier's own statement checks, derives every shape from the AIR and every prepared root from the program, and never reads a proof. It prices each group in permutations and partitions the groups over the leaves (heaviest first, onto the least-loaded leaf).
  • A leaf:
  • Nodes are the STARK block's (block_node, shared byte for byte with STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013). They check the id, the state and the output equal across children, add the shares, and the top asserts zero.
  • Final check: verify_block_tree(elf, statement, top) derives the top program and checks the top proof against it. Every preset is pinned inside: the base's options, the block format, the tree's options, the leaf rule and the fan-in.
  • Tests:
    • every new check has a negative, and each negative a mutation showing that check is the one refusing;
    • the leaf's challenges equal the host's;
    • a tree over another partition is refused at the final check.
block 25368371 FAST 413: first build, serial FAST 414: pipelined (mean of 2 arms)
base (streamed, 9 groups, 27 chains, 12 prepared tables) 21.00 s 20.98 s
plan + leaf emission 0.23 s + 2.38 s, after the base inside the base (done at 12.2 s of 21.0)
leaves (3) 7.04 s, one at a time 3.36 s, three at a time
nodes 2 levels: 3.41 + 1.91 s one top over 3 children: 1.36 s
recursion 12.36 s 4.75 s
whole block 33.37 s 25.72 s
tree verify (derives every program from the ELF and the statement) accepted, 4.97 s accepted, 2.95 s

Every proof was verified by the harness off the clock. The base is unchanged by the emission running beside phase B (20.98 s vs 21.18 s in the arms without the change).

Next:

  • one prepared stack per group (group 8 carries 12 prepared chains);
  • the base levers: build_traces over segmented lists, and KECCAK_RND streamed.

Base levers after the A/B (FAST 417–422, each S/P on one binary)

lever base status
builder: windows concatenated in parallel −0.52 s (417), −0.21 s (419, replication in band) on by default
KECCAK_RND chunks streamed −0.14 s, no effect (this block's keccak work is at its end) off by default
LT's MEMW-derived ops streamed per window (FAST 420) +1.27 s, regression: the layout thread binds (9 more LT chunks, layout busy +2.47 s); LT heights deterministic across runs off by default
streamed chunks laid out on 3 threads, prepared columns after the groups (FAST 421, 422; BIG 390) −0.54 s (421), −0.22 s (422 replication); but the extra host memory scales with the block: +2.2 / +4.2 / +8.7 GiB max RSS at 1.0 / 1.3 / 1.8× (BIG 390) reverted at 11e2de5; on again since 59e9890, bounded (E3)
the rest of the run packed as it is laid out, AIR order (FAST 422) +0.21 s, regression: groups 6–8 still close together off

Memory (D-MEMORY, with i-mem / i-mem2)

stage effect status
M1: the builder drops each streamed chunk's ops as it leaves (drop_streamed_ops) FAST 501: 1× base −0.74 s, peak 47.4 → 36.6 GiB; the 1.20× block 25490321 now proves (49.2 GiB, OOM before). BIG 391 lean gate on the exact head: same statement shape kept vs dropped, both verify, 1.20× at 49.05 GiB on (f4aafea)
E1 lean walk + E4 paged executor memory (D-EXEC, lane i-exec) base −0.45 s (FAST 602); median windows phase 42.4 → 25.2 s (harness, FAST 599) on (5b4e6b4)
E1 v2: smaller per-step records (CPU op 96 B, MEMW_A ops as 48 B rows; D-EXEC, lane i-exec) byte-identical: proof digest equal to d9d0ac5's under the fixed trace hash + deterministic grind (FAST 426) on (773f134)
E3: layout workers bounded to K + 1 chunks unpacked (layout_ahead, D-EXEC) BIG 392: max RSS vs workers 0 +0.23 / −0.06 / +0.16 GiB at 1.0 / 1.3 / 1.8× (flat). FAST 425: no effect on time (whole 20.91 s both), since phase A no longer waits on layout on (3 workers since 59e9890)
one prepared stack per group prepared cost +0.49 → +0.14 s; recursion −0.57 s in
W: the rest laid out in ≤ 2 GiB waves; KECCAK_RND built as its 2^16-row tables at the finish (lane i-m4b) BIG 562: max RSS 88.88 → 79.54 GiB at 4.13×, 29.84 → 27.10 at 1×; FAST 852 + 853: base −0.14 s on (70eee3e)
b2: the finish packs each table as it builds it; phase A commits those tables narrow, never transposing them (lane i-m4b) BIG 563: max RSS 78.84 → 52.10 GiB at 4.13× (89.15 with W and b2 off), 28.53 → 22.45 at 1×; FAST 855 + 856: base −0.36 s on (365e3ab)

Narrow trace storage (M4, lane i-m4)

Prover-only; the proof bytes are the same. The proof digest equals the head's before M4 (07d1bd43…) under the fixed trace hash + deterministic grind (FAST 820).

  • What changes: between phase A's commit and phase B, each table of at least 2^16 cells is kept on the host packed at the width its values need, 1, 2, 4 or 8 bytes a column (BlockOptions::narrow, production Narrowing::CARD).
    • The pack runs on the card from the group's resident store, on a thread of its own, after the group's commit.
    • Phase B uploads the packed table and widens it on the card. A host reader widens on the host.
    • A wrong width map is refused by the kept-tree check (test plus mutation).
  • Time-neutral: whole block −0.04 s at 1×.

Max RSS, wide → narrow (BIG 560; the traces themselves shrink about 4×, e.g. 24.89 → 6.20 GiB at 1×):

block wide narrow
25368371 (1.0×) 36.23 GiB 29.47 GiB
25453112 (1.8×) 63.52 GiB 46.10 GiB
25482821 (2.66×) 80.80 GiB 52.83 GiB
25410821 (4.13×) — 86.93 GiB, proves and verifies

Before M4, #1014 ran out of memory at 3.03× (i-m4); the narrow line puts the 120.7 GiB edge near 5.5–6×.

Phase A uploads the next group beside the commit (3′, lane i-noepoch-w2)

Prover-only; the proof bytes are the same (digest 07d1bd43… at the default, FAST 832; the commits, their order and their bytes are unchanged).

  • What changes: phase A used to upload, commit and retire each group in turn, so the card waited on every group's upload. Now the current group's commit runs on a thread of its own while phase A's thread takes the next group and uploads its columns (BlockOptions::upload_ahead, on by default; BLOCK_WHIR_UPLOAD_AHEAD=0 is the control).
    • The upload starts only after the commit has asked the card for its room, so it takes what the ledger has left; a store the ledger refuses goes up after the commit, as before (no store or commit room was refused in any run).
  • Time (FAST 832, A B B A on one binary): base −0.60 s, phase A −0.66 s, whole block −0.61 s; 46 % of the upload seconds hidden. Groups 4–5 cannot hide theirs: the inline layout thread closes them after the previous commit ends, which is the next term (E3, next section: on since 59e9890).
  • Memory: one more group's columns held at the peak — the previous group's wide columns and its packed bytes wait for its pack's install while the current group commits (FAST 834, BLOCK MEM terms). At most one group (≤ ≈ 3.8 GiB), constant per group, not growing with the block; max RSS +1.8 to +2.3 GiB at 1× (FAST 832, 833, 834).

Streamed chunks laid out on three threads (E3 on by default, lane i-noepoch-w2)

Prover-only; the proof bytes are the same (digest 07d1bd43… at the default, FAST 838 and 835; the groups, their order and the packing are unchanged).

  • What changes: layout_workers 0 → 3 in BlockOptions::production (59e9890). The streamed chunks are laid out on three threads, with at most K + 1 = 3 of them unpacked at a time (layout_ahead, E3), instead of on the builder's one layout thread. BLOCK_WHIR_LAYOUT_WORKERS=0 (the inline layout, 01da99f's) is the control.
  • Why it pays now: with phase A uploading ahead (3′), the inline layout closed groups 4–5 only after the previous commit had ended, so the card waited on them; three workers close them in time.
  • Time (FAST 838, A B B A on two binaries, A = 01da99f, B = 59e9890): base −0.54 s (t −14.7), phase A −0.58 s, whole block −0.48 s (t −12.4; 20.23 → 19.75 s); phase B +0.06 s (one job, below what a single job resolves). FAST 835 (one binary, workers 0 vs 3) read whole −1.00 s.
  • Memory: max RSS −1.51 GiB at 1× (31.99 → 30.48 GiB, FAST 838; 835 read −0.68 GiB). The bound is what the earlier revert lacked: unbounded, the workers' host memory grew with the block (BIG 390); bounded, it stayed flat at 1.0 / 1.3 / 1.8× (BIG 392, measured before M4 and 3′, not re-run on this head).
  • 59e9890 also caps a readout stamp (an upload's "paid" seconds at the upload's own length); no proving change.

Phase B's kept-top paths re-hashed in parallel (lever 1, lane i-noepoch-w2)

Prover-only; the proof bytes are the same (digest 07d1bd43… at the default, FAST 840 at 733557a — 82f9046 adds only the knob; the paths, their order and the refusal are unchanged).

  • What changes: phase B serves each revived commitment's first-round paths from the kept tree top. The leaves under each queried block are gathered from the card, re-hashed on the host, and the block's subtree root is checked against the kept node. That re-hash ran one block at a time on the prover's thread while the card waited: 27 gaps of ≈ 37 ms, 1.01 s at 1× (nsys timeline, FAST 839). The blocks are now re-hashed in parallel and kept in block order; the root check still refuses before any path is returned. BLOCK_WHIR_REHASH_SERIAL=1 restores the serial re-hash (a measurement knob).
  • Time (FAST 843 + 844 pooled, 6 + 6 + 6 runs: one binary with the knob, serial S vs parallel P, plus 59e9890's binary as the reference R): re-hash 1.01 → 0.07 s; phase B −0.96 s (P − R; P − S −0.93 s, t −30.5); whole block −0.82 s (P − R, t −5.5; 19.72 → 18.90 s); phase A unchanged (this build vs R +0.09 s, t +1.0); max RSS unchanged.
  • FAST 840 (two binaries, 2 + 2 runs) read phase A +0.71 s, from one run that stalled in the layout; the code runs only in phase B, and the one-binary check shows no phase-A effect.
  • Also in: BLOCK OPEN SPLIT (with LAMBDA_VM_BASE_SPLIT=1), phase B's openings by host stage and the kept-top gather and re-hash seconds.

The tree's leaves execute while their artifacts are built (R2-i, lane i-noepoch-w2)

Prover-only; the proofs are the same (the top proof's digest equal with and without it, 0da6fea9…, under the fixed trace hash + deterministic grind, FAST 841).

  • What changes: WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014's tree is proved by the W3 harness (prove_tree_pipelined, the way a prover would run it; there is no production tree driver yet). Its level 0 built all three leaves' artifacts first (three serial card holds, 0.24 s at 1×) and only then let the leaves execute and fill their traces (0.46 s until the first was ready) while the card sat idle (nsys timeline, FAST 839). Execution and fill do not read the artifacts; only the prove does. lfm_prove is now cut into its host half (lfm_execute_and_fill) and its card half (LfmFilled::prove, which asserts the artifacts' hasher is the one the traces were filled for), in the same order; the harness builds the artifacts on a thread of their own and each leaf executes and fills beside them, waiting for its artifacts only to prove. W3_EXEC_BESIDE_ARTIFACTS=0 is the control (the order before).
  • Time (FAST 841 + 842 pooled, 8 + 8 runs on one binary): level 0 −0.24 s (t −17.8), recursion −0.23 s, whole block −0.24 s (t −5.2; 18.78 → 18.54 s; per job −0.29 / −0.20); base and max RSS unchanged.
  • The first leaf now takes the card 0.48 s after the tree starts, not 0.76 s (one run of each arm, FAST 841's card-hold trace).

A second whole-block field. The W3 readout prints, on the clock between the base and the tree, the block report, a second block frame for the LT heights and the group tables: 0.13–0.14 s at 1× that a prover would not spend. "Whole block" keeps its meaning (every number above includes them); from d169edd on, W3 RECURSION also prints "whole excl. harness readouts" beside it (at d169edd, FAST 841 + 842 pooled: whole 18.54 s, whole excl. harness readouts 18.40 s).

The rest laid out in waves; KECCAK_RND built as its tables (W, lane i-m4b)

Prover-only; the proof bytes are the same (digest 07d1bd43… under the fixed trace hash + deterministic grind: FAST 854 at 0bcc1fc, and FAST 855's W arm on this head's code).

  • What changes:
    • The tables the finish builds (the "rest") were turned into columns all at once, and a table's column copy exists before its rows are freed. So for a few seconds most of the rest was held twice, and WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014's memory log put the 1× and 4.13× peaks there (BIG 561). The rest is now laid out in AIR order, in waves of at most 2 GiB of rows, each wave in parallel (BlockOptions::rest_layout_bytes). The tables, their order and the groups are the same (test).
    • The finish built KECCAK_RND as one table of 1,480 columns, which split_keccak_rnd then copied into its 2^16-row tables. At 4.13× that copy was a second peak of the same height. The finish now builds the 2^16-row tables directly (BlockOptions::finish_keccak_rnd_chunks), the same tables as the split's (test), and hands none out during the windows, which would change the groups.
    • BLOCK_WHIR_REST_LAYOUT=all and BLOCK_WHIR_KR_FINISH_CHUNKS=0 restore the old behaviour.
  • Memory (BIG 562, one binary at 0bcc1fc, which 70eee3e merges with WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014's later commits; A = both off, W = this; three layout workers; jemalloc never purges, the record posture):
block A W
25368371 (1.0×) 29.84 GiB 27.10 GiB
25482821 (2.66×) — 49.02 GiB
25410821 (4.13×) 88.88 GiB 79.54 GiB
  • The live peak (jemalloc active) at 4.13× fell 72.9 → 63.7 GiB and now sits at the finish's end.
  • Not as pre-registered, on magnitude: the bands were ≤ 76 GiB and a ≥ 10 GiB drop at 4.13×, and ≤ 27 GiB at 1×. The resident-set gain came from the KECCAK_RND copy alone. The waves bound the live column copies, but the resident set still grows ≈ 16 GiB during the layout, because the freed row buffers stay with the allocator while the column copies take new pages. That growth and the finish's 8-byte tables are the next cut's target (finish packing, being measured).
  • On the max-RSS line, the 120.7 GiB edge moves from ≈ 5.9× to ≈ 6.4–6.7× (a linear estimate).
  • Time (FAST 852 + 853 pooled, one binary per job, three layout workers): base −0.14 s against both off, phase A −0.12 s.
    • The time comes from the finish ending without the KECCAK_RND copy.
    • The waves add ≈ 0.3 s to the rest's layout, which phase A does not wait on.
    • Packing the rest as it is laid out (pack_rest_as_laid_out) was re-measured with the waves: no gain (+0.01 s), so it stays off.
  • Also in:
    • the block memory log, LAMBDA_VM_BLOCK_MEMLOG=1 (off by default). It prints the host memory term by term every half second and at each phase mark, with the line at jemalloc's active peak and each arena's bytes (those two in the lib tests, which install jemalloc).
    • the readouts BLOCK REST TABLES (each rest table's layout end and group) and each group's commit start (from@).

The tree's first finished leaf executes during phase B (R2-ii, lane i-noepoch-w2)

Prover-only; the proofs are the same (base digest 07d1bd43… and the top proof's digest equal with and without it, 0da6fea9… — FAST 841's — under the fixed trace hash + deterministic grind: FAST 845, and FAST 847 again on fac261f, the merge with W).

  • What changes: phase B now hands each group's share of the proof (its argue, its opening, its prepared opening) to an observer as that group's opening ends (block_prove_on_forks_observed; the old entry point passes a no-op, and the proof is the same with any observer). A leaf's arena is built from its groups' words alone (group_arena_words + leaf_arena; block_leaf_arena is built from them). The W3 harness turns the groups into words as they arrive; the first leaf whose groups are all opened while another group is still to come (leaf 0 = groups 1, 5, 6 at 1×, done after group 6) executes and fills its traces on a thread of its own, about 1.5 s before the base ends, and the tree picks it up. Off the clock its arena is checked against the finished proof's. W3_LEAF_DURING_PHASE_B=0 is the control.
  • Time (FAST 845 + 846 pooled, 8 + 8 runs on one binary, at c5d9cec before the merge with W): whole block −0.24 s (18.63 → 18.39 s; per job −0.21 / −0.28); phase B +0.02 s; max RSS unchanged. The card's wait before the first leaf's prove fell from 0.23 s to 0.003 s; level 0 gained 0.15 s of it, since the first prove now runs beside the other two leaves' execution and fill.
  • Memory: the host's high-water is set in phase A at every size (1×: 25.5 GiB active at its peak, 7.2 GiB active by group 6's end with 31.0 GiB resident; 4.13×: 63.7 GiB at the finish's end, i-m4b), so one leaf's traces during phase B's tail fit in what phase A already holds.

The finish's tables packed as they are built (b2, lane i-m4b)

Prover-only; the proof bytes are the same (digest 07d1bd43… under the fixed trace hash + deterministic grind at both arms: FAST 855 at 0cebe16, and FAST 857 again on 365e3ab, the merge with R2-ii).

  • What changes:
    • After W, the finish still built its tables 8 bytes per cell, and phase A transposed each one into columns before packing it on the card. At 4.13× that layout took 24.6 GiB of wide tables and grew the resident set by ≈ 16 GiB (BIG 562).
    • The finish now packs each table into the narrow layout as soon as it is built (BlockOptions::pack_finished). KECCAK_RND is packed in waves of four 2^16-row tables. Phase A takes those tables as narrow columns and uploads them as they are, so the card packs only the groups' other tables. Nothing packed is transposed.
    • KECCAK, ECSM and ECDAS stay wide (BLOCK REST PACKED counts the rest: 187 of 191 tables at 4.13×).
    • A packed table's preprocessed columns are checked word for word against the AIR's before it is committed; a wrong one is refused (test, with a mutation).
    • BLOCK_WHIR_PACK_FINISHED=0 restores W.
  • Memory (BIG 563, one binary at 0cebe16; O = W and b2 off, A = W, B = W + b2; the memory log on; jemalloc never purges):
block O A (W) B (W + b2)
25368371 (1.0×) — 28.53 GiB 22.45 GiB
25482821 (2.66×) — — 37.64 GiB
25410821 (4.13×) 89.15 GiB 78.84 GiB 52.10 GiB
  • At 4.13× the live peak (jemalloc active) fell 63.7 → 49.7 GiB. The resident set now sits within 2 GiB of it (retention +2.0 GiB, from +14.7), and the rest's layout adds +1.15 GiB to it instead of +16.1. Every pre-registered row is in but one: the 1× live peak, at 21.0 GiB against a band of ≤ 19, because at 1× it falls in the middle of the finish, where the commit pipeline's in-flight copies bind.
  • On the max-RSS line (52.10 GiB at 4.13×, 12.82 + 9.47 GiB per 1×), the 120.7 GiB edge moves from ≈ 6.4–6.7× to ≈ 11.1–11.4× (a linear estimate for the base alone; the median block is 9.78×).
  • Next in line at the 4.13× peak: the committed groups held narrow for phase B (18.5 GiB, ≈ 5.9 GiB per 1×), then the finish's tables still being built (16.5 GiB).
  • Time (FAST 855 + 856 pooled, 8 + 8 runs on one binary): base −0.36 s (t −6.2; per job −0.40 / −0.32), all of it in phase A: the rest is laid out in 0.13 s instead of 0.57.

Measurement posture: the card's pool now retains freed memory (baseline shift)

Not a prover change; the proofs are the same (the top proof's digest 0da6fea9… under both postures, fixed trace hash + deterministic grind, FAST 848; 351a773 adds only a test readout).

  • What the env set, and why: every WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014 box run (the canonical env since FAST 416) set LAMBDA_VM_MEMPOOL_RELEASE_MB=0. The card's stream-ordered memory pool then hands its freed blocks back to the driver at each synchronize, so a VRAM sampler's total − free reads the live working set. The code default is to retain every freed block (DEFAULT_MEMPOOL_RELEASE_THRESHOLD_BYTES = u64::MAX in math-cuda), and that is what a prover runs.
  • What release-0 cost: each synchronize paid for the release, and the next large allocation mapped the memory again. In FAST 839's trace, the time spent inside cuStreamSynchronize and cuMemAllocAsync while the card was idle was ≈ 0.25 s in phase A, 0.42 s in phase B (0.26 s of it in the encode: one 18–33 ms gap per group) and 0.34 s in the tree.
  • Measured (FAST 848 + 849, one binary at 351a773, release-0 vs the default, 8 + 8 runs): whole block 17.91 → 17.27 s, −0.64 s (t −6.0). Phase B −0.44 s (t −28, against the 0.42 s above); phase A −0.07 s (t −0.7; phase A's runs scatter); recursion −0.12 s (level 0 −0.09 s); base −0.51 s; max RSS −0.16 GiB.
  • Gates, every run: the tree verified; the log's posture line matched its arm; no host fallback, device commit error, decline or ledger refusal; the card's peaks 24.15 GiB (the ledger's reservation) and 22.8 GiB (the pool's live high-water), read from the ledger and the pool, not from total − free (the new W3 DEVICE line).
  • From 351a773 on, WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014's numbers are taken at the code default. Every earlier number in this description was taken under release-0 and stays as measured. VRAM figures come from the ledger or the pool's high-water, never from total − free.

The prover refuses a partition over the group maximum (lane i-m4b)

Prover-only; the proof bytes are the same (digest 07d1bd43…, FAST 858). The prover now refuses a block over BlockFormat::max_groups as its groups close, with the verifier's own InvalidTableCounts error, instead of proving a block the verifier refuses. The median block (90 groups against the cap of 64) spent its whole ≈ 190 s base on such a proof (BIG 564). A test covers both prover paths, with a mutation for each check.

…d pair of its own

The per-table scheduler's driver threads are not rayon workers, so every one
of them staged through the shared slot-0 pinned slab. Uploads went one
single-buffered chunk at a time, and a driver copying a retained LDE out held
the slab's mutex through the whole host copy while the others queued with the
card idle. The slab also grew to the largest retained LDE's next power of two
and stayed pinned for the rest of the process.

By default the row-major commit's trace upload and its retained-LDE download
go through a pair of 32 MiB pinned buffers lent to that one transfer. That
covers both upload sites, the row-major expansion and the column-major
engine's.
- htod_staged: the host fills one buffer while the previous chunk's DMA drains
  the other. It returns once the last chunk is queued; each buffer's event
  guards it for the next borrower.
- dtoh_staged_into: two chunks in flight, each landed chunk copied straight
  into the Vec's spare capacity with no zero fill, and the in-place transpose
  queued right behind the last chunk's read.
- At most 8 pairs (512 MiB pinned, 16 allocations), made on demand and never
  grown or freed; a transfer beyond that waits for a pair.

LAMBDA_VM_STAGING_SHARED_SLAB=1 keeps the shared slab. The setting is read
once and named on stderr (`[gpu] transfer staging: ...`). Both tree drivers
print the staging counters after the base and at the end: bytes and host
seconds per path, pairs, waits, and the shared slabs' pinned footprint.

Measured on block 25368371 (FAST, RTX 5090), ABBA palindromes at 5cbdf06,
on the pre-engine base 169b668 behind a temporary knob. STARK: 121.30 ->
115.50 s (-5.80, A spread 1.20); the base 48.3 -> 44.0 s; under nsys the
base's card idle 9.89 -> 5.32 s; host peak -4.41 GiB, which is the slab:
4.01 GiB after the base against 0.13. WHIR: 107.60 -> 105.50 s (-2.10);
level 0 -1.3 s. Program identities unchanged. On this base the column-major
engine's upload takes the same pairs; that site is new here and not in the
measurement above.

Tests. Device: exact round trips across chunk boundaries, the closure
contract, the buffer-reuse hazard, 12 threads over 8 pairs, and root / host
LDE / handle parity of the base, ext3, split-tree and column-major engine
commits through either staging, with the staged path's bytes counted.
Card-free: the slab footprint and the staging line's shape.
Level 0 opened with host work alone: every first-round wrap's prologue at
once (reconstruct, emit, arenas, and on WHIR the epoch harvest), with
nothing on the card. None of it needs the card or the global proof, only the
ELF and the epoch proofs the base finished long before.

By default the tree drivers now start a lead-in before the base. Its helpers
(two by default) wait until the base reports its epoch count, then build the
prologues of wraps 0..want (want = level 0's first pool round) from copies of
the leading epoch proofs, with the same functions the pool calls, so the
programs are the pool's own. Level 0 takes each prologue instead of building
it. A prologue no helper started is built by the pool as before, and a
panicking one is handed back. Nothing in the lead-in takes the card permit
or holds device memory of its own. It uses the base's DECODE derivations as
the base shares them: the STARK commitment, and on WHIR the root and the
prepared opening, whose derivation is a device commit.

The base reports through an EpochObserver installed for the calling thread
(with_epoch_observer): the epoch count from the producer as soon as the
final epoch is executed, each proved epoch, and the DECODE work. That holds
on both preparation schedules and both pipelines; with no observer
installed, the pipeline is unchanged. Level 0 still takes its own DECODE
derivations from the base, and LFM_TREE_REDERIVE_DECODE=1 still re-derives
them there; the lead-in's copies are only for its prologues. BaseDecode now
holds the Arc the base shares. An I4 schedule test reads epoch positions
through the slice-based epoch_chain_position.

LFM_TREE_PROLOGUES_AT_LEVEL0=1 builds the prologues at level 0's start
instead; LFM_TREE_TAIL_PROLOGUES and LFM_TREE_TAIL_HELPERS size the lead-in.
The driver prints `L0 PROLOGUES: ...` either way, and at level 0's start how
many prologues were ready.

Measured on block 25368371 (FAST, RTX 5090), ABBA palindromes at 5cbdf06,
on 169b668 behind a temporary knob, before the base-prep and DECODE-handoff
changes. WHIR: 107.60 -> 104.40 s (-3.20, A spread 1.20); the lead-in
3.19 -> 0.07 s; level 0 -3.1 s; base unchanged; device peak +592 MiB, from
the harvest's MLE evaluations running unreserved beside the base's tail.
STARK: 121.30 -> 118.65 s (-2.65); the lead-in 5.54 -> 0.22 s and level 0
-6.1 s, but the base +3.4 s, because the two helpers slow the base's epoch
proofs. Program identities unchanged. On this base the DECODE handoff
already removes part of the lead-in, so the gain here is smaller than above.

Tests. Card-free: the hand-off (order, the count gate, handing back an
unstarted, failed or context-failed prologue, waiting on one in progress,
close). Fixture scale, on both bases: the observer sees the count once and
every epoch byte for byte, and a prologue built from the leading epochs emits
the pool's own program and arenas.
…te does

Every staged_transfers test forces its path with the thread override, and
only the block-scale tree drivers read the lead-in's setting, so no card-free
test read either setting from the environment. One test each now does, and
prints the line that names it:
- staged_transfers: staging_pairs_enabled() is !LAMBDA_VM_STAGING_SHARED_SLAB,
  with its `[gpu] transfer staging: ...` line;
- the tree tests: lead_in_enabled() is !LFM_TREE_PROLOGUES_AT_LEVEL0, with an
  `L0 PROLOGUES setting: ...` line.
The gate runs each with --nocapture under the default and under the opt-out
and counts the named lines.
…nned staging, level 0's first wrap prologues built in the base's tail

I6: each row-major commit transfer is staged through a pinned pair of its
own instead of the worker's shared slab; opt-out
LAMBDA_VM_STAGING_SHARED_SLAB=1, named once on stderr ([gpu] transfer
staging: ...). I7: level 0's first wrap prologues are built by helpers in
the base's tail; opt-out LFM_TREE_PROLOGUES_AT_LEVEL0=1, sized by
LFM_TREE_TAIL_PROLOGUES / LFM_TREE_TAIL_HELPERS (L0 PROLOGUES: ...). No
proof byte moves. Measured on the lane's base (IDLE-B box2): I6 STARK
-5.80 s, WHIR -2.10 s; I7 WHIR -3.20 s, STARK -2.65 s net.
…xes K3/K4/K5

Brings the HASH lane's six commits onto candidate C2 (7d41668): the
half-warp Merkle tops (K3), the work-queue grind (K4), the limb-multiply
permutation variants (K5), each still behind its LAMBDA_VM_GAP_* knob,
their parity tests and host KATs, the serialised grind-counter tests and
the queue grid's context fix. No conflict; the next commits make the
three fixes the defaults.
…imb permutation by default

The three RPX device fixes measured EFFECTIVE on the WHIR block (job 160,
wt300-307: K3 -3.90 s, K4 -6.45 s, K5 variant 5 -10.05 s against 107.15 s,
identities identical in every arm) are now the defaults. Each keeps an
opt-out, and its default lives in one constant in `rpx_paths`, so a
pipeline that needs one off flips one line:

- LAMBDA_VM_RPX_WARP_MERKLE (WARP_MERKLE_DEFAULT): narrow levels and the
  tail on rpx_merkle_level_warp / rpx_merkle_tail_warp; =0 walks a
  thread per parent (rpx_merkle_level, rpx_merkle_tail).
- LAMBDA_VM_RPX_GRIND_QUEUE (GRIND_QUEUE_DEFAULT): rpx_grind_search_queue
  on a card-filling grid; =0 is rpx_grind_search on LAMBDA_VM_GRIND_GRID.
- LAMBDA_VM_RPX_LIMB_PERMUTE (LIMB_PERMUTE_DEFAULT): every RPX kernel from
  rpx_v5.cubin (32-bit limb multiply, square_n unrolled by four); =0
  loads rpx_v0.cubin, the 64-bit multiply.

Each variable takes 0 or 1 (anything else aborts) and each switch prints
one line on first use, "[gpu] RPX Merkle: ...", "[gpu] RPX grind: ...",
"[gpu] RPX permutation: ...", naming the path and whether it came from
the pipeline default or the variable. The queue's grid line becomes
"[gpu] RPX grind queue: grid ...". build.rs now builds rpx.cu twice
(variants 0 and 5) instead of five times.

The LAMBDA_VM_GAP_* knobs and the gap_hash module are gone. The parity
tests move to prover/tests/rpx_device_paths.rs, named for what they
compare (the per-parent walk, the stride grind), and the temporary
wording leaves the kernels, the host KATs and the docs.
The PROVE SPLIT line's "r4_grind (n/airs on device)" took its delta of
gpu_lde::gpu_grind_calls(), the keccak arm's counter. An RPX grind that
runs on the device counts in gpu_grind_calls_rpx(), so under RPX every
line read 0/airs while every table ground on the card: on the WHIR block
(job 160) the root proof printed 0/11 beside the harness's own count of
11 RPX device grinds for it.

device_grinds_now() sums both arms and report() takes its delta of that.
rpx_grind_device gains a test that reads it around one RPX device grind
(the keccak-only count reads 0 there and fails).
The result lines still carried the campaign's fix ids (K3, K4, K5) and
called the old paths "shipped", which stopped meaning anything once the
new paths became the defaults. They now say what is compared: the warp
walk against the per-parent walk, the queue grind against the stride
grind, the limb primitives and the permutation variants. Output only;
every check is unchanged and both binaries still pass.
…er half-warp, a queue grind, the limb permutation

K3: a Merkle level and the tail compress one permutation per half-warp
(opt-out LAMBDA_VM_RPX_WARP_MERKLE=0). K4: the device grind claims nonces
from a work queue (LAMBDA_VM_RPX_GRIND_QUEUE=0). K5: the whole RPX module
runs the limb-multiply permutation variant (LAMBDA_VM_RPX_LIMB_PERMUTE=0).
Each opt-out accepts 0 or 1 and names itself once on first use. Every
digest, root and grind nonce search is byte-identical to the previous
kernels (host known-answer tests and device parity). The prove split now
counts RPX device grinds. Measured on the pre-K1 base: WHIR K3 -3.90 s,
K4 -6.45 s, K5 -10.05 s; STARK K3 -1.00 s, K4 -0.50 s, K5 -7.95 s.
The DEEP and out-of-domain denominators were inverted by a global
Montgomery scan: compute_denoms plus five scan kernels, each a full pass
over the domain with prefix and suffix scratch, although every row needs
only its own few inverses. By default now:

- compute_and_invert_denoms_ext3_dev runs one kernel,
  invert_denoms_rowwise_ext3_k{1..8}: each thread builds its row's
  denominators and inverts them in registers with one base-field
  inversion (adjugate over norm, the norms batched by Montgomery's
  trick, kernels/ext3_inv.cuh). More than 8 per row keep the scan.
- the fully resident R4 DEEP inverts its own row's 1 + K denominators
  (deep_composition_ext3_fused_m{1..4}) and needs no inverse buffer; it
  falls back to the buffered kernel above 3 points or on any
  precondition miss.
- the single-point OOD sums (the R3 composition parts: 1-2 columns, so
  1-2 blocks on the card) run on the row-chunked multi kernel, and the
  multi kernels' chunk count loses its 64 cap.

The values are the same field elements; raw limbs may differ by p, which
nothing downstream observes. Measured on block 25368371 on one RTX 5090,
ABBA behind a switch on 169b668: STARK (one_row=auto) -0.70 s whole
run against a 0.50 s A spread and -2.35 GiB device peak; WHIR kernels
-49.8 % with the wall inside the noise.

LAMBDA_VM_DEEP_INV_LEGACY=1 restores all three (the scan, the buffered
DEEP, one block per OOD column), read once per process with a banner,
as LAMBDA_VM_LDE_LEGACY does for the LDE. The legacy paths stay public
for tests/deep_inv_parity.rs, whose tests name both paths per call;
tests/deep_inv_setting.rs checks that the process setting is followed.
gpu_fused_deep_calls() counts the fused dispatch, and
cuda_path_integration asserts it follows the setting: a table that
silently fell back to the buffered kernel would still verify.
recompute_lde_produces_byte_identical_proofs compares the bytes of two
proves of one instance, one per residency mode, at the test options'
grinding factor of 1. Under `parallel` the CPU nonce search is rayon's
find_any (crypto::grinding::generate_nonce), so the two proves can
return different valid nonces; the nonce is absorbed before the queries
are drawn, and every opening after it moves. The test fails whenever the
two searches disagree, whichever residency mode runs: on a laptop it
failed 11/20 at d1dc455 and 15/20 at 7d41668, and 20/20 passed with
LAMBDA_VM_DETERMINISTIC_GRIND=1 (the smallest nonce) or without
`parallel` (a sequential find).

The residency tests now prove at grinding factor 0, as zf_golden_tests
already does for the same reason. The residency mode acts on the main
LDE, which the grind never reads, so the comparison loses nothing it
could catch.
…t field

A WHIR round has three proof-of-work slots (folding, out-of-domain, query),
and the proof carries three nonces a round whatever the bits. A slot whose
grind has zero bits, and the last round's out-of-domain slot, is carried and
never read, so any value in it verifies. A query-only grind (P2) would leave
two such unbound fields a round.

ChainFormat gains `nonces: NonceLayout`:
- Three (the default): today's format, byte for byte.
- Spent: a round carries only the nonces its grinds spend. The in-guest
  arena has no word for an unspent nonce; ChainShape::carries is the one
  place the layout is written, and the word count, the hints and the arena
  words all read it. The host verifier refuses a nonzero value in a host
  field the layout does not carry (Error::UnspentNonce). RoundNonces keeps
  its three fields, so both layouts share one proof type and Three keeps its
  bytes.

GrindBits::query_only(bits) grinds before the query positions only. The
query count reads the query grind alone, so it does not move.

No production config uses Spent or a query-only grind yet, so every proof,
program and pin is unchanged. The legacy layout's chain programs, arenas and
proof bytes are pinned against values printed at 0428c39, and the
transcript closed form now prices only the grinds a config spends.
Each WHIR base-chain round ground 20 bits before three challenges. Only the
query grind buys proven bits as placed: the folding grind sits before the
round's first sumcheck message, so the first folding challenge is redrawn by
varying that message at one hash a try, and the out-of-domain grind follows
the out-of-domain point. The new ZF lever `whir_grind` therefore defaults to
`query`: GrindBits::query_only(20) under NonceLayout::Spent, one grind and
one nonce word a round, 518 grinds a block instead of 1,472 at stack 27. The
query count reads the query grind alone and stays 112. The proven bits per
phase do not move: chain minimum 130.393 at stack 27, pipeline minimum
128.946.

LAMBDA_VM_ZF_WHIR_GRIND=all is the opt-out: GrindBits::uniform(20) under
NonceLayout::Three, the production config from before this commit. A test
pins it against a literal, and its chain programs against the values printed
at 0428c39. The banner gains `whir_grind=`. No univariate option reads the
lever, so no STARK proof, program or id moves.

Re-blessed: the production default chain's pins now describe the P2 chain
(12 grind permutations, 16,411 permutations, 32,590 arena words,
150,258 / 202,873 rows). Its previous pins move unchanged to the opt-out's
test. The banner strings in zf_format's tests gain the new key.
The ZfFormat::DEFAULT doc quoted a block timing for whir_grind=query from an
earlier measurement. A measured number in a comment goes stale, so the doc
now says what the lever does and why it loses no proven bits.

The same reasoning in whir_chain's module header, GrindBits::query_only and
WhirGrind gave the out-of-domain grind's reason as "it follows the
out-of-domain point". That is half of it: the grind sits right before the
batching challenge and does guard it. Dropping it costs nothing because the
batching challenge has far more bits than the target without any grind.

Comments only.
evaluate_many_base uploaded the rest of the point for its fold loop even when
nothing was left to fold, which asks the driver for a zero-byte allocation at
one variable. The loop and its uploads now run only from two variables up; at
two and more the same launches run in the same order.

The parity test gains the short shapes the argue's device-columns knob sends
to this path: 2^2 to 2^15 rows, up to 2,000 columns, and two chunks of the
evaluation budget at 2^15.
…DA_VM_ARGUE_DEVICE_COLUMNS)

The claim reduce evaluates every column of a table at its reduced point. The
batched device evaluation took a table only once each column was 2^16 rows
tall, a threshold set for columns that had to be uploaded first. The epoch's
columns are already resident, so for a resident table the host cost is its
cells: the widest precompile, 1,480 columns of 2^15 rows, was walked one
column after another on one thread, 263-303 ms a table, about 1.1 s of the
WHIR base's argue idle on block 25368371 (D-GFS, G2).

Under LAMBDA_VM_ARGUE_DEVICE_COLUMNS=1 a resident table goes to the card once
it holds 2^16 cells, and the columns left to the host are spread over the pool
when each is a host loop. Off by default; off is today's path, line for line.
The values are the columns' multilinear extensions at the point, so the proof
does not change.

- host_evaluate_calls() counts the columns walked on the host, beside
  evaluate_calls(); under LAMBDA_VM_BASE_SPLIT=1 every prove prints
  `ARGUE COLUMNS #k: on the card C · on the host H · xchecked X || device
  columns on|off · xcheck on|off`, the A/B's mechanism line.
- LAMBDA_VM_ARGUE_XCHECK=1 recomputes every card value on the host and fails
  the prove on a mismatch, naming the column. It is the block's identity gate:
  two proves of a block never share bytes (six table builders order rows by
  HashMap iteration), so the comparison has to be in-process.
- force_column_value_fault arms a one-cell corruption so the checks can be
  shown to fail. Never armed outside a test.

Tests: the host walk proves the same with the knob on (unit); on a device,
card against host at 2^1..2^15 rows and up to 2,000 columns, resident and
uploaded; the reduce's proof, point and transcript knob off/on with the path
each arm took; the threshold; and the fault seen by the identity, the verifier
and the cross-check (tests/argue_device_columns.rs, box only).
…he card

The CPU, ADD and MUL fixture at a height where LAMBDA_VM_ARGUE_DEVICE_COLUMNS
has work to move: CPU 2^14 x 5 and ADD 2^14 x 4 (2^16 cells, the threshold)
go to the card once resident, MUL 2^13 x 4 stays on the host. multi_prove with
the knob off and on must give the same bincode bytes, table by table and whole,
the same transcript state and next challenge, and both proofs verify. On a
device each arm is shown to have taken its own path (9 columns on the card
with the knob on, none off); without one, the host walk is compared with the
same walk spread over the pool.

The negative control arms the one-cell fault: the identity must then name
CPU, the first table the card values, and the proof must fail the claim
reduce's check. It needs a device and says SKIPPED without one.
…r's copies

A table statement took its preprocessed count from the column copies the
verifier held. An AIR that declares its preprocessed columns by count only
(`with_preprocessed`, as every LFM chip does) has no column builder, so its
statement counted zero: with no prepared opening its program columns were
bound by nothing, and with one the honest proof was refused
(`settled > preprocessed.len()`).

`TableStatement` now carries `num_preprocessed`. `statement_with_preprocessed`
sets it to the number of copies, so every existing caller is unchanged and no
proof byte moves. The new `statement_with_prepared_prefix(count)` holds no
copies and takes the AIR's count; `check_preprocessed` refuses a prefix that
is neither settled by a prepared opening nor recomputed, and `multi_verify`
bounds the settled count by `num_preprocessed`.

Tests on the three-table fixture, with the CPU table's first three columns
declared preprocessed and its reversed trace as a forged program that still
balances the bus: the forgery verifies under a count-zero statement (the
control), and is refused with the count and no opening, with an opening over
the pinned columns, and with an opening that settles only part of the prefix;
an honest proof with its opening verifies.
Pure WHIR, D-WHIR §2: a recursion proof (wrap, node, root) can be proved by
the base's multilinear prover instead of one STARK per table. The LFM chips
are AirWithBuses with a single-source constraint IR and bus interactions, so
this is `multilinear_table::multi_prove` over the program's tables plus glue:

- the prepared stack: every table's preprocessed prefix (the instruction
  groups) committed once per program as one stacked WHIR commitment and opened
  at each table's reduced point — DECODE's mechanism over several tables;
- the statement (tag, program_id_w, version, word count, each public word as
  four felts, the heights, the chain config, the pad), absorbed before any
  challenge; indices are implicit and a claim out of position is refused;
- program_id_w, folding the prepared roots, the heights and counts, the hasher,
  the chip set and the WHIR format.

Every statement is built with `statement_with_prepared_prefix` at the AIR's
own preprocessed count, and a prepared plan that does not settle exactly
`0..count` of every table is refused before any argument runs (the §2.4 trap).
The hash is RPX by type (`RpxWhir`, `DefaultTranscript<E, RpxTranscriptHash>`),
never the `LAMBDA_VM_WHIR_HASH` knob. Policy A: one group, the prefix committed
in the main stack and in the prepared stack. The chip set must be the WHIR
recursion one (no keccak, BLAKE3 or BITWISE); an LFM_HASH split is one more
table.

`LAMBDA_VM_LFM_PROVER=stark|whir` selects the prover (default stark, bannered
on every setting, an unknown value aborts); nothing reads it yet, so every
proof is today's.

Tests: TrivialV0 round-trips at the production config (grinding on) and the
registry programs round-trip or are refused by chip set; a forged program (one
constant changed, same shape) verifies under count-zero statements — the trap —
and is refused by the W-LFM verifier without an opening and with an opening
over the honest stack; a deleted opening, a plan missing a table, a restated
height, a tampered or reordered public word are refused; a KAT against an
independent transcript shows the prepared roots are absorbed before z. Tests
other than the anchor run ungrinded through a cfg(test) switch in the one
config derivation.
D-WHIR §3: the wrap's `whir_epoch_program` body as a node leg, with the W-LFM
statement over the child's hinted public words, the child's prepared-stack
roots interned in the roots block, every table's plan `settled = count,
route None`, the closure against the claimed words' LfmPublic balance, the
main group walk, and the prepared opening last with each table's prefix at
its own reduced point. The roots block, table walk, group walk and stacked
opening are the wrap's own emitters.

The public words are hinted as four felts each: the statement's Pack rows
read every lane as a base token, so a hinted word with nonzero upper lanes
has no satisfying assignment. The plan is `WhirLfmPlan::build`, which refuses
at emit time a prepared plan that does not settle every AIR prefix exactly.

`emit_whir_node` is `emit_node` with W-legs (bindings and publishes read only
the legs' lanes); `emit_node` and every STARK emitter are untouched. The cost
form `whir_leg_cost` composes the landed forms, including a shape-only count
of each table's proof words, and `shape_only_artifacts` evaluates it at a
program shape with no commitment.

Tests: the leg executes an honest TrivialV0 child at the production config and
publishes the host's (z, alpha); a mutated published felt, carried root, bus
output, GKR word, column value, main-chain and prepared-chain final value are
each refused; the forged program is refused; a plan missing a table is refused
at emit time; F1 is exact by kind on a real child (ops 214,566, constants 150
both ways, hints 42,349, permutations 17,203) and each table's shape word
count equals its arena length; pins at the D-WHIR §3.2 shapes (wrap 0: 39,795
permutations, 398,742 ops, 75,432 hints), within three units of the design's
instrument. The leg program proved as a W-LFM proof is box-scale (16.9 M
cells, stack n25) and is ignored here.
shift_evals(x, k)[y] = eq_evals(x)[(y - k) mod 2^n]: for a corner y, shift_k(x, y)
is the indicator of x = y - k extended multilinearly in x, which is eq(x, y - k).
The claim reduce's tables on the card (LAMBDA_VM_ARGUE_DEVICE_TABLES) build
eq(alpha) once and read each offset's shift table as a rotated copy, so the
carry recursion must agree with the rotation value for value: every offset of
the small cubes, offsets past the cube, base field and extension.
…ix out

Policy B of a W-LFM proof (D-WHIR §2.4) commits each table's preprocessed
prefix only in the prepared stack: the main stack holds the value columns,
and the prefix's claims are settled by the prepared opening alone — the same
binding, since that opening already settles them under policy A.

- stacked_eval: `ColumnsAt` names a group's columns in a resident store,
  `From(first)` (today's contiguous run) or `Map` (one store column each);
  `commit_mapped` / `prove_mapped` take it, and `commit` / `prove` delegate
  with `From`, so every existing caller is unchanged.
- multilinear_table: `CommittedTables::commit_grouped_settled` stacks each
  table's columns past its settled prefix, reading them out of the resident
  store (which keeps every table contiguous for its argument) through a
  column map; `commit_grouped` is it with no prefix. `multi_prove` refuses,
  before any transcript work, a proof whose prepared opening does not settle
  exactly the prefixes left out, and opens each group on its own columns'
  claims. `multi_verify_settled` takes the exclusion as the verifier's format
  choice; `multi_verify` is it with none.

Tests: policy B round-trips on the three-table fixture; the forged prefix is
refused under B; a B proof read as A and an A proof read as B are refused; a
prover whose opening does not match the exclusion is refused.
`PrepPolicy` (A: `Both`, B: `PreparedOnly`) is a W-LFM format choice stamped
on the artifacts and folded into program_id_w; `LAMBDA_VM_LFM_WHIR_PREP=
both|prepared` selects it (default both, bannered, an unknown value aborts),
and `build_whir_artifacts_under` takes it explicitly so one process builds a
program both ways. The prepared stack does not depend on it.

Under B the plan's main shapes are the value columns, the prover commits with
`commit_grouped_settled`, the host verifier runs `multi_verify_settled`, and
the W-leg's group walk opens each table's columns past its prefix (the
prepared opening still takes the full walk). The cost form follows the main
shapes.

Tests: TrivialV0 round-trips under B, a B proof read as A and an A proof read
as B are refused, the forged program is refused under B on the host and in
the leg, F1 is exact under B (ops 200,986, constants 149, hints 39,995,
permutations 15,894), and the §3.2 pins under B (wrap 0: 38,440 permutations,
384,825 ops, 73,961 hints; the design's instrument had 38,441 / 384,823 /
73,961).
Under LAMBDA_VM_BASE_SPLIT=1 the W-LFM prover pushes the base's split record
(prep, absorb, commit, prove, wall, and the inner challenge / argue / open
slots multi_prove accumulates) under a new index, `LFM_INDEX`, which the line
names `W-LFM` — so a recursion prove prints `WHIR PROVE SPLIT W-LFM` and no
reader of the base's table can take it for an epoch. Stages are timed
directly rather than through `stage_done`, whose `BASE EPOCH` line would be
misread. Nothing changes with the variable unset.
`w1_one_wrap_and_one_node_proved_both_ways`, ignored (box tier, cuda, the
production block): proves the RPX base, then
- W1a: wraps 0..3 under STARK, wrap 0 under W-LFM policy A and B, wraps 1..3
  under policy B;
- W1b: L1N0 (arity 3 over the STARK wraps, as the tree emits it) under STARK
  and W-LFM A and B;
- W1c: the pure-WHIR L1 node emitted over the policy-B WHIR wraps with
  W-legs, proved under W-LFM A and B, and the W-leg cost form per child;
- W1d: the census of an L2-shaped parent over three STARK L1N0 children and
  over three W1c children.

Each prove prints one `W1 PROVE` line: build, prove and its split, host verify,
proof bytes, host peak, the stacks (W-LFM) and the identity. The STARK wraps'
and L1N0's identities are the default's byte gate against the record's tree
at the base sha (wt800).
The W1 test held every wrap's full W-LFM build, prepared-stack codewords
included (device-resident), while only the artifacts and the proof feed the
pure-WHIR node and the parent census. `w1_prove_whir` now returns the
artifacts and drops the prover's stack after the host verify.
The security gate reads each chain's shape (num_vars, blowup, fold schedule,
caps, queries, grinds) from GAPB CHAIN lines, and the census instrument that
prints them (a9ffbf48a) is not on this branch. W1 prints the same line, from
the artifacts' config, for every main and prepared polynomial a W-LFM proof
opens, tagged with the proof it belongs to.
Pieces for LAMBDA_VM_ARGUE_DEVICE_TABLES; nothing calls them until the
multilinear side does, and every existing path is unchanged.

- Extra and DeviceFactors::session_with: a weight a resident sumcheck adds is
  a table to upload, as before (session delegates, byte for byte), or the
  point of an eq table, built into its slab with eq_expand_into.
- reduce_session: the claim reduce's session with its tables built from the
  epoch's resident columns. eq(alpha) is built once; each offset's shift table
  is two device-to-device copies of it, since shift_k(alpha, y) is
  eq(alpha, y - k) at a corner (pinned by multilinear's
  a_shift_table_is_the_eq_table_rotated). Each offset's batched column is the
  new batched_column_ext3 kernel over the resident run: one thread per row,
  base x ext3 per member, the host's sum in another order.
- SumcheckSession::factor and set_cell: read one factor, write one cell, for
  the cross-check and the fault hooks.
…AMBDA_VM_ARGUE_DEVICE_TABLES)

The WHIR base's ARGUE builds two families of challenge-dependent tables on
the host, on the rayon pool, then uploads them pageable: the zerocheck's
weights eq(r) and eq(row), and the claim reduce's shift tables and batched
columns. On the head's trace the ARGUE thread waits 3.56 s in those joins
(the producer's prep holds the pool) and the uploads are 20 GB a block,
1.16 s copy-only (I-GFS.md §6).

Under LAMBDA_VM_ARGUE_DEVICE_TABLES=1 they are built where they are folded,
from a few kilobytes: the points, the columns each offset reads, and their
weights. Off by default; off is today's path.

- A2: batch::Weight is a table or the point of an eq table. multi_prove passes
  the zerocheck's two weights as points under the knob; prove_resident keeps a
  point until the card builds it with its rounds, or builds it on the host if
  the card turns them down. prove_core takes Weights; prove_statements wraps
  its tables.
- A3: claim_reduce::prove asks gpu::prove_reduce_resident first. Its gates are
  prove_sumcheck's over the same tables, plus a resident run; its rounds stop
  at the same crossover and the host finishes over the folded tables through
  the same loop. A decline is before the transcript moves.

The values are the host's in exact arithmetic, so the proof is too. Under
LAMBDA_VM_ARGUE_XCHECK every table the card built is compared with the host's
before the first round and a mismatch fails the prove, naming it;
force_table_fault arms a one-cell corruption so that check, the identity and
the verifier can be shown to see one. Under LAMBDA_VM_BASE_SPLIT each prove
prints `ARGUE TABLES #k: built on the card N · xchecked X || device tables
on|off · xcheck on|off`.

Tests: the host arm (a point proves what its table proves); on a device
(tests/argue_device_tables.rs, box only), the reduce tables and weights
against the host at many shapes and offsets, and a prove_core knob matrix
with shifted reads, closures and nothing resident, verified, with the path
each arm took and the negative controls; in stark, the tall fixture's whole
argument identical with tables on, alone and with the columns knob, and its
negative control.
SumcheckSession::values makes one synchronous device-to-host copy per factor.
values_gathered reads the same bytes with one launch of the new
gather_factor_heads_ext3 kernel (each factor's first len cells, through the
session's pointer table, into one buffer) and one copy back. Nothing calls it
until multilinear does, under LAMBDA_VM_ARGUE_LEAN_READS.

The parity test compares the two reads bit for bit after rounds and folds, for
uploaded sessions and sessions over resident factors with an eq weight and a
table added, from 1 factor to 1,482.
Production lays the streamed chunks out on 3 workers since 59e9890; two
test harness comments still described 0 (the inline layout) as production.
Comments only.
…kes them narrow (cut b2)

After b1 the 4.13x high-water is the finish itself (BIG 561: at p5's end the
finish holds its tables, 24.6 GiB at eight bytes a cell, beside held_narrow and
its op lists). b2 keeps those tables packed from the moment each is generated.

- multilinear: NarrowColumns::pack_row_major packs a row-major trace column by
  column (#1013's NarrowMain::pack, the same layout), and
  TraceData::new_narrow holds packed columns from the start.
- stark: TraceTable::pack_main_narrow / narrow_main / take_narrow_main (the
  packed main trace, no 64-bit copy kept); CommittedTable::from_narrow (a
  layout whose committed columns are the main columns in order).
- trace builder: StreamSkip::pack and WindowedTraceBuilder::pack_finished_tables:
  each table the finish generates is packed as it is made, chunk by chunk.
  KECCAK, ECSM, ECDAS and an unchunked KECCAK_RND stay wide (the block cuts
  them after the build).
- block whir: BlockOptions::pack_finished (production on,
  BLOCK_WHIR_PACK_FINISHED=0|1): a packed table is laid out narrow
  (table_of_narrow: the packed columns become the table's, its preprocessed
  columns checked word for word against the program's), with no transposition
  and no wide column made. The BLOCK REST PACKED line counts them.
- phase A: a group holding a narrow table is uploaded packed and widened on the
  card (upload_group), the commit reads handles (a host fallback widens), and
  M4's packer skips what is narrow already. The memory log counts held bytes.

The tables, groups and proof bytes are the same: a test proves
test_keccak_multi with b1 and b2 on and off, inline and on three layout
workers, and compares the partition, the statement and (deterministic grind)
the bytes; a packed table whose first preprocessed column differs in one word
is refused (mutation: with the comparison off, that test fails).
With the finish packing its tables (b2), each KECCAK_RND chunk's pack is serial
on the task that built it, and one task building and packing every chunk in
turn would make KECCAK_RND the finish's last table by more than it already is
(BIG 561: gen_keccak_rnds ends 4.6 s after LT at 4.13x). Packed chunks are now
built and packed in parallel waves of four, at most four of them (about
3 GiB) wide at once; the tables and their order are the same, and the unpacked
path is unchanged.
…es are laid out

#1014 now lays the streamed chunks out on three workers (0498480), and phase
A waits on the groups the rest closes, which close only once the whole rest
is laid out. With b1's waves the AIR order's prefix is laid out first, so
pack_rest_as_laid_out (the rest's sink, groups identical by test) is now on in
production: each wave's tables go to the packer as the wave ends, and a group
closes once its own tables are laid out. Laid out all at once, the first table
in AIR order was ready only with the last, which is why packing as laid out
read +0.21 s before (FAST 422).

The real-block tests' BLOCK_WHIR_PACK_REST=0|1 now overrides the production
choice instead of replacing it with off.

Readouts for the finish's place in phase A: each group's commit start
(from@, beside done@), and every rest table's layout end and group in AIR order
(BLOCK REST TABLES).
…r and re-hash

For the timeline of #1014's phase B: with LAMBDA_VM_BASE_SPLIT=1 the block
records, over phase B, the WHIR chains' six host stages (grind, sumcheck,
fold, commit_folded, ood, queries), the round wall, the query split (sample,
tree rebuild, coset gather, assemble), and the seconds the first-round paths
from the kept tree tops spend gathering their blocks and re-hashing them on
the host (two new counters in whir_commit). BLOCK OPEN SPLIT prints them.
Readout only: no schedule or proof byte changes.
…arallel

A revived commitment opens its first-round paths from the kept tree top: the
leaves under each queried block are gathered and re-hashed on the host, and
each block's subtree root is checked against the kept node. That loop ran one
block at a time on the prover's thread while the card waited: on #1014's 1x
block it is 27 gaps of about 37 ms in phase B's openings, 1.01 s in all (nsys
timeline, FAST 839), with nothing else on the host.

Each block's subtree depends only on its own leaves, so the blocks are now
re-hashed in parallel. The subtrees stay in block order, the paths are
assembled from them as before, and the first block in that order whose root
is not the kept node is the one refused, as the serial walk would. The proof
bytes do not change.

A new test opens a tall stack with 48 queries (tens of distinct blocks a
round) after retire and revive at drops 3 and 4: the opening equals the kept
one's byte for byte, and a revive over other columns still refuses.
… at the finish)

FAST 852 + 853 at three layout workers, 8 + 8 arms against 59e9890's
behaviour: b1 with the sink −0.124 s, b1's waves without it −0.135 s, the sink
alone +0.21 s. The groups do close earlier with the sink, but at 1x the
committer is still busy when they do. pack_rest_as_laid_out goes back to off;
the readouts and the override stay.
BLOCK_WHIR_REHASH_SERIAL=1 makes the kept-top paths re-hash their queried
blocks one at a time, as before the parallel re-hash, in a parallel build.
FAST 840 measured the parallel re-hash on two binaries: phase B -0.93 s, but
phase A +0.71 s, where the change never runs. The knob lets one binary carry
both arms, so a build effect and a real one can be told apart. The bytes are
the same either way (the stacked_eval byte tests pass with it set); read once
per process.
…d the top's digest

The W3 harness prints several readouts on the clock between the base and the
tree (the block report, a second block frame for the LT heights, the group
tables): 0.13 s at 1x that a prover would not spend (FAST 839). The whole
block keeps its meaning, since every earlier W3 number includes them; a second
field beside it now gives the whole block without them.

It also prints a blake3 digest of the top proof's bytes, so two runs under
LAMBDA_VM_FIXED_TRACE_HASH=1 and LAMBDA_VM_DETERMINISTIC_GRIND=1 can be
compared byte for byte (off the clock).
…built

The W3 harness proves #1014's tree the way a prover would run it. Its level 0
built all three leaves' artifacts first (serial card holds, 0.24 s at 1x) and
only then let the leaves execute and fill their traces (0.46 s until the
first was ready), with the card idle for the second part (nsys timeline,
FAST 839). Execution and fill do not read the artifacts; only the prove does.

lfm_prove_with_residency is cut into its two halves, unchanged in order:
lfm_execute_and_fill (the host phase) and LfmFilled::prove (the card phase,
asserting the artifacts' hasher is the one the traces were filled for). The
harness now builds the leaves' artifacts on a thread of their own and each
leaf executes and fills beside them, waiting for its artifacts only to prove.
W3_EXEC_BESIDE_ARTIFACTS=0 is the control (every leaf waits for every
artifact first, the order before). The proofs are the same.
…rver

block_prove_on_forks_observed calls an observer after each group's opening
with that group's share of the proof (GroupOpened: the roots, its batched
argue or its tables' proofs, its opening, its prepared opening).
block_prove_on_forks delegates with a no-op, and the streamed block prove
takes the observer beside its statement observer
(prove_block_whir_observed_groups). The observer only reads: the proof is the
same with any observer. It lets the block tree's first leaf, whose groups
are all opened before phase B ends, start before the base is done.
block_leaf_arena is split into group_arena_words (one group's words from its
own share of the proof: its tables' proofs or batched argue, its opening, its
prepared opening) and leaf_arena (the block's roots, then the leaf's groups in
order). The arena is the same; the shape checks and their refusals stay,
with a group's prepared opening now required to match the plan's one
prepared stack for that group.
The W3 harness now listens to phase B's groups as they open. The first leaf
whose groups are all opened while another group is still to come builds its
arena from them and executes and fills its traces on a thread of its own; the
tree's worker for that leaf joins it instead of executing after the base.
After R2-i the card still idled about 0.23 s at 1x before the first leaf's
prove, waiting for that leaf's execution and fill (FAST 841 + 842).
W3_LEAF_DURING_PHASE_B=0 is the control.

Off the clock, the early leaf's arena is checked against the finished proof's
(W3 EARLY LEAF), and a box-tier test proves small streamed blocks under both
argue formats and checks that the groups phase B hands over build every
leaf's arena.
…s its tables, the block memory log) into nw2/r2ii

No conflicts. multilinear_block.rs: the memory log's phase-B marks and the
group observer sit side by side at the end of each group (the observer first,
then the group's columns are released and the mark taken). block_whir.rs and
the W3 harness: the new BlockOptions fields and the memory log's plumbing
next to the group observer's.
… proof to an observer; the block tree's first leaf executes during phase B) into m4b/b2

No conflict. R2-ii's two full BlockOptions literals in lfm/whir_block_tests.rs
already carry b2's pack_finished: true, so its hand-over tests run with the
finish's tables packed, as production does.
… the ledger and the pool count them

Off the clock, after the tree: the commit/encode and opening host fallbacks,
device commit errors, argue device fallbacks, GKR tree refusals, fused
declines, the reservation high-water and the memory pool's live high-water.
A run that changes the pool's release posture needs the card's peak from the
ledger and the pool, not from total - free, which reads a retaining pool's
freed blocks as used.
…its groups close

The verifier refuses a statement with more groups than BlockFormat::max_groups
(block_frame), but the prover never checked: on the median block (90 groups
against 64) it spent its whole base on a proof the verifier refuses (BIG 564).

One check, check_group_count, now serves the verifier and both prover paths:
the streamed packer refuses the group past the maximum as it closes, and the
whole-run path refuses after sizing its groups. The error is the verifier's,
word for word. Prover-only: a block within the maximum proves as before.
MauroToscano added a commit that referenced this pull request Oct 3, 2026
…14fea1

crypto/stark/src/narrow.rs and crypto/stark/src/spill.rs are #1013's files,
byte-identical to mem/stark-f1 @ e5214fea1 (i-mem3's F1: the page-backed
Bytes, the streaming Digester, the fused writer, over the S1/S2 store). #1013
is their one source of truth: #1014 never edits them, re-imports any change
from there, and they merge as one when both PRs reach main.

Wiring only: the two modules in lib.rs (dead code allowed there, since some
parts serve #1013's prover alone), math's page-bytes feature, and libc
unconditional as the store needs it (it was behind disk-spill). #1013's prover
glue (TraceTable::spill_main) and its spill tests are not imported; #1014's
adapter and tests follow.
MauroToscano added a commit that referenced this pull request Oct 3, 2026
#1013 landed 4cb0a07 (the writer's digest / copy / pwrite step timings on
SpillStats, SpilledMain::is_resident). crypto/stark/src/spill.rs is again
byte-identical to noepoch/stark @ edddc68; narrow.rs was already. Readouts
only: nothing #1014 calls changed.
MauroToscano added a commit that referenced this pull request Oct 3, 2026
… reserve

LAMBDA_VM_BLOCK_SPILL now reads as #1013's (prover/src/block.rs SpillPolicy …
spill_decision @ edddc68): auto (and unset) | off | always | <GiB> resident
budget. auto spills a committed table once the host's bytes (the larger of
VmHWM and the cgroup's memory.current), the reserve and the table would pass
the target (LAMBDA_VM_BLOCK_SPILL_TARGET_GIB, else the cgroup's memory.max
less 10 GiB, else MemTotal less 10 GiB). Phase A decides per table and never
past the writers' queue; groups 0 and 1 count but never spill.

The reserve is #1014's own, inferred from the median block (BIG 565): the
finish's p5 transient (25.5 GiB) and the tree's leaf programs alive through
phase B (16.3 GiB), about 1.25 GiB per total G cells or 1.65 per G cells
committed so far, plus 6 GiB for phase B's bump and the read-back window.

BlockSpill carries the policy's choice (SpillWanted) and the kept bytes and
committed cells it reads. A block that fits spills nothing; the store opens
for every policy but off. Tests: the knob and the decision (unit), a budget
the block fits in spills nothing and one of zero spills (block).
MauroToscano added a commit that referenced this pull request Oct 3, 2026
…@ 278e6a8

The rule #1014 copies from #1013 read only cgroup v2's memory.max and
memory.current. On a v1 host (FAST) the target fell through to MemTotal
less 10 GiB, and the host's charge was not read at all. #1013 fixed both
in 278e6a8; this re-copies spill_target_bytes, spill_target_from,
cgroup_memory and host_bytes_now from that sha, with its two tests (the
cgroup files from fake trees, and the target rule), cited as before.

The only changes from #1013's text are the citations and the tests'
temporary directory, which gets a name of its own so the two copies of
the test never share one if they ever run in one process.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant