WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress - #1014
Draft
MauroToscano wants to merge 1306 commits into
Draft
MauroToscano wants to merge 1306 commits into
MauroToscano wants to merge 1306 commits into
Conversation
…d pair of its own The per-table scheduler's driver threads are not rayon workers, so every one of them staged through the shared slot-0 pinned slab. Uploads went one single-buffered chunk at a time, and a driver copying a retained LDE out held the slab's mutex through the whole host copy while the others queued with the card idle. The slab also grew to the largest retained LDE's next power of two and stayed pinned for the rest of the process. By default the row-major commit's trace upload and its retained-LDE download go through a pair of 32 MiB pinned buffers lent to that one transfer. That covers both upload sites, the row-major expansion and the column-major engine's. - htod_staged: the host fills one buffer while the previous chunk's DMA drains the other. It returns once the last chunk is queued; each buffer's event guards it for the next borrower. - dtoh_staged_into: two chunks in flight, each landed chunk copied straight into the Vec's spare capacity with no zero fill, and the in-place transpose queued right behind the last chunk's read. - At most 8 pairs (512 MiB pinned, 16 allocations), made on demand and never grown or freed; a transfer beyond that waits for a pair. LAMBDA_VM_STAGING_SHARED_SLAB=1 keeps the shared slab. The setting is read once and named on stderr (`[gpu] transfer staging: ...`). Both tree drivers print the staging counters after the base and at the end: bytes and host seconds per path, pairs, waits, and the shared slabs' pinned footprint. Measured on block 25368371 (FAST, RTX 5090), ABBA palindromes at 5cbdf06, on the pre-engine base 169b668 behind a temporary knob. STARK: 121.30 -> 115.50 s (-5.80, A spread 1.20); the base 48.3 -> 44.0 s; under nsys the base's card idle 9.89 -> 5.32 s; host peak -4.41 GiB, which is the slab: 4.01 GiB after the base against 0.13. WHIR: 107.60 -> 105.50 s (-2.10); level 0 -1.3 s. Program identities unchanged. On this base the column-major engine's upload takes the same pairs; that site is new here and not in the measurement above. Tests. Device: exact round trips across chunk boundaries, the closure contract, the buffer-reuse hazard, 12 threads over 8 pairs, and root / host LDE / handle parity of the base, ext3, split-tree and column-major engine commits through either staging, with the staged path's bytes counted. Card-free: the slab footprint and the staging line's shape.
Level 0 opened with host work alone: every first-round wrap's prologue at once (reconstruct, emit, arenas, and on WHIR the epoch harvest), with nothing on the card. None of it needs the card or the global proof, only the ELF and the epoch proofs the base finished long before. By default the tree drivers now start a lead-in before the base. Its helpers (two by default) wait until the base reports its epoch count, then build the prologues of wraps 0..want (want = level 0's first pool round) from copies of the leading epoch proofs, with the same functions the pool calls, so the programs are the pool's own. Level 0 takes each prologue instead of building it. A prologue no helper started is built by the pool as before, and a panicking one is handed back. Nothing in the lead-in takes the card permit or holds device memory of its own. It uses the base's DECODE derivations as the base shares them: the STARK commitment, and on WHIR the root and the prepared opening, whose derivation is a device commit. The base reports through an EpochObserver installed for the calling thread (with_epoch_observer): the epoch count from the producer as soon as the final epoch is executed, each proved epoch, and the DECODE work. That holds on both preparation schedules and both pipelines; with no observer installed, the pipeline is unchanged. Level 0 still takes its own DECODE derivations from the base, and LFM_TREE_REDERIVE_DECODE=1 still re-derives them there; the lead-in's copies are only for its prologues. BaseDecode now holds the Arc the base shares. An I4 schedule test reads epoch positions through the slice-based epoch_chain_position. LFM_TREE_PROLOGUES_AT_LEVEL0=1 builds the prologues at level 0's start instead; LFM_TREE_TAIL_PROLOGUES and LFM_TREE_TAIL_HELPERS size the lead-in. The driver prints `L0 PROLOGUES: ...` either way, and at level 0's start how many prologues were ready. Measured on block 25368371 (FAST, RTX 5090), ABBA palindromes at 5cbdf06, on 169b668 behind a temporary knob, before the base-prep and DECODE-handoff changes. WHIR: 107.60 -> 104.40 s (-3.20, A spread 1.20); the lead-in 3.19 -> 0.07 s; level 0 -3.1 s; base unchanged; device peak +592 MiB, from the harvest's MLE evaluations running unreserved beside the base's tail. STARK: 121.30 -> 118.65 s (-2.65); the lead-in 5.54 -> 0.22 s and level 0 -6.1 s, but the base +3.4 s, because the two helpers slow the base's epoch proofs. Program identities unchanged. On this base the DECODE handoff already removes part of the lead-in, so the gain here is smaller than above. Tests. Card-free: the hand-off (order, the count gate, handing back an unstarted, failed or context-failed prologue, waiting on one in progress, close). Fixture scale, on both bases: the observer sees the count once and every epoch byte for byte, and a prologue built from the leading epochs emits the pool's own program and arenas.
…te does Every staged_transfers test forces its path with the thread override, and only the block-scale tree drivers read the lead-in's setting, so no card-free test read either setting from the environment. One test each now does, and prints the line that names it: - staged_transfers: staging_pairs_enabled() is !LAMBDA_VM_STAGING_SHARED_SLAB, with its `[gpu] transfer staging: ...` line; - the tree tests: lead_in_enabled() is !LFM_TREE_PROLOGUES_AT_LEVEL0, with an `L0 PROLOGUES setting: ...` line. The gate runs each with --nocapture under the default and under the opt-out and counts the named lines.
…nned staging, level 0's first wrap prologues built in the base's tail I6: each row-major commit transfer is staged through a pinned pair of its own instead of the worker's shared slab; opt-out LAMBDA_VM_STAGING_SHARED_SLAB=1, named once on stderr ([gpu] transfer staging: ...). I7: level 0's first wrap prologues are built by helpers in the base's tail; opt-out LFM_TREE_PROLOGUES_AT_LEVEL0=1, sized by LFM_TREE_TAIL_PROLOGUES / LFM_TREE_TAIL_HELPERS (L0 PROLOGUES: ...). No proof byte moves. Measured on the lane's base (IDLE-B box2): I6 STARK -5.80 s, WHIR -2.10 s; I7 WHIR -3.20 s, STARK -2.65 s net.
…xes K3/K4/K5 Brings the HASH lane's six commits onto candidate C2 (7d41668): the half-warp Merkle tops (K3), the work-queue grind (K4), the limb-multiply permutation variants (K5), each still behind its LAMBDA_VM_GAP_* knob, their parity tests and host KATs, the serialised grind-counter tests and the queue grid's context fix. No conflict; the next commits make the three fixes the defaults.
…imb permutation by default The three RPX device fixes measured EFFECTIVE on the WHIR block (job 160, wt300-307: K3 -3.90 s, K4 -6.45 s, K5 variant 5 -10.05 s against 107.15 s, identities identical in every arm) are now the defaults. Each keeps an opt-out, and its default lives in one constant in `rpx_paths`, so a pipeline that needs one off flips one line: - LAMBDA_VM_RPX_WARP_MERKLE (WARP_MERKLE_DEFAULT): narrow levels and the tail on rpx_merkle_level_warp / rpx_merkle_tail_warp; =0 walks a thread per parent (rpx_merkle_level, rpx_merkle_tail). - LAMBDA_VM_RPX_GRIND_QUEUE (GRIND_QUEUE_DEFAULT): rpx_grind_search_queue on a card-filling grid; =0 is rpx_grind_search on LAMBDA_VM_GRIND_GRID. - LAMBDA_VM_RPX_LIMB_PERMUTE (LIMB_PERMUTE_DEFAULT): every RPX kernel from rpx_v5.cubin (32-bit limb multiply, square_n unrolled by four); =0 loads rpx_v0.cubin, the 64-bit multiply. Each variable takes 0 or 1 (anything else aborts) and each switch prints one line on first use, "[gpu] RPX Merkle: ...", "[gpu] RPX grind: ...", "[gpu] RPX permutation: ...", naming the path and whether it came from the pipeline default or the variable. The queue's grid line becomes "[gpu] RPX grind queue: grid ...". build.rs now builds rpx.cu twice (variants 0 and 5) instead of five times. The LAMBDA_VM_GAP_* knobs and the gap_hash module are gone. The parity tests move to prover/tests/rpx_device_paths.rs, named for what they compare (the per-parent walk, the stride grind), and the temporary wording leaves the kernels, the host KATs and the docs.
The PROVE SPLIT line's "r4_grind (n/airs on device)" took its delta of gpu_lde::gpu_grind_calls(), the keccak arm's counter. An RPX grind that runs on the device counts in gpu_grind_calls_rpx(), so under RPX every line read 0/airs while every table ground on the card: on the WHIR block (job 160) the root proof printed 0/11 beside the harness's own count of 11 RPX device grinds for it. device_grinds_now() sums both arms and report() takes its delta of that. rpx_grind_device gains a test that reads it around one RPX device grind (the keccak-only count reads 0 there and fails).
The result lines still carried the campaign's fix ids (K3, K4, K5) and called the old paths "shipped", which stopped meaning anything once the new paths became the defaults. They now say what is compared: the warp walk against the per-parent walk, the queue grind against the stride grind, the limb primitives and the permutation variants. Output only; every check is unchanged and both binaries still pass.
…er half-warp, a queue grind, the limb permutation K3: a Merkle level and the tail compress one permutation per half-warp (opt-out LAMBDA_VM_RPX_WARP_MERKLE=0). K4: the device grind claims nonces from a work queue (LAMBDA_VM_RPX_GRIND_QUEUE=0). K5: the whole RPX module runs the limb-multiply permutation variant (LAMBDA_VM_RPX_LIMB_PERMUTE=0). Each opt-out accepts 0 or 1 and names itself once on first use. Every digest, root and grind nonce search is byte-identical to the previous kernels (host known-answer tests and device parity). The prove split now counts RPX device grinds. Measured on the pre-K1 base: WHIR K3 -3.90 s, K4 -6.45 s, K5 -10.05 s; STARK K3 -1.00 s, K4 -0.50 s, K5 -7.95 s.
The DEEP and out-of-domain denominators were inverted by a global
Montgomery scan: compute_denoms plus five scan kernels, each a full pass
over the domain with prefix and suffix scratch, although every row needs
only its own few inverses. By default now:
- compute_and_invert_denoms_ext3_dev runs one kernel,
invert_denoms_rowwise_ext3_k{1..8}: each thread builds its row's
denominators and inverts them in registers with one base-field
inversion (adjugate over norm, the norms batched by Montgomery's
trick, kernels/ext3_inv.cuh). More than 8 per row keep the scan.
- the fully resident R4 DEEP inverts its own row's 1 + K denominators
(deep_composition_ext3_fused_m{1..4}) and needs no inverse buffer; it
falls back to the buffered kernel above 3 points or on any
precondition miss.
- the single-point OOD sums (the R3 composition parts: 1-2 columns, so
1-2 blocks on the card) run on the row-chunked multi kernel, and the
multi kernels' chunk count loses its 64 cap.
The values are the same field elements; raw limbs may differ by p, which
nothing downstream observes. Measured on block 25368371 on one RTX 5090,
ABBA behind a switch on 169b668: STARK (one_row=auto) -0.70 s whole
run against a 0.50 s A spread and -2.35 GiB device peak; WHIR kernels
-49.8 % with the wall inside the noise.
LAMBDA_VM_DEEP_INV_LEGACY=1 restores all three (the scan, the buffered
DEEP, one block per OOD column), read once per process with a banner,
as LAMBDA_VM_LDE_LEGACY does for the LDE. The legacy paths stay public
for tests/deep_inv_parity.rs, whose tests name both paths per call;
tests/deep_inv_setting.rs checks that the process setting is followed.
gpu_fused_deep_calls() counts the fused dispatch, and
cuda_path_integration asserts it follows the setting: a table that
silently fell back to the buffered kernel would still verify.
recompute_lde_produces_byte_identical_proofs compares the bytes of two proves of one instance, one per residency mode, at the test options' grinding factor of 1. Under `parallel` the CPU nonce search is rayon's find_any (crypto::grinding::generate_nonce), so the two proves can return different valid nonces; the nonce is absorbed before the queries are drawn, and every opening after it moves. The test fails whenever the two searches disagree, whichever residency mode runs: on a laptop it failed 11/20 at d1dc455 and 15/20 at 7d41668, and 20/20 passed with LAMBDA_VM_DETERMINISTIC_GRIND=1 (the smallest nonce) or without `parallel` (a sequential find). The residency tests now prove at grinding factor 0, as zf_golden_tests already does for the same reason. The residency mode acts on the main LDE, which the grind never reads, so the comparison loses nothing it could catch.
…t field A WHIR round has three proof-of-work slots (folding, out-of-domain, query), and the proof carries three nonces a round whatever the bits. A slot whose grind has zero bits, and the last round's out-of-domain slot, is carried and never read, so any value in it verifies. A query-only grind (P2) would leave two such unbound fields a round. ChainFormat gains `nonces: NonceLayout`: - Three (the default): today's format, byte for byte. - Spent: a round carries only the nonces its grinds spend. The in-guest arena has no word for an unspent nonce; ChainShape::carries is the one place the layout is written, and the word count, the hints and the arena words all read it. The host verifier refuses a nonzero value in a host field the layout does not carry (Error::UnspentNonce). RoundNonces keeps its three fields, so both layouts share one proof type and Three keeps its bytes. GrindBits::query_only(bits) grinds before the query positions only. The query count reads the query grind alone, so it does not move. No production config uses Spent or a query-only grind yet, so every proof, program and pin is unchanged. The legacy layout's chain programs, arenas and proof bytes are pinned against values printed at 0428c39, and the transcript closed form now prices only the grinds a config spends.
Each WHIR base-chain round ground 20 bits before three challenges. Only the query grind buys proven bits as placed: the folding grind sits before the round's first sumcheck message, so the first folding challenge is redrawn by varying that message at one hash a try, and the out-of-domain grind follows the out-of-domain point. The new ZF lever `whir_grind` therefore defaults to `query`: GrindBits::query_only(20) under NonceLayout::Spent, one grind and one nonce word a round, 518 grinds a block instead of 1,472 at stack 27. The query count reads the query grind alone and stays 112. The proven bits per phase do not move: chain minimum 130.393 at stack 27, pipeline minimum 128.946. LAMBDA_VM_ZF_WHIR_GRIND=all is the opt-out: GrindBits::uniform(20) under NonceLayout::Three, the production config from before this commit. A test pins it against a literal, and its chain programs against the values printed at 0428c39. The banner gains `whir_grind=`. No univariate option reads the lever, so no STARK proof, program or id moves. Re-blessed: the production default chain's pins now describe the P2 chain (12 grind permutations, 16,411 permutations, 32,590 arena words, 150,258 / 202,873 rows). Its previous pins move unchanged to the opt-out's test. The banner strings in zf_format's tests gain the new key.
The ZfFormat::DEFAULT doc quoted a block timing for whir_grind=query from an earlier measurement. A measured number in a comment goes stale, so the doc now says what the lever does and why it loses no proven bits. The same reasoning in whir_chain's module header, GrindBits::query_only and WhirGrind gave the out-of-domain grind's reason as "it follows the out-of-domain point". That is half of it: the grind sits right before the batching challenge and does guard it. Dropping it costs nothing because the batching challenge has far more bits than the target without any grind. Comments only.
evaluate_many_base uploaded the rest of the point for its fold loop even when nothing was left to fold, which asks the driver for a zero-byte allocation at one variable. The loop and its uploads now run only from two variables up; at two and more the same launches run in the same order. The parity test gains the short shapes the argue's device-columns knob sends to this path: 2^2 to 2^15 rows, up to 2,000 columns, and two chunks of the evaluation budget at 2^15.
…DA_VM_ARGUE_DEVICE_COLUMNS) The claim reduce evaluates every column of a table at its reduced point. The batched device evaluation took a table only once each column was 2^16 rows tall, a threshold set for columns that had to be uploaded first. The epoch's columns are already resident, so for a resident table the host cost is its cells: the widest precompile, 1,480 columns of 2^15 rows, was walked one column after another on one thread, 263-303 ms a table, about 1.1 s of the WHIR base's argue idle on block 25368371 (D-GFS, G2). Under LAMBDA_VM_ARGUE_DEVICE_COLUMNS=1 a resident table goes to the card once it holds 2^16 cells, and the columns left to the host are spread over the pool when each is a host loop. Off by default; off is today's path, line for line. The values are the columns' multilinear extensions at the point, so the proof does not change. - host_evaluate_calls() counts the columns walked on the host, beside evaluate_calls(); under LAMBDA_VM_BASE_SPLIT=1 every prove prints `ARGUE COLUMNS #k: on the card C · on the host H · xchecked X || device columns on|off · xcheck on|off`, the A/B's mechanism line. - LAMBDA_VM_ARGUE_XCHECK=1 recomputes every card value on the host and fails the prove on a mismatch, naming the column. It is the block's identity gate: two proves of a block never share bytes (six table builders order rows by HashMap iteration), so the comparison has to be in-process. - force_column_value_fault arms a one-cell corruption so the checks can be shown to fail. Never armed outside a test. Tests: the host walk proves the same with the knob on (unit); on a device, card against host at 2^1..2^15 rows and up to 2,000 columns, resident and uploaded; the reduce's proof, point and transcript knob off/on with the path each arm took; the threshold; and the fault seen by the identity, the verifier and the cross-check (tests/argue_device_columns.rs, box only).
…he card The CPU, ADD and MUL fixture at a height where LAMBDA_VM_ARGUE_DEVICE_COLUMNS has work to move: CPU 2^14 x 5 and ADD 2^14 x 4 (2^16 cells, the threshold) go to the card once resident, MUL 2^13 x 4 stays on the host. multi_prove with the knob off and on must give the same bincode bytes, table by table and whole, the same transcript state and next challenge, and both proofs verify. On a device each arm is shown to have taken its own path (9 columns on the card with the knob on, none off); without one, the host walk is compared with the same walk spread over the pool. The negative control arms the one-cell fault: the identity must then name CPU, the first table the card values, and the proof must fail the claim reduce's check. It needs a device and says SKIPPED without one.
…r's copies A table statement took its preprocessed count from the column copies the verifier held. An AIR that declares its preprocessed columns by count only (`with_preprocessed`, as every LFM chip does) has no column builder, so its statement counted zero: with no prepared opening its program columns were bound by nothing, and with one the honest proof was refused (`settled > preprocessed.len()`). `TableStatement` now carries `num_preprocessed`. `statement_with_preprocessed` sets it to the number of copies, so every existing caller is unchanged and no proof byte moves. The new `statement_with_prepared_prefix(count)` holds no copies and takes the AIR's count; `check_preprocessed` refuses a prefix that is neither settled by a prepared opening nor recomputed, and `multi_verify` bounds the settled count by `num_preprocessed`. Tests on the three-table fixture, with the CPU table's first three columns declared preprocessed and its reversed trace as a forged program that still balances the bus: the forgery verifies under a count-zero statement (the control), and is refused with the count and no opening, with an opening over the pinned columns, and with an opening that settles only part of the prefix; an honest proof with its opening verifies.
Pure WHIR, D-WHIR §2: a recursion proof (wrap, node, root) can be proved by the base's multilinear prover instead of one STARK per table. The LFM chips are AirWithBuses with a single-source constraint IR and bus interactions, so this is `multilinear_table::multi_prove` over the program's tables plus glue: - the prepared stack: every table's preprocessed prefix (the instruction groups) committed once per program as one stacked WHIR commitment and opened at each table's reduced point — DECODE's mechanism over several tables; - the statement (tag, program_id_w, version, word count, each public word as four felts, the heights, the chain config, the pad), absorbed before any challenge; indices are implicit and a claim out of position is refused; - program_id_w, folding the prepared roots, the heights and counts, the hasher, the chip set and the WHIR format. Every statement is built with `statement_with_prepared_prefix` at the AIR's own preprocessed count, and a prepared plan that does not settle exactly `0..count` of every table is refused before any argument runs (the §2.4 trap). The hash is RPX by type (`RpxWhir`, `DefaultTranscript<E, RpxTranscriptHash>`), never the `LAMBDA_VM_WHIR_HASH` knob. Policy A: one group, the prefix committed in the main stack and in the prepared stack. The chip set must be the WHIR recursion one (no keccak, BLAKE3 or BITWISE); an LFM_HASH split is one more table. `LAMBDA_VM_LFM_PROVER=stark|whir` selects the prover (default stark, bannered on every setting, an unknown value aborts); nothing reads it yet, so every proof is today's. Tests: TrivialV0 round-trips at the production config (grinding on) and the registry programs round-trip or are refused by chip set; a forged program (one constant changed, same shape) verifies under count-zero statements — the trap — and is refused by the W-LFM verifier without an opening and with an opening over the honest stack; a deleted opening, a plan missing a table, a restated height, a tampered or reordered public word are refused; a KAT against an independent transcript shows the prepared roots are absorbed before z. Tests other than the anchor run ungrinded through a cfg(test) switch in the one config derivation.
D-WHIR §3: the wrap's `whir_epoch_program` body as a node leg, with the W-LFM statement over the child's hinted public words, the child's prepared-stack roots interned in the roots block, every table's plan `settled = count, route None`, the closure against the claimed words' LfmPublic balance, the main group walk, and the prepared opening last with each table's prefix at its own reduced point. The roots block, table walk, group walk and stacked opening are the wrap's own emitters. The public words are hinted as four felts each: the statement's Pack rows read every lane as a base token, so a hinted word with nonzero upper lanes has no satisfying assignment. The plan is `WhirLfmPlan::build`, which refuses at emit time a prepared plan that does not settle every AIR prefix exactly. `emit_whir_node` is `emit_node` with W-legs (bindings and publishes read only the legs' lanes); `emit_node` and every STARK emitter are untouched. The cost form `whir_leg_cost` composes the landed forms, including a shape-only count of each table's proof words, and `shape_only_artifacts` evaluates it at a program shape with no commitment. Tests: the leg executes an honest TrivialV0 child at the production config and publishes the host's (z, alpha); a mutated published felt, carried root, bus output, GKR word, column value, main-chain and prepared-chain final value are each refused; the forged program is refused; a plan missing a table is refused at emit time; F1 is exact by kind on a real child (ops 214,566, constants 150 both ways, hints 42,349, permutations 17,203) and each table's shape word count equals its arena length; pins at the D-WHIR §3.2 shapes (wrap 0: 39,795 permutations, 398,742 ops, 75,432 hints), within three units of the design's instrument. The leg program proved as a W-LFM proof is box-scale (16.9 M cells, stack n25) and is ignored here.
shift_evals(x, k)[y] = eq_evals(x)[(y - k) mod 2^n]: for a corner y, shift_k(x, y) is the indicator of x = y - k extended multilinearly in x, which is eq(x, y - k). The claim reduce's tables on the card (LAMBDA_VM_ARGUE_DEVICE_TABLES) build eq(alpha) once and read each offset's shift table as a rotated copy, so the carry recursion must agree with the rotation value for value: every offset of the small cubes, offsets past the cube, base field and extension.
…ix out Policy B of a W-LFM proof (D-WHIR §2.4) commits each table's preprocessed prefix only in the prepared stack: the main stack holds the value columns, and the prefix's claims are settled by the prepared opening alone — the same binding, since that opening already settles them under policy A. - stacked_eval: `ColumnsAt` names a group's columns in a resident store, `From(first)` (today's contiguous run) or `Map` (one store column each); `commit_mapped` / `prove_mapped` take it, and `commit` / `prove` delegate with `From`, so every existing caller is unchanged. - multilinear_table: `CommittedTables::commit_grouped_settled` stacks each table's columns past its settled prefix, reading them out of the resident store (which keeps every table contiguous for its argument) through a column map; `commit_grouped` is it with no prefix. `multi_prove` refuses, before any transcript work, a proof whose prepared opening does not settle exactly the prefixes left out, and opens each group on its own columns' claims. `multi_verify_settled` takes the exclusion as the verifier's format choice; `multi_verify` is it with none. Tests: policy B round-trips on the three-table fixture; the forged prefix is refused under B; a B proof read as A and an A proof read as B are refused; a prover whose opening does not match the exclusion is refused.
`PrepPolicy` (A: `Both`, B: `PreparedOnly`) is a W-LFM format choice stamped on the artifacts and folded into program_id_w; `LAMBDA_VM_LFM_WHIR_PREP= both|prepared` selects it (default both, bannered, an unknown value aborts), and `build_whir_artifacts_under` takes it explicitly so one process builds a program both ways. The prepared stack does not depend on it. Under B the plan's main shapes are the value columns, the prover commits with `commit_grouped_settled`, the host verifier runs `multi_verify_settled`, and the W-leg's group walk opens each table's columns past its prefix (the prepared opening still takes the full walk). The cost form follows the main shapes. Tests: TrivialV0 round-trips under B, a B proof read as A and an A proof read as B are refused, the forged program is refused under B on the host and in the leg, F1 is exact under B (ops 200,986, constants 149, hints 39,995, permutations 15,894), and the §3.2 pins under B (wrap 0: 38,440 permutations, 384,825 ops, 73,961 hints; the design's instrument had 38,441 / 384,823 / 73,961).
Under LAMBDA_VM_BASE_SPLIT=1 the W-LFM prover pushes the base's split record (prep, absorb, commit, prove, wall, and the inner challenge / argue / open slots multi_prove accumulates) under a new index, `LFM_INDEX`, which the line names `W-LFM` — so a recursion prove prints `WHIR PROVE SPLIT W-LFM` and no reader of the base's table can take it for an epoch. Stages are timed directly rather than through `stage_done`, whose `BASE EPOCH` line would be misread. Nothing changes with the variable unset.
`w1_one_wrap_and_one_node_proved_both_ways`, ignored (box tier, cuda, the production block): proves the RPX base, then - W1a: wraps 0..3 under STARK, wrap 0 under W-LFM policy A and B, wraps 1..3 under policy B; - W1b: L1N0 (arity 3 over the STARK wraps, as the tree emits it) under STARK and W-LFM A and B; - W1c: the pure-WHIR L1 node emitted over the policy-B WHIR wraps with W-legs, proved under W-LFM A and B, and the W-leg cost form per child; - W1d: the census of an L2-shaped parent over three STARK L1N0 children and over three W1c children. Each prove prints one `W1 PROVE` line: build, prove and its split, host verify, proof bytes, host peak, the stacks (W-LFM) and the identity. The STARK wraps' and L1N0's identities are the default's byte gate against the record's tree at the base sha (wt800).
The W1 test held every wrap's full W-LFM build, prepared-stack codewords included (device-resident), while only the artifacts and the proof feed the pure-WHIR node and the parent census. `w1_prove_whir` now returns the artifacts and drops the prover's stack after the host verify.
The security gate reads each chain's shape (num_vars, blowup, fold schedule, caps, queries, grinds) from GAPB CHAIN lines, and the census instrument that prints them (a9ffbf48a) is not on this branch. W1 prints the same line, from the artifacts' config, for every main and prepared polynomial a W-LFM proof opens, tagged with the proof it belongs to.
Pieces for LAMBDA_VM_ARGUE_DEVICE_TABLES; nothing calls them until the multilinear side does, and every existing path is unchanged. - Extra and DeviceFactors::session_with: a weight a resident sumcheck adds is a table to upload, as before (session delegates, byte for byte), or the point of an eq table, built into its slab with eq_expand_into. - reduce_session: the claim reduce's session with its tables built from the epoch's resident columns. eq(alpha) is built once; each offset's shift table is two device-to-device copies of it, since shift_k(alpha, y) is eq(alpha, y - k) at a corner (pinned by multilinear's a_shift_table_is_the_eq_table_rotated). Each offset's batched column is the new batched_column_ext3 kernel over the resident run: one thread per row, base x ext3 per member, the host's sum in another order. - SumcheckSession::factor and set_cell: read one factor, write one cell, for the cross-check and the fault hooks.
…AMBDA_VM_ARGUE_DEVICE_TABLES) The WHIR base's ARGUE builds two families of challenge-dependent tables on the host, on the rayon pool, then uploads them pageable: the zerocheck's weights eq(r) and eq(row), and the claim reduce's shift tables and batched columns. On the head's trace the ARGUE thread waits 3.56 s in those joins (the producer's prep holds the pool) and the uploads are 20 GB a block, 1.16 s copy-only (I-GFS.md §6). Under LAMBDA_VM_ARGUE_DEVICE_TABLES=1 they are built where they are folded, from a few kilobytes: the points, the columns each offset reads, and their weights. Off by default; off is today's path. - A2: batch::Weight is a table or the point of an eq table. multi_prove passes the zerocheck's two weights as points under the knob; prove_resident keeps a point until the card builds it with its rounds, or builds it on the host if the card turns them down. prove_core takes Weights; prove_statements wraps its tables. - A3: claim_reduce::prove asks gpu::prove_reduce_resident first. Its gates are prove_sumcheck's over the same tables, plus a resident run; its rounds stop at the same crossover and the host finishes over the folded tables through the same loop. A decline is before the transcript moves. The values are the host's in exact arithmetic, so the proof is too. Under LAMBDA_VM_ARGUE_XCHECK every table the card built is compared with the host's before the first round and a mismatch fails the prove, naming it; force_table_fault arms a one-cell corruption so that check, the identity and the verifier can be shown to see one. Under LAMBDA_VM_BASE_SPLIT each prove prints `ARGUE TABLES #k: built on the card N · xchecked X || device tables on|off · xcheck on|off`. Tests: the host arm (a point proves what its table proves); on a device (tests/argue_device_tables.rs, box only), the reduce tables and weights against the host at many shapes and offsets, and a prove_core knob matrix with shifted reads, closures and nothing resident, verified, with the path each arm took and the negative controls; in stark, the tall fixture's whole argument identical with tables on, alone and with the columns knob, and its negative control.
SumcheckSession::values makes one synchronous device-to-host copy per factor. values_gathered reads the same bytes with one launch of the new gather_factor_heads_ext3 kernel (each factor's first len cells, through the session's pointer table, into one buffer) and one copy back. Nothing calls it until multilinear does, under LAMBDA_VM_ARGUE_LEAN_READS. The parity test compares the two reads bit for bit after rounds and folds, for uploaded sessions and sessions over resident factors with an eq weight and a table added, from 1 factor to 1,482.
Production lays the streamed chunks out on 3 workers since 59e9890; two test harness comments still described 0 (the inline layout) as production. Comments only.
…kes them narrow (cut b2) After b1 the 4.13x high-water is the finish itself (BIG 561: at p5's end the finish holds its tables, 24.6 GiB at eight bytes a cell, beside held_narrow and its op lists). b2 keeps those tables packed from the moment each is generated. - multilinear: NarrowColumns::pack_row_major packs a row-major trace column by column (#1013's NarrowMain::pack, the same layout), and TraceData::new_narrow holds packed columns from the start. - stark: TraceTable::pack_main_narrow / narrow_main / take_narrow_main (the packed main trace, no 64-bit copy kept); CommittedTable::from_narrow (a layout whose committed columns are the main columns in order). - trace builder: StreamSkip::pack and WindowedTraceBuilder::pack_finished_tables: each table the finish generates is packed as it is made, chunk by chunk. KECCAK, ECSM, ECDAS and an unchunked KECCAK_RND stay wide (the block cuts them after the build). - block whir: BlockOptions::pack_finished (production on, BLOCK_WHIR_PACK_FINISHED=0|1): a packed table is laid out narrow (table_of_narrow: the packed columns become the table's, its preprocessed columns checked word for word against the program's), with no transposition and no wide column made. The BLOCK REST PACKED line counts them. - phase A: a group holding a narrow table is uploaded packed and widened on the card (upload_group), the commit reads handles (a host fallback widens), and M4's packer skips what is narrow already. The memory log counts held bytes. The tables, groups and proof bytes are the same: a test proves test_keccak_multi with b1 and b2 on and off, inline and on three layout workers, and compares the partition, the statement and (deterministic grind) the bytes; a packed table whose first preprocessed column differs in one word is refused (mutation: with the comparison off, that test fails).
With the finish packing its tables (b2), each KECCAK_RND chunk's pack is serial on the task that built it, and one task building and packing every chunk in turn would make KECCAK_RND the finish's last table by more than it already is (BIG 561: gen_keccak_rnds ends 4.6 s after LT at 4.13x). Packed chunks are now built and packed in parallel waves of four, at most four of them (about 3 GiB) wide at once; the tables and their order are the same, and the unpacked path is unchanged.
…es are laid out #1014 now lays the streamed chunks out on three workers (0498480), and phase A waits on the groups the rest closes, which close only once the whole rest is laid out. With b1's waves the AIR order's prefix is laid out first, so pack_rest_as_laid_out (the rest's sink, groups identical by test) is now on in production: each wave's tables go to the packer as the wave ends, and a group closes once its own tables are laid out. Laid out all at once, the first table in AIR order was ready only with the last, which is why packing as laid out read +0.21 s before (FAST 422). The real-block tests' BLOCK_WHIR_PACK_REST=0|1 now overrides the production choice instead of replacing it with off. Readouts for the finish's place in phase A: each group's commit start (from@, beside done@), and every rest table's layout end and group in AIR order (BLOCK REST TABLES).
…r and re-hash For the timeline of #1014's phase B: with LAMBDA_VM_BASE_SPLIT=1 the block records, over phase B, the WHIR chains' six host stages (grind, sumcheck, fold, commit_folded, ood, queries), the round wall, the query split (sample, tree rebuild, coset gather, assemble), and the seconds the first-round paths from the kept tree tops spend gathering their blocks and re-hashing them on the host (two new counters in whir_commit). BLOCK OPEN SPLIT prints them. Readout only: no schedule or proof byte changes.
…arallel A revived commitment opens its first-round paths from the kept tree top: the leaves under each queried block are gathered and re-hashed on the host, and each block's subtree root is checked against the kept node. That loop ran one block at a time on the prover's thread while the card waited: on #1014's 1x block it is 27 gaps of about 37 ms in phase B's openings, 1.01 s in all (nsys timeline, FAST 839), with nothing else on the host. Each block's subtree depends only on its own leaves, so the blocks are now re-hashed in parallel. The subtrees stay in block order, the paths are assembled from them as before, and the first block in that order whose root is not the kept node is the one refused, as the serial walk would. The proof bytes do not change. A new test opens a tall stack with 48 queries (tens of distinct blocks a round) after retire and revive at drops 3 and 4: the opening equals the kept one's byte for byte, and a revive over other columns still refuses.
…) into m4b/b2
… at the finish) FAST 852 + 853 at three layout workers, 8 + 8 arms against 59e9890's behaviour: b1 with the sink −0.124 s, b1's waves without it −0.135 s, the sink alone +0.21 s. The groups do close earlier with the sink, but at 1x the committer is still busy when they do. pack_rest_as_laid_out goes back to off; the readouts and the override stay.
BLOCK_WHIR_REHASH_SERIAL=1 makes the kept-top paths re-hash their queried blocks one at a time, as before the parallel re-hash, in a parallel build. FAST 840 measured the parallel re-hash on two binaries: phase B -0.93 s, but phase A +0.71 s, where the change never runs. The knob lets one binary carry both arms, so a build effect and a real one can be told apart. The bytes are the same either way (the stacked_eval byte tests pass with it set); read once per process.
…d the top's digest The W3 harness prints several readouts on the clock between the base and the tree (the block report, a second block frame for the LT heights, the group tables): 0.13 s at 1x that a prover would not spend (FAST 839). The whole block keeps its meaning, since every earlier W3 number includes them; a second field beside it now gives the whole block without them. It also prints a blake3 digest of the top proof's bytes, so two runs under LAMBDA_VM_FIXED_TRACE_HASH=1 and LAMBDA_VM_DETERMINISTIC_GRIND=1 can be compared byte for byte (off the clock).
…built The W3 harness proves #1014's tree the way a prover would run it. Its level 0 built all three leaves' artifacts first (serial card holds, 0.24 s at 1x) and only then let the leaves execute and fill their traces (0.46 s until the first was ready), with the card idle for the second part (nsys timeline, FAST 839). Execution and fill do not read the artifacts; only the prove does. lfm_prove_with_residency is cut into its two halves, unchanged in order: lfm_execute_and_fill (the host phase) and LfmFilled::prove (the card phase, asserting the artifacts' hasher is the one the traces were filled for). The harness now builds the leaves' artifacts on a thread of their own and each leaf executes and fills beside them, waiting for its artifacts only to prove. W3_EXEC_BESIDE_ARTIFACTS=0 is the control (every leaf waits for every artifact first, the order before). The proofs are the same.
…rver block_prove_on_forks_observed calls an observer after each group's opening with that group's share of the proof (GroupOpened: the roots, its batched argue or its tables' proofs, its opening, its prepared opening). block_prove_on_forks delegates with a no-op, and the streamed block prove takes the observer beside its statement observer (prove_block_whir_observed_groups). The observer only reads: the proof is the same with any observer. It lets the block tree's first leaf, whose groups are all opened before phase B ends, start before the base is done.
block_leaf_arena is split into group_arena_words (one group's words from its own share of the proof: its tables' proofs or batched argue, its opening, its prepared opening) and leaf_arena (the block's roots, then the leaf's groups in order). The arena is the same; the shape checks and their refusals stay, with a group's prepared opening now required to match the plan's one prepared stack for that group.
The W3 harness now listens to phase B's groups as they open. The first leaf whose groups are all opened while another group is still to come builds its arena from them and executes and fills its traces on a thread of its own; the tree's worker for that leaf joins it instead of executing after the base. After R2-i the card still idled about 0.23 s at 1x before the first leaf's prove, waiting for that leaf's execution and fill (FAST 841 + 842). W3_LEAF_DURING_PHASE_B=0 is the control. Off the clock, the early leaf's arena is checked against the finished proof's (W3 EARLY LEAF), and a box-tier test proves small streamed blocks under both argue formats and checks that the groups phase B hands over build every leaf's arena.
…s its tables, the block memory log) into nw2/r2ii No conflicts. multilinear_block.rs: the memory log's phase-B marks and the group observer sit side by side at the end of each group (the observer first, then the group's columns are released and the mark taken). block_whir.rs and the W3 harness: the new BlockOptions fields and the memory log's plumbing next to the group observer's.
… proof to an observer; the block tree's first leaf executes during phase B) into m4b/b2 No conflict. R2-ii's two full BlockOptions literals in lfm/whir_block_tests.rs already carry b2's pack_finished: true, so its hand-over tests run with the finish's tables packed, as production does.
… the ledger and the pool count them Off the clock, after the tree: the commit/encode and opening host fallbacks, device commit errors, argue device fallbacks, GKR tree refusals, fused declines, the reservation high-water and the memory pool's live high-water. A run that changes the pool's release posture needs the card's peak from the ledger and the pool, not from total - free, which reads a retaining pool's freed blocks as used.
…its groups close The verifier refuses a statement with more groups than BlockFormat::max_groups (block_frame), but the prover never checked: on the median block (90 groups against 64) it spent its whole base on a proof the verifier refuses (BIG 564). One check, check_group_count, now serves the verifier and both prover paths: the streamed packer refuses the group past the maximum as it closes, and the whole-run path refuses after sizing its groups. The error is the verifier's, word for word. Prover-only: a block within the maximum proves as before.
MauroToscano
added a commit
that referenced
this pull request
Oct 3, 2026
…14fea1 crypto/stark/src/narrow.rs and crypto/stark/src/spill.rs are #1013's files, byte-identical to mem/stark-f1 @ e5214fea1 (i-mem3's F1: the page-backed Bytes, the streaming Digester, the fused writer, over the S1/S2 store). #1013 is their one source of truth: #1014 never edits them, re-imports any change from there, and they merge as one when both PRs reach main. Wiring only: the two modules in lib.rs (dead code allowed there, since some parts serve #1013's prover alone), math's page-bytes feature, and libc unconditional as the store needs it (it was behind disk-spill). #1013's prover glue (TraceTable::spill_main) and its spill tests are not imported; #1014's adapter and tests follow.
MauroToscano
added a commit
that referenced
this pull request
Oct 3, 2026
… reserve LAMBDA_VM_BLOCK_SPILL now reads as #1013's (prover/src/block.rs SpillPolicy … spill_decision @ edddc68): auto (and unset) | off | always | <GiB> resident budget. auto spills a committed table once the host's bytes (the larger of VmHWM and the cgroup's memory.current), the reserve and the table would pass the target (LAMBDA_VM_BLOCK_SPILL_TARGET_GIB, else the cgroup's memory.max less 10 GiB, else MemTotal less 10 GiB). Phase A decides per table and never past the writers' queue; groups 0 and 1 count but never spill. The reserve is #1014's own, inferred from the median block (BIG 565): the finish's p5 transient (25.5 GiB) and the tree's leaf programs alive through phase B (16.3 GiB), about 1.25 GiB per total G cells or 1.65 per G cells committed so far, plus 6 GiB for phase B's bump and the read-back window. BlockSpill carries the policy's choice (SpillWanted) and the kept bytes and committed cells it reads. A block that fits spills nothing; the store opens for every policy but off. Tests: the knob and the decision (unit), a budget the block fits in spills nothing and one of zero spills (block).
MauroToscano
added a commit
that referenced
this pull request
Oct 3, 2026
…@ 278e6a8 The rule #1014 copies from #1013 read only cgroup v2's memory.max and memory.current. On a v1 host (FAST) the target fell through to MemTotal less 10 GiB, and the host's charge was not read at all. #1013 fixed both in 278e6a8; this re-copies spill_target_bytes, spill_target_from, cgroup_memory and host_bytes_now from that sha, with its two tests (the cgroup files from fake trees, and the target rule), cited as before. The only changes from #1013's text are the citations and the tests' temporary directory, which gets a name of its own so the two copies of the test never share one if they ever run in one process.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One WHIR (multilinear) proof per block, with no epochs (the prove-and-retire / VADCOP shape). This is a second prover next to #1010's epoch-based one. #1010 stays the reference and this branch does not touch it. The branch starts from #1010's head
f3d359998; compare againstf3d359998to see only this work.What it is
prover::block_whir::prove_block_whir/verify_block_whir. The whole block is oneMultiProof:(z, α, β)are drawn once.S_post ‖ g: upload again, argue (LogUp-GKR), re-encode with the NTT alone, open. The dropped tree levels are rebuilt from the queried cosets and checked against the kept nodes.WindowedTraceBuilder(noepoch/windowed-builder, also used by STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013). Every full chunk of CPU, MEMW_R, MEMW_A, MEMW, LOAD, LT, SHIFT and STORE is laid out and committed as soon as it exists.⚠ Format changes (all approved by Mauro, 10-02, but §6, the port of #1013's S0a; for the cryptography team's end review; ledger rows W-0…W-6, W-8 in CRYPTO-REVIEW.md)
The block is one WHIR proof (
BlockWhirProof, oneMultiProof) over all of the block's tables.z, α, β.S_post ‖ g).Every parameter below is a verifier-side constant. None is read from the proof.
BLOCK_GROUP_POLYSBLOCK_MAX_GROUPSBLOCK_MAX_TABLE_VARSBLOCK_MAX_KECCAK_RNDBLOCK_KECCAK_RND_MAX_VARSBLOCK_MAX_ECDASBLOCK_ECDAS_MAX_VARSBLOCK_MAX_KECCAKBLOCK_KECCAK_MAX_VARSBLOCK_MAX_ECSMBLOCK_ECSM_MAX_VARSArgueFormat::BATCHEDLEAF_PERMS_CAP,BLOCK_FAN_IN1. The group partition is in the statement
What. The streamed prover packs tables into groups in the order their chunks complete. It writes the partition (per group, its tables' indices) into the statement. The transcript absorbs it before any root.
Verifier. The verifier checks:
It then rebuilds every group's stack layout itself.
Security. Proven bits are unchanged.
BLOCK_MAX_TABLE_VARS(change 3), so no partition can make a chain taller than 2^27.Proof bytes vary from run to run. Arrival order depends on thread timing, and the hash-ordered tables' row order (LT, EQ, BYTEWISE, BRANCH, MUL, DVRM) follows a hash state that is random per process. Mauro, 10-02: "the proof not having the same bytes it's fine, they never had the same bytes anyways".
LAMBDA_VM_FIXED_TRACE_HASH=1(fixed hash keys) andLAMBDA_VM_DETERMINISTIC_GRIND=1. Under both, two processes prove the same bytes (FAST 423).2. Prepared openings, one stack per group (approved by Mauro, 10-02: "it's a standard technique, go ahead")
In short. This is #1010's existing prepared DECODE opening, carried into the block format. It is extended to the dense genesis pages and stacked per group.
What. The recursion guest cannot evaluate DECODE's five 2^20 preprocessed columns, nor the dense genesis pages (millions of rows), on its own. So each group's prepared tables (DECODE and the dense genesis pages chosen by
genesis_stack_plan) have their leading preprocessed columns stacked into one commitment.verify_block_whir, andverify_block_tree's plan. The prover's own roots only shortcut the prover's side.z.Negatives, on a guest with two dense genesis pages. Each is refused by the host verifier and by the recursion leaf, and admitted when the openings are skipped (the mutation):
Wrongly committed stacks are refused at the roots block: two pages swapped, a page holding another page's columns, another program's DECODE.
Cost. One stack per group costs +0.14 s of base (FAST 419; one commitment per table cost +0.49 s) and saves 0.57 s of recursion, by taking 90 k permutations out of the leaves.
BlockFormat::prepared = falseturns the openings off, for measurement only.3. ECDAS split; table heights capped
What.
split_ecdas).(round, op).Why. At a median block, one ECDAS table (2^19) argues on a 12 GiB tree. At p90 (2^20) it stacks into five polynomials on a 24 GiB tree.
Gates. FAST 534 and 535: the negatives, 3 mutations caught.
4. Recursion leaf: the leaf cap (G3) and the inverse share (W1)
LEAF_PERMS_CAP. With no fixed leaf count, it adds leaves while the heaviest leaf is over the cap. The block's own partition is unchanged.p · ediv(1, q), which has no satisfying assignment for q = 0 whatever p is. The formerediv(p, q)left the share free at p = q = 0 (≈ 2^-160 under GKR soundness).LFM_WHIR_SHARE_INVERSE=0keeps the former form.5. Batched argue, on by default since d9d0ac5 (
BlockFormat::argue = ArgueFormat::BATCHED; Mauro, 10-02: "batch the constraints")What. Each group's tables are argued together on the group's fork:
Verifier side. The verifier derives the bins from the statement's shapes and its own cap; the variant is never read from the proof. The proof carries one
BatchedArgueper group inBlockWhirProof.argues, andproof.tablesis empty under this format. The recursion leaves verify the same format (lane i-batch2, N-4).Security at the block's measured inputs (D-BATCH §3.2's formulas; re-checked at FAST 424's inputs: |T| ≤ 60 tables a group, ≤ 51 trees a bin, ≤ 2^28.87 input cells, bus messages ≤ 204 elements, N_C ≤ 413, D_max 4, n ≤ 21; L = 300):
Every term is above the WHIR phase, so the proof minimum is unchanged: 128.946 under the campaign accounting (130.393 for the WHIR fold under the calculator of record).
Measured.
The knob.
ArgueFormat::PerTable(one argument per table, the format before) stays measurable, withBLOCK_WHIR_ARGUE=per-tablein the real-block tests. At the per-table format, the MultiProof, the prepared openings and the partition are digest-equal to the head before the batched merge (BIG 393).6. KECCAK and ECSM split; their heights capped (the port of #1013's S0a; landed at 3cfffe2; approved by the lead under Mauro's any-block target, Mauro to confirm)
What.
split_keccak,split_ecsm), as ECDAS is at 2^17.Why. A keccak-heavy block at the gas limit makes about 2^21 permutation calls. One KECCAK table that tall stacks into eight polynomials of 2^27, against a group budget of three. At its cap each table fits one polynomial.
Bytes. Every block measured so far has KECCAK ≤ 2^17 rows and ECSM ≤ 2^11, so it builds the same tables and the same proof (FAST 830: digest 07d1bd43… unchanged on block 25368371).
Gates. FAST 830: KECCAK and ECSM forced into 4 tables each on block 25368371 prove and verify, base and tree; the count and height negatives; mutation C caught.
One-binary A/B: the block tree vs #1010's epoch tree (FAST 416, noepoch/whir @ 454dad6)
Eight arms E B B E E B B E on the same binary (md5 checked after every arm):
Base, block 25368371 on FAST (each run against #1010's epoch base on the same binary)
Recursion (W3): proved and verified on block 25368371 (FAST 413)
lfm::whir_block::WhirBlockPlan) runs the host verifier's own statement checks, derives every shape from the AIR and every prepared root from the program, and never reads a proof. It prices each group in permutations and partitions the groups over the leaves (heaviest first, onto the least-loaded leaf).block_node, shared byte for byte with STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013). They check the id, the state and the output equal across children, add the shares, and the top asserts zero.verify_block_tree(elf, statement, top)derives the top program and checks the top proof against it. Every preset is pinned inside: the base's options, the block format, the tree's options, the leaf rule and the fan-in.Every proof was verified by the harness off the clock. The base is unchanged by the emission running beside phase B (20.98 s vs 21.18 s in the arms without the change).
Next:
Base levers after the A/B (FAST 417–422, each S/P on one binary)
Memory (D-MEMORY, with i-mem / i-mem2)
drop_streamed_ops)layout_ahead, D-EXEC)Narrow trace storage (M4, lane i-m4)
Prover-only; the proof bytes are the same. The proof digest equals the head's before M4 (07d1bd43…) under the fixed trace hash + deterministic grind (FAST 820).
BlockOptions::narrow, productionNarrowing::CARD).Max RSS, wide → narrow (BIG 560; the traces themselves shrink about 4×, e.g. 24.89 → 6.20 GiB at 1×):
Before M4, #1014 ran out of memory at 3.03× (i-m4); the narrow line puts the 120.7 GiB edge near 5.5–6×.
Phase A uploads the next group beside the commit (3′, lane i-noepoch-w2)
Prover-only; the proof bytes are the same (digest 07d1bd43… at the default, FAST 832; the commits, their order and their bytes are unchanged).
BlockOptions::upload_ahead, on by default;BLOCK_WHIR_UPLOAD_AHEAD=0is the control).Streamed chunks laid out on three threads (E3 on by default, lane i-noepoch-w2)
Prover-only; the proof bytes are the same (digest 07d1bd43… at the default, FAST 838 and 835; the groups, their order and the packing are unchanged).
layout_workers0 → 3 inBlockOptions::production(59e9890). The streamed chunks are laid out on three threads, with at most K + 1 = 3 of them unpacked at a time (layout_ahead, E3), instead of on the builder's one layout thread.BLOCK_WHIR_LAYOUT_WORKERS=0(the inline layout, 01da99f's) is the control.Phase B's kept-top paths re-hashed in parallel (lever 1, lane i-noepoch-w2)
Prover-only; the proof bytes are the same (digest 07d1bd43… at the default, FAST 840 at 733557a — 82f9046 adds only the knob; the paths, their order and the refusal are unchanged).
BLOCK_WHIR_REHASH_SERIAL=1restores the serial re-hash (a measurement knob).BLOCK OPEN SPLIT(withLAMBDA_VM_BASE_SPLIT=1), phase B's openings by host stage and the kept-top gather and re-hash seconds.The tree's leaves execute while their artifacts are built (R2-i, lane i-noepoch-w2)
Prover-only; the proofs are the same (the top proof's digest equal with and without it, 0da6fea9…, under the fixed trace hash + deterministic grind, FAST 841).
prove_tree_pipelined, the way a prover would run it; there is no production tree driver yet). Its level 0 built all three leaves' artifacts first (three serial card holds, 0.24 s at 1×) and only then let the leaves execute and fill their traces (0.46 s until the first was ready) while the card sat idle (nsys timeline, FAST 839). Execution and fill do not read the artifacts; only the prove does.lfm_proveis now cut into its host half (lfm_execute_and_fill) and its card half (LfmFilled::prove, which asserts the artifacts' hasher is the one the traces were filled for), in the same order; the harness builds the artifacts on a thread of their own and each leaf executes and fills beside them, waiting for its artifacts only to prove.W3_EXEC_BESIDE_ARTIFACTS=0is the control (the order before).A second whole-block field. The W3 readout prints, on the clock between the base and the tree, the block report, a second block frame for the LT heights and the group tables: 0.13–0.14 s at 1× that a prover would not spend. "Whole block" keeps its meaning (every number above includes them); from d169edd on,
W3 RECURSIONalso prints "whole excl. harness readouts" beside it (at d169edd, FAST 841 + 842 pooled: whole 18.54 s, whole excl. harness readouts 18.40 s).The rest laid out in waves; KECCAK_RND built as its tables (W, lane i-m4b)
Prover-only; the proof bytes are the same (digest 07d1bd43… under the fixed trace hash + deterministic grind: FAST 854 at 0bcc1fc, and FAST 855's W arm on this head's code).
BlockOptions::rest_layout_bytes). The tables, their order and the groups are the same (test).split_keccak_rndthen copied into its 2^16-row tables. At 4.13× that copy was a second peak of the same height. The finish now builds the 2^16-row tables directly (BlockOptions::finish_keccak_rnd_chunks), the same tables as the split's (test), and hands none out during the windows, which would change the groups.BLOCK_WHIR_REST_LAYOUT=allandBLOCK_WHIR_KR_FINISH_CHUNKS=0restore the old behaviour.pack_rest_as_laid_out) was re-measured with the waves: no gain (+0.01 s), so it stays off.LAMBDA_VM_BLOCK_MEMLOG=1(off by default). It prints the host memory term by term every half second and at each phase mark, with the line at jemalloc's active peak and each arena's bytes (those two in the lib tests, which install jemalloc).BLOCK REST TABLES(each rest table's layout end and group) and each group's commit start (from@).The tree's first finished leaf executes during phase B (R2-ii, lane i-noepoch-w2)
Prover-only; the proofs are the same (base digest 07d1bd43… and the top proof's digest equal with and without it, 0da6fea9… — FAST 841's — under the fixed trace hash + deterministic grind: FAST 845, and FAST 847 again on fac261f, the merge with W).
block_prove_on_forks_observed; the old entry point passes a no-op, and the proof is the same with any observer). A leaf's arena is built from its groups' words alone (group_arena_words+leaf_arena;block_leaf_arenais built from them). The W3 harness turns the groups into words as they arrive; the first leaf whose groups are all opened while another group is still to come (leaf 0 = groups 1, 5, 6 at 1×, done after group 6) executes and fills its traces on a thread of its own, about 1.5 s before the base ends, and the tree picks it up. Off the clock its arena is checked against the finished proof's.W3_LEAF_DURING_PHASE_B=0is the control.The finish's tables packed as they are built (b2, lane i-m4b)
Prover-only; the proof bytes are the same (digest 07d1bd43… under the fixed trace hash + deterministic grind at both arms: FAST 855 at 0cebe16, and FAST 857 again on 365e3ab, the merge with R2-ii).
BlockOptions::pack_finished). KECCAK_RND is packed in waves of four 2^16-row tables. Phase A takes those tables as narrow columns and uploads them as they are, so the card packs only the groups' other tables. Nothing packed is transposed.BLOCK REST PACKEDcounts the rest: 187 of 191 tables at 4.13×).BLOCK_WHIR_PACK_FINISHED=0restores W.Measurement posture: the card's pool now retains freed memory (baseline shift)
Not a prover change; the proofs are the same (the top proof's digest 0da6fea9… under both postures, fixed trace hash + deterministic grind, FAST 848; 351a773 adds only a test readout).
LAMBDA_VM_MEMPOOL_RELEASE_MB=0. The card's stream-ordered memory pool then hands its freed blocks back to the driver at each synchronize, so a VRAM sampler's total − free reads the live working set. The code default is to retain every freed block (DEFAULT_MEMPOOL_RELEASE_THRESHOLD_BYTES = u64::MAXin math-cuda), and that is what a prover runs.cuStreamSynchronizeandcuMemAllocAsyncwhile the card was idle was ≈ 0.25 s in phase A, 0.42 s in phase B (0.26 s of it in the encode: one 18–33 ms gap per group) and 0.34 s in the tree.W3 DEVICEline).The prover refuses a partition over the group maximum (lane i-m4b)
Prover-only; the proof bytes are the same (digest 07d1bd43…, FAST 858). The prover now refuses a block over
BlockFormat::max_groupsas its groups close, with the verifier's ownInvalidTableCountserror, instead of proving a block the verifier refuses. The median block (90 groups against the cap of 64) spent its whole ≈ 190 s base on such a proof (BIG 564). A test covers both prover paths, with a mutation for each check.