Skip to content

perf(keccak256): absorb unaligned operands without staging allocations - #3070

Draft
Qumeric wants to merge 5 commits into
develop-v2.x.0-oldfrom
perf/keccak-xorin-unaligned
Draft

Qumeric wants to merge 5 commits into
develop-v2.x.0-oldfrom
perf/keccak-xorin-unaligned

Conversation

@Qumeric

@Qumeric Qumeric commented Jul 25, 2026 •

Copy link
Copy Markdown
Contributor

native_xorin currently handles unaligned pointers and partial words with up to two heap allocations and three copies per absorb. The guest bump allocator never reclaims that memory. Short inputs and incremental hashing commonly take this path because neither the caller's pointer nor the current state offset is necessarily aligned.

Replace those allocations with byte XORs around an aligned XORIN middle:

  1. XOR the prefix up to the state's next 8-byte boundary.
  2. Issue XORIN for the middle, staging its input on the stack only when misaligned.
  3. XOR the trailing partial word.

The fallback uses one stack buffer and at most 14 byte XORs. Overlapping operands snapshot the input before writing, preserving read-before-write semantics; the snapshot's offset makes its middle aligned without recursion or another copy. Accesses stay within the supplied ranges, with no padding requirement. The supported maximum remains 136 bytes, checked in the fallback and in debug builds at entry. The aligned fast path retains identical guest instructions and no stack frame.

Coverage includes known digests, chunked absorption, rate boundaries, and 1,944 direct XORIN cases covering independent alignments, aliasing, overlap in both directions, adjacency, and canaries. The guest program runs in the interpreter, RVR, and CUDA proving tests. A separate differential harness passed 149,056 cases against a byte-wise snapshot reference, covering every length from 0 to 136.

Fresh benchmarks on Ethereum mainnet block 24001988 compare base 56d5e4d with implementation 6a4a53a. Both use openvm-eth fd543064 with identical dependency versions apart from OpenVM, guest toolchain openvm-1.94.1, RVR, and g7.4xlarge with jemalloc.

Metric Base PR Change
Metered instructions 574,390,542 565,581,700 −1.534%
Unpadded main cells 41,947,336,228 40,932,581,006 −2.419%
Metered unpadded memory bytes 690,000,962,509 676,278,916,325 −1.989%
App segments 60 58 −2
Padded proof main cells 52,891,749,454 50,499,067,164 −4.524%
Application proving time 87.944 s 82.449 s −6.248%

Metered comparison · Application proving comparison. Proving time excludes RVR compilation; this is one paired run on one block, not a stable throughput estimate. OpenVM's aligned Keccak microbenchmark exercises the unchanged fast path.

@Qumeric
Qumeric marked this pull request as draft July 25, 2026 07:22
@github-actions

This comment has been minimized.

@github-actions

This comment has been minimized.

@Qumeric
Qumeric force-pushed the perf/keccak-xorin-unaligned branch from 71b32e5 to 0544fab Compare July 25, 2026 08:39
@github-actions

This comment has been minimized.

@github-actions

This comment has been minimized.

@Qumeric
Qumeric force-pushed the perf/keccak-xorin-unaligned branch from 97ef510 to 33ada7b Compare July 29, 2026 07:50
@github-actions

This comment has been minimized.

@Qumeric
Qumeric force-pushed the perf/keccak-xorin-unaligned branch from 33ada7b to 6e274c3 Compare July 29, 2026 08:30
@github-actions

This comment has been minimized.

`native_xorin` fell back to allocating aligned copies of both the state
slice and the input whenever either pointer or the length was not a
multiple of 8, costing two allocations and up to three copies of the
data per absorb. The allocations go through the guest bump allocator,
so they are never reclaimed.

Decompose the unaligned case instead: XOR the bytes below the buffer's
next word boundary and past its last whole word in software (at most 14
bytes total), and absorb the whole aligned words in between with one
XORIN. Only a misaligned input still needs staging, into a stack buffer
sized for one rate block, and only for the instruction's part.

Overlapping operands keep the instruction's semantics: XORIN reads both
ranges before writing (see the executor in
extensions/keccak256/circuit/src/xorin/execution.rs), so the input is
snapshotted up front when the ranges intersect. Staging the snapshot at
the buffer's misalignment preserves relative alignment, letting the
recursive call reach the instruction without a second copy.

The aligned path is unchanged and stays assertion-free; the rate bound
that protects the fixed staging buffers is asserted on the cold path.
The new alignment test program hashes every input misalignment against
every chunking, pins the overlap semantics, and runs in the VM.
@Qumeric
Qumeric force-pushed the perf/keccak-xorin-unaligned branch from 6e274c3 to f71d9d6 Compare September 9, 2026 08:39
@github-actions

This comment has been minimized.

@github-actions

github-actions Bot commented Sep 9, 2026

Copy link
Copy Markdown
Contributor
group app.proof_time_ms app.cycles leaf.proof_time_ms
fibonacci 477 4,000,051 233
keccak 7,640 14,365,133 1,626
sha2_bench 4,412 11,167,961 532
regex 779 4,090,656 217
ecrecover 209 112,210 189
pairing 248 592,827 173
kitchen_sink 2,225 1,979,971 471

Note: cells_used metrics omitted because CUDA tracegen does not expose unpadded trace heights.

Commit: 6a4a53a

Benchmark Workflow

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant