Conversation
Qumeric
marked this pull request as draft
July 25, 2026 07:22
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
Qumeric
force-pushed
the
perf/keccak-xorin-unaligned
branch
from
July 25, 2026 08:39
71b32e5 to
0544fab
Compare
This comment has been minimized.
This comment has been minimized.
Qumeric
force-pushed
the
perf/keccak-xorin-unaligned
branch
from
July 29, 2026 07:40
0544fab to
97ef510
Compare
This comment has been minimized.
This comment has been minimized.
Qumeric
force-pushed
the
perf/keccak-xorin-unaligned
branch
from
July 29, 2026 07:50
97ef510 to
33ada7b
Compare
This comment has been minimized.
This comment has been minimized.
Qumeric
force-pushed
the
perf/keccak-xorin-unaligned
branch
from
July 29, 2026 08:30
33ada7b to
6e274c3
Compare
This comment has been minimized.
This comment has been minimized.
`native_xorin` fell back to allocating aligned copies of both the state slice and the input whenever either pointer or the length was not a multiple of 8, costing two allocations and up to three copies of the data per absorb. The allocations go through the guest bump allocator, so they are never reclaimed. Decompose the unaligned case instead: XOR the bytes below the buffer's next word boundary and past its last whole word in software (at most 14 bytes total), and absorb the whole aligned words in between with one XORIN. Only a misaligned input still needs staging, into a stack buffer sized for one rate block, and only for the instruction's part. Overlapping operands keep the instruction's semantics: XORIN reads both ranges before writing (see the executor in extensions/keccak256/circuit/src/xorin/execution.rs), so the input is snapshotted up front when the ranges intersect. Staging the snapshot at the buffer's misalignment preserves relative alignment, letting the recursive call reach the instruction without a second copy. The aligned path is unchanged and stays assertion-free; the rate bound that protects the fixed staging buffers is asserted on the cold path. The new alignment test program hashes every input misalignment against every chunking, pins the overlap semantics, and runs in the VM.
Qumeric
force-pushed
the
perf/keccak-xorin-unaligned
branch
from
September 9, 2026 08:39
6e274c3 to
f71d9d6
Compare
This comment has been minimized.
This comment has been minimized.
Contributor
Note: cells_used metrics omitted because CUDA tracegen does not expose unpadded trace heights. Commit: 6a4a53a |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
native_xorincurrently handles unaligned pointers and partial words with up to two heap allocations and three copies per absorb. The guest bump allocator never reclaims that memory. Short inputs and incremental hashing commonly take this path because neither the caller's pointer nor the current state offset is necessarily aligned.Replace those allocations with byte XORs around an aligned XORIN middle:
The fallback uses one stack buffer and at most 14 byte XORs. Overlapping operands snapshot the input before writing, preserving read-before-write semantics; the snapshot's offset makes its middle aligned without recursion or another copy. Accesses stay within the supplied ranges, with no padding requirement. The supported maximum remains 136 bytes, checked in the fallback and in debug builds at entry. The aligned fast path retains identical guest instructions and no stack frame.
Coverage includes known digests, chunked absorption, rate boundaries, and 1,944 direct XORIN cases covering independent alignments, aliasing, overlap in both directions, adjacency, and canaries. The guest program runs in the interpreter, RVR, and CUDA proving tests. A separate differential harness passed 149,056 cases against a byte-wise snapshot reference, covering every length from 0 to 136.
Fresh benchmarks on Ethereum mainnet block 24001988 compare base
56d5e4dwith implementation6a4a53a. Both use openvm-ethfd543064with identical dependency versions apart from OpenVM, guest toolchainopenvm-1.94.1, RVR, andg7.4xlargewith jemalloc.Metered comparison · Application proving comparison. Proving time excludes RVR compilation; this is one paired run on one block, not a stable throughput estimate. OpenVM's aligned Keccak microbenchmark exercises the unchanged fast path.