AMD EPYC 9R14 (Zen 4): the full speed panel

This directory holds the third host of the chart, an AMD EPYC 9R14 (Zen 4, Genoa). Every timed chart row was re-run there with the public timing reproduction, following its README. That covers all 42 manifest rows and the six chart rows outside the manifest. PolyXOR128's raw variant, the four HalftimeHash styles and a ChainHash backend comparison were run alongside. The Xeon and M2 tabs are unchanged except for the ChainHash rows, which were refreshed to the same revision (records/chainhash/refresh-a0116ea).

Host

CPU AMD EPYC 9R14, family 25 model 17 (Zen 4); a 16-vCPU KVM guest, one thread per core
Caches 1 MiB L2 per core; two 32 MiB L3s, one shared by CPUs 0–7 and one by CPUs 8–15
ISA AVX-512 (F, VL, BW, DQ, IFMA, VBMI), VPCLMULQDQ, VAES, GFNI, SHA-NI
Clock TSC 2.600 GHz (tsc_known_freq, clocksource tsc). A busy core runs at 3.66 GHz (boost), measured with a dependent-add loop; a user-mode perf_event_open cycle counter confirms 1.0002 cycles per add.
OS Rocky Linux 9.8, kernel 5.14.0-687
Compiler GCC 11.5.0 (Red Hat 11.5.0-14), -O3 -march=native (resolves to znver4); Rust 1.98.1 for PolyXOR128 (backend hash_blocks_avx512)
Pinning nice -n 10 taskset -c 8-15 for every timing, and for the builds

SMHasher3's x86 timer counts TSC ticks. B/cycle on this tab therefore means bytes per 2.600 GHz tick. Wall-clock GB/s is 2.6 × B/cycle, and bytes per core cycle at 3.66 GHz is 0.71 × B/cycle. The printed GiB/s assumes 3.5 GHz, as on the other hosts. The Xeon's TSC runs at 2.900 GHz, and the M2 uses a calibrated clock, so compare hashes within one tab.

Protocol

This is the reproduction's x86 rule, the same as on the Xeon. Each row gets two complete serial passes of unmodified SMHasher3 NAME --test=Speed. The cell takes the higher bulk average (262144-byte keys, eight alignments) and, independently, the lower 1–31-byte average.

Each start waits until no other SMHasher3 process runs on CPUs 8–15. Every 5 s the runner samples for overlapping runs; none was excluded. Another job ran on CPUs 0–7, the other L3, during the measurements, so no machine-wide load threshold was applied. Load1 at launch ranged 0.95–6.35 and is recorded per run.

The rows outside the manifest ran on the same binary with the same selection. There, a whole-machine check discarded and repeated three runs that overlapped another SMHasher3 process (evidence/extra/discarded.jsonl). No bulk spread exceeds 2%. The largest small-key spreads are 3.5% (HalftimeHash-512) and 3.2% (polyxor-128.raw), so no row needed the

5% re-run.

The ChainHash pin of the reproduction moved to a0116ea during the run. ChainHash does not enter any other row. The build was repeated at the new pin, and VerifyAll was asserted: 0x66672BD6 and 0x1FCA728C, unchanged. The chainhash and chainhash-128 rows were then re-timed on that binary (evidence/panel-chainhash). Their earlier values at 30c0111 are kept in speeds_zen4.json#/meta.

Results

Verification is SMHasher3's VerifyAll value on this build. Sanity is the reproduction's native Sanity record. The four Sanity failures (foldhash ×2, shipped HalftimeHash24 and Marvin32) are the ones the reproduction README documents for x86. The two HalftimeHash24 wrappers have zero registered constants. Their computed values equal the Xeon's (0x8A105352 shipped, 0x3F1372EA fixed).

Row Registration Bulk B/tick GB/s at TSC 2.6 GHz 1–31 B cycles Spread bulk / small Verification Sanity
city CityHash-64 6.82 17.7 34.57 0.00% / 0.00% 0x5FABC5C5 PASS PASS
farm FarmHash-64.NA 6.74 17.5 34.48 0.00% / 0.00% 0xEBC4A679 PASS PASS
murmur MurmurHash3-128 2.95 7.7 36.88 0.00% / 0.00% 0x6384BA69 PASS PASS
mx3 mx3.v3 4.34 11.3 32.52 0.00% / 0.18% 0x7B287B65 PASS PASS
fasthash-64 fasthash-64 2.85 7.4 26.14 0.00% / 0.08% 0xA16231A7 PASS PASS
fasthash-32 fasthash-32 2.85 7.4 27.60 0.00% / 0.00% 0xE9481AFC PASS PASS
muse MuseAir 9.95 25.9 13.27 0.00% / 0.00% 0xF89F1683 PASS PASS
muse-v2 MuseAir-v2 7.33 19.1 19.92 0.14% / 0.05% 0x7140CABC PASS PASS
komi komihash 7.71 20.0 21.65 0.00% / 0.00% 0x8157FF6D PASS PASS
t1ha t1ha2-64 7.00 18.2 28.59 0.00% / 0.00% 0x8F16C948 PASS PASS
a5 a5hash 3.80 9.9 11.04 0.26% / 0.00% 0xADDE79B3 PASS PASS
a5wide a5hash-128 10.00 26.0 19.36 0.00% / 0.00% 0x89406B11 PASS PASS
rapid3 rapidhash 11.61 30.2 19.84 0.09% / 0.00% 0x1FDC65EE PASS PASS
foldhash-fast foldhash-fast 11.41 29.7 14.03 0.80% / 0.07% 0xA167BA5E PASS FAIL
foldhash-quality foldhash-quality 11.25 29.2 17.47 0.18% / 0.06% 0xE269316E PASS FAIL
mum mum3.exact.unroll3 8.80 22.9 17.30 1.97% / 0.06% 0x8BD72B8C PASS PASS
mir mir.exact 3.25 8.5 24.50 0.62% / 0.00% 0x00A393C8 PASS PASS
xxh3-64 XXH3-64 18.81 48.9 20.38 0.16% / 0.05% 0x1AAEE62C PASS PASS
xxh3-128 XXH3-128 18.79 48.9 23.66 0.80% / 0.04% 0x288DAA94 PASS PASS
highway HighwayHash-64 3.88 10.1 59.71 0.00% / 0.00% 0xF3246108 PASS PASS
spooky SpookyHash2-64 6.65 17.3 34.76 0.00% / 0.03% 0x972C4BDC PASS PASS
pengyhash pengyhash 6.15 16.0 52.48 0.00% / 0.02% 0x861A1254 PASS PASS
nmhash32 NMHASH 11.43 29.7 32.07 0.00% / 0.00% 0x12A30553 PASS PASS
nmhash32x NMHASHX 11.45 29.8 20.73 0.09% / 0.05% 0xA8580227 PASS PASS
gx gxhash-64 30.79 80.1 31.28 0.69% / 0.00% 0x48F84240 PASS PASS
ahash rust-ahash 0.78 2.0 72.07 0.00% / 0.03% 0x3BF4383B PASS PASS
siphash-1-3 SipHash-1-3 1.03 2.7 62.06 0.00% / 0.00% 0x8936B193 PASS PASS
siphash-2-4 SipHash-2-4 0.55 1.4 86.30 0.00% / 0.03% 0x57B661ED PASS PASS
halftime24 HalftimeHash24-shipped 15.96 41.5 54.51 0.06% / 0.02% 0x8A105352 unverifiable (zero constant) FAIL
go-maphash GoMapHash 22.38 58.2 74.12 0.00% / 0.00% 0x710289AB PASS PASS
abseil-hash AbseilHash-default 9.77 25.4 14.06 0.21% / 0.00% 0x07203CDB PASS PASS
dotnet-marvin Marvin32 1.14 3.0 22.00 0.00% / 0.05% 0x306A5169 PASS FAIL
polymur polymurhash 6.13 15.9 28.32 0.99% / 0.00% 0x0722B1A7 PASS PASS
poly1305 poly1305-hash 5.05 13.1 195.97 0.20% / 0.06% 0xBD015C42 PASS PASS
ghash ghash 8.91 23.2 1418.97 0.22% / 0.49% 0x1F397201 PASS PASS
umash UMASH-64 9.10 23.7 26.91 0.00% / 0.15% 0x36A264CD PASS PASS
umash128 UMASH-128 6.39 16.6 31.13 0.00% / 0.03% 0x63857D05 PASS PASS
clhash CLhash 10.55 27.4 31.23 0.00% / 0.06% 0x2E554CB4 PASS PASS
chainhash chainhash 35.09 91.2 62.46 0.00% / 0.03% 0x66672BD6 PASS PASS
halftime24-fixed HalftimeHash24-fixed 12.44 32.3 71.63 0.16% / 0.13% 0x3F1372EA unverifiable (zero constant) PASS
chainhash128 chainhash-128 19.32 50.2 116.78 1.79% / 0.02% 0x1FCA728C PASS PASS
polyxor polyxor-128 19.79 51.5 129.96 0.00% / 0.03% 0xA9574CA8 PASS PASS
(extra) wyhash 11.20 29.1 17.46 0.27% / 0.06% 0x9DAE7DD3 PASS —
(extra) XXH-64 4.14 10.8 35.71 0.00% / 0.00% 0x8F8224C4 PASS —
(extra) XXH-32 2.85 7.4 27.31 0.00% / 0.22% 0x6FD78385 PASS —
(extra) MurmurHash2-64 2.85 7.4 26.60 0.00% / 0.04% 0x1F0D3804 PASS —
(extra) MurmurHash2-32 1.42 3.7 22.52 0.00% / 0.00% 0x27864C1E PASS —
(extra) MurmurHash2a 1.42 3.7 25.80 0.00% / 0.08% 0x7FBD4396 PASS —
(extra) HalftimeHash-64 3.37 8.8 65.03 0.00% / 0.34% 0xED42E424 PASS —
(extra) HalftimeHash-128 11.35 29.5 63.92 0.18% / 0.42% 0x952DF141 PASS —
(extra) HalftimeHash-256 17.12 44.5 63.98 0.41% / 0.53% 0x912330EA PASS —
(extra) HalftimeHash-512 17.45 45.4 70.48 0.40% / 3.49% 0x1E0F99EA PASS —
(extra) polyxor-128.raw 19.83 51.6 96.43 0.00% / 3.15% 0x23A6E8BB PASS —

ChainHash backends on Zen 4

The headers dispatch to ZMM when AVX-512 is present. Forcing each x86 backend in one binary (speeds_zen4_chainhash_backends.json):

Registration Bulk B/tick 1–31 B cycles
chainhash (run-time dispatch) 35.14 61.78
chainhash.zmm 35.15 61.75
chainhash.ymm 20.14 61.72
chainhash.xmm 10.07 61.79
chainhash-128 (run-time dispatch) 19.35 116.58
chainhash-128.zmm 19.38 127.77
chainhash-128.ymm 12.40 127.53
chainhash-128.xmm 6.19 127.94

ZMM is 1.75× (ChainHash) and 1.56× (ChainHash-128) faster than YMM in bulk on Zen 4. Zen 4 executes 512-bit VPCLMULQDQ as two 256-bit halves, but the wider loop still wins, so the ZMM dispatch is right on this CPU.

The forced ChainHash-128 registrations pass the backend as a compile-time constant. Their short-input cost (about 128 cycles) is higher than the dispatched entry point's 116.6; bulk is unaffected.

Files

File What it is
speeds_zen4.json The 42 manifest rows: selected cells, both passes, spreads, load, Sanity records (reproduction benchmark.py output)
speeds_zen4_extra.json wyhash, XXH64, XXH32, MurmurHash64A, MurmurHash2, MurmurHash2A, the four HalftimeHash styles, polyxor-128.raw; with per-length 1–31 B costs
speeds_zen4_chainhash_backends.json The backend comparison above, with 0a03c63 and PolyXOR128/XXH3-128 controls
provenance.json CPU, clocks, kernel, compiler and flags, rustc, patch hashes, binary SHA-256s, pin set, load
verifyall-30c0111.txt, verifyall-a0116ea.txt VerifyAll of the two reproduction builds
evidence/panel/, evidence/panel-chainhash/, evidence/extra/, evidence/backends/ Raw --test=Speed output per row and pass (NAME.runN.txt), run and gate records, Sanity logs
bench/ The TSC and core-clock probe with its before and after output, and the provenance collector
merge_epyc.py Writes these cells into the article's data.json and records/speeds.json

Reproduce with the repository's README (--host EPYC9R14 --cpus 8-15 --nice 10). The extra rows use records/chainhash/refresh-a0116ea/bench/run_speed.py on the same binary.