Repository navigation
feat: Rust micro-benchmarks for the hot path - #7
Merged
Merged
Conversation
Adds benches/micro.rs (criterion) covering ring writes, insert_batch, per-reservation cost, FifoSampler, the consumer's sample and drain_round, over four pytree shapes: atari, small, medium, large. Each group drives the crate's own methods directly, with no network or drainer runtime. Source payloads are never all-zero: an unwritten calloc buffer maps the kernel's zero page and made drain_round/large read cache instead of DRAM, ~20% faster than reality. Run with `just bench-rs [filter]`; about a minute for the full suite. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds Rust micro-benchmarks for the hot path in
benches/micro.rs, so changes to the store, sampler and ingress can be measured before they land. There's no network and no drainer runtime: each group calls the crate's own methods directly.memcpy/{shape}PytreeRingBuf::slot_mut+ the store's copy loopinsert_batch; the gap between them is the synchronisation costinsert_batch/{shape}Store::insert_batch, thenStore::samplereservation/{1,16,256}insert_batchwith 1-byte samples, in chunks of that sizefifo_sampler/{1,16,256}FifoSampler::commit_batch+selectSamplergets compared againstsample/{shape}Store::insert_sync(untimed refill), thenStore::sampledrain_round/{shape}ingress::drain_roundover 32 connections, buffers recycled through free poolsShapes:
atari(5 arrays, 55 KiB),small(32 × 8 B),medium(8 × 1 KiB) andlarge(the 44-array, 996.5 KiB pytree thatbench_distributed.pysends). Stores are built the wayPyServer::newbuilds them.Run it with
just bench-rs [regex filter]. The full suite takes about a minute and peaks at about 800 MB resident, almost all of it fromlarge.docs/src/development.mddescribes the groups and how to compare against a saved baseline (--save-baseline main/--baseline main).Two findings from building it
smallis limited by the cost of each copy, not by bandwidth (~2.5 GiB/s). One batch is 8,192 copies of 8 bytes. The array sizes are only known at runtime, so everycopy_nonoverlappingcompiles to acall memcpy@GLIBC(checked in the disassembly), and everyslot_mutre-does a bounds check and a load through theUnsafeCell. Throwaway variants of thememcpyloop onsmall:slot_mut, runtime length)slot_mut, constant 8-byte lengthHoisting the pointers barely helps while the opaque
memcpycall is still there, because the call forces the reloads anyway. Once the call is gone, hoisting gives a further 1.75×. That points to two candidate optimisations for many-small-leaf pytrees: copy small, fixed-size leaves with inline copies, and resolve per-array base pointers once (as in the earlierperf/low-level-optimisationsWIP). This PR doesn't change either.drain_round/largewas ~20% faster thaninsert_batch/largeeven though it does more work. This was a bug in the benchmark itself. Its source buffers werevec![0u8; n]: calloc'd and never written, so every page mapped the kernel's shared zero page, and the copy read cache instead of DRAM. The transport'sread_exactalways writes its buffers, so payloads now are never all zeros. Throwaway check at ~1 MB/sample: zeroed sources 27.1 ms vs written 34.5 ms. After the fix,drain_round/largeis 25.9 ms, level withinsert_batch/largeat 26.3 ms.Numbers (Ryzen 7 5800X, one run)
memcpyinsert_batchdrain_roundsamplereservation1 / 16 / 256: 5.83 µs, 1.76 µs, 1.45 µs per batch.fifo_sampler1 / 16 / 256: 771 ns, 63 ns, 18 ns per batch.Results differ by several percent between separate invocations (
atarimoved about 15% between two runs), so compare a change against a saved baseline rather than across sessions.Test plan
cargo fmt --checkcargo clippy --all-targets --all-features -- -D warningscargo bench --bench micro: all 22 benches runjust bench-rs fifo_samplerfilters as documented🤖 Generated with Claude Code