Skip to content

feat: Rust micro-benchmarks for the hot path - #7

Merged
sash-a merged 1 commit into
mainfrom
feat/rust-micro-benches
Oct 7, 2026
Merged

sash-a merged 1 commit into
mainfrom
feat/rust-micro-benches

Conversation

@sash-a

@sash-a sash-a commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

Summary

Adds Rust micro-benchmarks for the hot path in benches/micro.rs, so changes to the store, sampler and ingress can be measured before they land. There's no network and no drainer runtime: each group calls the crate's own methods directly.

Group Calls Measures
memcpy/{shape} PytreeRingBuf::slot_mut + the store's copy loop The floor under insert_batch; the gap between them is the synchronisation cost
insert_batch/{shape} Store::insert_batch, then Store::sample The drainer's insert path for one batch
reservation/{1,16,256} insert_batch with 1-byte samples, in chunks of that size Per-reservation cost (CAS, commit, metrics) without the memcpy
fifo_sampler/{1,16,256} FifoSampler::commit_batch + select The baseline any other Sampler gets compared against
sample/{shape} Store::insert_sync (untimed refill), then Store::sample The consumer's release + select + zero-copy view
drain_round/{shape} ingress::drain_round over 32 connections, buffers recycled through free pools A full drainer round

Shapes: atari (5 arrays, 55 KiB), small (32 × 8 B), medium (8 × 1 KiB) and large (the 44-array, 996.5 KiB pytree that bench_distributed.py sends). Stores are built the way PyServer::new builds them.

Run it with just bench-rs [regex filter]. The full suite takes about a minute and peaks at about 800 MB resident, almost all of it from large. docs/src/development.md describes the groups and how to compare against a saved baseline (--save-baseline main / --baseline main).

Two findings from building it

small is limited by the cost of each copy, not by bandwidth (~2.5 GiB/s). One batch is 8,192 copies of 8 bytes. The array sizes are only known at runtime, so every copy_nonoverlapping compiles to a call memcpy@GLIBC (checked in the disassembly), and every slot_mut re-does a bounds check and a load through the UnsafeCell. Throwaway variants of the memcpy loop on small:

Variant Time per batch
As in the store (slot_mut, runtime length) 22.0 µs
Base pointers hoisted, runtime length 20.8 µs
slot_mut, constant 8-byte length 15.6 µs
Base pointers hoisted, constant length 8.9 µs

Hoisting the pointers barely helps while the opaque memcpy call is still there, because the call forces the reloads anyway. Once the call is gone, hoisting gives a further 1.75×. That points to two candidate optimisations for many-small-leaf pytrees: copy small, fixed-size leaves with inline copies, and resolve per-array base pointers once (as in the earlier perf/low-level-optimisations WIP). This PR doesn't change either.

drain_round/large was ~20% faster than insert_batch/large even though it does more work. This was a bug in the benchmark itself. Its source buffers were vec![0u8; n]: calloc'd and never written, so every page mapped the kernel's shared zero page, and the copy read cache instead of DRAM. The transport's read_exact always writes its buffers, so payloads now are never all zeros. Throwaway check at ~1 MB/sample: zeroed sources 27.1 ms vs written 34.5 ms. After the fix, drain_round/large is 25.9 ms, level with insert_batch/large at 26.3 ms.

Numbers (Ryzen 7 5800X, one run)

atari small medium large
memcpy 1.03 ms 21.7 µs 40.1 µs 25.9 ms
insert_batch 956 µs 26.5 µs 43.6 µs 26.3 ms
drain_round 1.05 ms 29.6 µs 45.9 µs 25.9 ms
sample 31 ns 54 ns 34 ns 65 ns

reservation 1 / 16 / 256: 5.83 µs, 1.76 µs, 1.45 µs per batch. fifo_sampler 1 / 16 / 256: 771 ns, 63 ns, 18 ns per batch.

Results differ by several percent between separate invocations (atari moved about 15% between two runs), so compare a change against a saved baseline rather than across sessions.

Test plan

  • cargo fmt --check
  • cargo clippy --all-targets --all-features -- -D warnings
  • cargo bench --bench micro: all 22 benches run
  • just bench-rs fifo_sampler filters as documented

🤖 Generated with Claude Code

Adds benches/micro.rs (criterion) covering ring writes, insert_batch,
per-reservation cost, FifoSampler, the consumer's sample and drain_round,
over four pytree shapes: atari, small, medium, large. Each group drives
the crate's own methods directly, with no network or drainer runtime.

Source payloads are never all-zero: an unwritten calloc buffer maps the
kernel's zero page and made drain_round/large read cache instead of DRAM,
~20% faster than reality.

Run with `just bench-rs [filter]`; about a minute for the full suite.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@sash-a
sash-a merged commit 62d2911 into main Oct 7, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant