Skip to content

ADR-0019: durable meta write through an O_DSYNC fd (one barrier per commit) - #86

Closed
qdequele wants to merge 2 commits into
mainfrom
qdequele/adr-0019-meta-dsync
Closed

qdequele wants to merge 2 commits into
mainfrom
qdequele/adr-0019-meta-dsync

Conversation

@qdequele

@qdequele qdequele commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

What this is

Implements ADR-0019 (Accepted by Quentin 2026-10-01): replace the two
fdatasyncs a default-mode commit issues (C3 data + C5 meta) with one
durable meta write through an O_DSYNC descriptor — LMDB's me_mfd scheme.
One barrier syscall per durable commit instead of two.

Builds on the campaign already merged in #83; branches cleanly off current
main (2 commits, 0 behind).

Mechanism (LMDB parity, clean-room from mdb.c)

  • At env open, a writable non-WRITE_MAP env opens the data file a second time
    O_WRONLY|O_DSYNC|O_CLOEXEC (file::open_meta_sync), stored on
    MmapBacking::meta_sync. Commit step C4 writes the meta page through it; the
    write returns only when the page is durable, so C5's separate fdatasync
    disappears in default mode.
  • New trait seam Backing::write_page_durable(pgno, psize, data) — default =
    write_at_page + sync_data (byte-identical to the old C4+C5 for backends
    without a synchronized write); MmapBacking overrides with the dsync-fd
    pwrite. Routing mirrors LMDB's mask: fused C4+C5 when sync_meta && !write_map; NO_META_SYNC/NO_SYNC route the meta through the plain fd
    unchanged; WRITE_MAP keeps msync C3/C5 (REC-12 untouched).
  • Scrub adopted (OQ3): on a failed/short durable meta write, the slot's
    previous mapped bytes are rewritten through the plain fd before the env is
    poisoned (REC-13), closing the "reopen before power loss reads an
    unacknowledged N" window.
  • Full-page meta write kept (OQ4): Option D (sector-0-only) is left as a
    bench-gated ledger candidate, not folded in.
  • Platforms: Option A on all platforms incl. macOS (Quentin's OQ2 call).
    On macOS O_DSYNC does not force the device cache, so the macOS meta
    barrier is weaker than today's sync_data — macOS is a dev platform, not a
    durability target; stated plainly in SPEC 06. (Reviewers: confirm the macOS
    path matches the approved decision.)

Crash ordering / safety

REC-7 intact: C3 (data fdatasync) still completes before the meta write begins,
and the meta write returning implies durability — so the old H3→H4 "meta
written, not yet durable" window ceases to exist in default mode (strictly
fewer reachable crash states
). The H3 assertion gets stricter (== N) in
default mode; the old {N−1, N} outcome stays fully exercised via fault capture
during the C4 write and under the relaxed modes.

New crash coverage: FaultBacking::write_page_durable (journal → observer
window → fold-self), a broken_dsync mutation self-test that trips REC-18
within 4 default-mode cuts (crash_mutation.rs), and crash_dsync.rs.

No on-disk format change. No API change.

Spec surface to review

docs/SPEC/04-txn-mvcc.md (TXN-61 C4/C5, H3/H4), docs/SPEC/06-recovery.md
(REC-6/7/9/13/17), docs/SPEC/01-flags.md (fd-routing note),
docs/adr/0019-meta-write-dsync.md, and the ADR-0004 cross-reference.

Gate

Gate Result
fmt --check clean
clippy --workspace --all-targets -D warnings clean
test --workspace all suites ok, 0 failures (incl. durability_barriers, crash_dsync, crash_mutation)
miri -p zerodb-core no UB (incl. durability_barriers)
loom 9 model checks ok
stress (180 s) pass
crash-test-full 10,048 cycles (659 image cuts + 1,980 SIGKILL, incl. ADR-0017 spill cycles), 0 violations
fuzz-quick diff_ops 22,705 runs / 684 s 0 divergences; fuzz_image_open ok

⚠️ Performance claim is NOT yet measured on target

This is a durable-commit latency change whose win is device-dependent (FUA vs
flush-emulated). Per the ADR bench plan, the three-column LMDB / before / after
numbers on Graviton4 + EBS io2 (the 0.91× / 1.51 ms-p99 gap this targets)
are pending and must land before merge. The correctness gate below stands on
its own; the perf justification does not yet.

Approved by Quentin 2026-10-01 (ADR-0019 Accepted).

🤖 Generated with Claude Code

Summary by CodeRabbit

  • Reliability
    • In default mode, commits are durable once the metadata write completes, combining steps that previously required a separate sync. On macOS, this does not guarantee that recent commits survive power loss.
    • If a durable metadata write fails, the database attempts to restore the previous metadata and prevents the failed commit from being published.
    • WRITE_MAP retains its separate sync step; NO_SYNC and NO_META_SYNC retain reduced durability.
  • Documentation
    • Updated durability and crash-recovery guidance to reflect these behaviors.

…rier per durable commit)

Implements ADR-0019 (Accepted 2026-10-01), LMDB me_mfd parity:

- zerodb-io: every writable non-WRITE_MAP env opens the data file a second
  time O_WRONLY|O_DSYNC|O_CLOEXEC (MmapBacking::meta_sync) — even under
  NO_SYNC/NO_META_SYNC, as the fork does; READ_ONLY/WRITE_MAP open none.
  MmapBacking::write_page_durable pwrites the meta through it, durable on
  return. All platforms, macOS included (maintainer decision; the weaker
  macOS O_DSYNC guarantee is stated in SPEC 06).
- zerodb-core: Backing::write_page_durable (default = write_at_page +
  sync_data, byte-identical to the old C4+C5 for test backings). The commit
  pipeline fuses C4+C5 through it when sync_meta && !write_map (LMDB's
  routing mask); C5's separate fdatasync is gone in default mode. WRITE_MAP
  keeps the explicit C4 map write + C5 msync (REC-12 untouched).
- LMDB's failed-write scrub (REC-13 as amended): on a failed/short durable
  meta write the slot's previous bytes are rewritten through the plain path
  best-effort, then the env is poisoned — a clean reopen before power loss
  cannot read back an unacknowledged commit.
- FaultBacking models the durable-on-return write: journal pending →
  in-flight observer window ({absent, torn, intact}) → live-view write →
  fold-SELF (never fold-all). Knobs: broken_dsync (mutation self-test, trips
  REC-18.4 on default-mode cuts), fail_durable_writes (scrub tests),
  set_durable_write_observer (the harness's in-flight H3 capture).
- Harness: image H3 cuts in default mode capture mid-durable-write; SIGKILL
  H3 assertion TIGHTENS to 'recovers to N' for default mode (other modes keep
  {N-1, N}); crash_mutation gains the broken_dsync tripwire test.
- Tests: zerodb-io fd-policy/routing/fallback units, fault-model units,
  durability_barriers rewritten for the new call shapes (default = one C3
  barrier + one durable write) plus scrub/poison coverage, and the
  crash_dsync integration test (routing per mode, scrub end-to-end with a
  real-file reopen seeing only the acked state).
- SPEC 04 TXN-61 C4/C5/H3 rows + TXN-64, SPEC 06 REC-6/REC-7 (+ macOS
  platform note)/REC-9/REC-13/REC-17/REC-20, SPEC 01 S6 fd-routing
  realization; ADR-0019 dated implementation note (OQ4: full-page write
  kept; OQ5: the fd lives in MmapBacking). No DIVERGENCES entry: parity.
- short_txn_census: CENSUS_SYNC=1 opens both engines without NO_SYNC for the
  strace syscall census.
@coderabbitai

coderabbitai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: d5c6839a-4dac-4ddd-91fd-e12c839c6140

📥 Commits

Reviewing files that changed from the base of the PR and between 2a69c79 and 469ec77.

📒 Files selected for processing (18)
  • crates/zerodb-core/src/env.rs
  • crates/zerodb-core/src/rwtxn.rs
  • crates/zerodb-core/tests/durability_barriers.rs
  • crates/zerodb-io/src/fault.rs
  • crates/zerodb-io/src/file.rs
  • crates/zerodb-io/src/lib.rs
  • crates/zerodb-oracle/examples/short_txn_census.rs
  • crates/zerodb-oracle/src/crash/image.rs
  • crates/zerodb-oracle/src/crash/sigkill.rs
  • crates/zerodb-oracle/tests/crash_dsync.rs
  • crates/zerodb-oracle/tests/crash_harness_smoke.rs
  • crates/zerodb-oracle/tests/crash_mutation.rs
  • docs/DECISIONS.md
  • docs/SPEC/01-flags.md
  • docs/SPEC/04-txn-mvcc.md
  • docs/SPEC/06-recovery.md
  • docs/adr/0004-write-path.md
  • docs/adr/0019-meta-write-dsync.md

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The commit pipeline now writes default-mode metadata through a durable backing operation. The file-backed implementation can use an O_DSYNC descriptor. Other durability modes retain their existing write and sync paths. Fault and crash tests cover the updated behavior.

Changes

Durable metadata commit path

Layer / File(s) Summary
Backing API and descriptor setup
crates/zerodb-core/src/env.rs, crates/zerodb-io/src/file.rs, crates/zerodb-io/src/lib.rs, docs/SPEC/01-flags.md, docs/adr/0019-meta-write-dsync.md
Backing adds write_page_durable. MmapBacking can use a second O_DSYNC descriptor for writable, non-WRITE_MAP environments. The descriptor policy and fallback behavior are documented and tested.
Commit ordering and barrier tests
crates/zerodb-core/src/rwtxn.rs, crates/zerodb-core/tests/durability_barriers.rs, docs/SPEC/04-txn-mvcc.md, docs/adr/0004-write-path.md, docs/adr/0019-meta-write-dsync.md
Default commits make the metadata write durable at C4 and publish the snapshot afterward. WRITE_MAP retains its C5 sync. Failed durable writes trigger a best-effort restoration of the previous metadata bytes and poison the environment. Tests check call order and failure behavior.
Fault model and crash validation
crates/zerodb-io/src/fault.rs, crates/zerodb-oracle/src/crash/*, crates/zerodb-oracle/tests/*, crates/zerodb-oracle/examples/short_txn_census.rs, docs/SPEC/06-recovery.md, docs/DECISIONS.md
The fault backend models durable writes, in-flight captures, injected failures, and demoted dsync. Crash tests check recovery outcomes, including failed writes and default-mode H3. The census example supports runs with and without NO_SYNC.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~60 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant CommitPipeline
  participant MmapBacking
  participant MetaSyncFD as O_DSYNC meta-sync fd
  CommitPipeline->>MmapBacking: sync_data for data pages
  CommitPipeline->>MmapBacking: write_page_durable for the meta page
  MmapBacking->>MetaSyncFD: write meta page
  MetaSyncFD-->>MmapBacking: durable write returns
  MmapBacking-->>CommitPipeline: write_page_durable returns
  CommitPipeline->>CommitPipeline: run H3 and H4, then publish snapshot
Loading

Merge Risk: ⚪ Minimal · up to 469ec

No actionable merge-blocking defect is established. Default commits use durable metadata writes while other modes retain their documented behavior; target-device performance measurements are still pending.

Security Architecture Review

Security architecture risk: 🟡 Moderate · up to 469ec

The new metadata write handle is not checked against the database handle. If the database pathname is replaced during opening, commits could write to a different file and report success without persisting metadata to the intended database. Exposure depends on who can modify the storage directory.

Retained concerns

  • Medium · security · inferred: The separately opened metadata descriptor is not bound to the validated database descriptor. An actor able to replace the database directory entry between opens could redirect durable metadata writes to another existing file writable by the process. Data writes and validation would still use the original file, allowing unintended file corruption and successful commits whose metadata is absent from the intended database. Previously, data and metadata writes shared one descriptor.
Security review details

Security Blast Radius

  • inferred — The file-identity concern requires control over pathname replacement during environment opening. Its potential target is an existing file the database process can open for writing, with metadata writes confined to the two slot offsets. Process privileges therefore bound exposure; tenant, service and deployment reachability remain unestablished.

Security Findings and Attack Paths

  • inferred — A pathname-replacement race can leave the primary descriptor attached to a valid database while the metadata descriptor follows a replacement entry. Validation then accepts the original mapping, but subsequent synchronized metadata writes use the replacement file. This is a source-supported conditional attack path, not a demonstrated deployment exploit.

Trust Boundaries and Controls

  • observed — The public open path canonicalizes the directory and restricts the data filename to one component. The I/O layer nevertheless resolves that pathname separately for each descriptor, and database validation examines the primary mapping rather than checking descriptor identity.

Resilience and Maintainability Implications

  • inferred — Restoration after a failed durable write is best-effort: its error is discarded, so scrub failure or interruption can still leave unacknowledged metadata visible after reopen. The prior failed-C5 path already permitted that outcome, and this PR adds mitigation rather than establishing a new exposure. The inspected recovery test covers successful restoration, not restoration failure.

Hardening Proposals

  • proposed — Before constructing the backing, compare the opened descriptors' device and inode identities and reject a mismatch. A controlled pathname-replacement test could verify that opening fails instead of creating split data and metadata targets.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 61.64% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 73 functions across 12 files. (6 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: durable meta writes through an O_DSYNC descriptor, with one barrier per commit. This matches the implementation and PR objectives.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 61.64% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 73 functions across 12 files. (6 skipped: 6 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@qdequele

qdequele commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

Parked — no measured performance benefit.

The change is correct and the gate is green (incl. crash-test-full 10k cycles, miri, loom, fuzz), and the mechanism provably fires — an strace census of 2000 durable single-put commits on EBS io2 (Graviton) shows durability barriers drop from 2 → 1 fdatasync per commit (4002 → 2001), with the meta write moving onto the O_DSYNC fd exactly as designed.

But it yields no latency improvement on any device we could measure:

fdatasync/commit commit latency
before (main) ~2.0 1.977 ms
after (ADR-0019) ~1.0 1.995 ms

commit/sync/* on the io2 engine-ladder (3 interleaved rounds, CGU1) was also flat: before ≈ after ≈ LMDB (0.99–1.00×), 0 improved / 0 regressed.

Root cause: on the io2 volume reachable here, fdatasync returns in ~7 µs — it isn't a real device-flush round-trip, so the second barrier was already free and removing it saves nothing. This box does not reproduce the flush-expensive regime the ADR targets.

Not verified: the ADR's named decisive workload — rust-storage-bench YCSB B --fsync, sustained/concurrent, on a volume that genuinely flushes (fdatasync in the ms range). That's the only regime where this could still show the cited 0.91× / p99 gap.

Branch qdequele/adr-0019-meta-dsync is kept. Reopen if/when durable-commit latency becomes a priority and a device/workload demonstrates the win. ADR-0019 stays Accepted (sound design); this is a perf-justification gap, not a correctness issue.

@qdequele qdequele closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant