Skip to content

memdb: compact the values left in mostly dead blocks - #53

Merged
mumtaz6 merged 2 commits into
masterfrom
selective-compaction
Oct 6, 2026
Merged

mumtaz6 merged 2 commits into
masterfrom
selective-compaction

Conversation

@mumtaz6

@mumtaz6 mumtaz6 commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Adds memdb.DB.Compact, which frees the WAL logs that long-kept values hold, and runs it in the server's message store.

The problem

  • Why logs stay: a block's logs stay in the WAL while it holds a value, and so do the logs of every block that deletes from it, transitively. A delete must outlive the put it deletes, or the deleted value comes back after a restart.
  • The effect: in a store that rewrites and deletes its keys, one long-kept value (a session row, or a message waiting on an absent subscriber) holds the logs of most blocks written after it.
  • Seen in practice: a server's message store on v0.7.0 kept every log since it started, about 12 a minute while idle. Another held 1,379 logs for 1,644 keys, and every one of its 558 blocks was held this way.

Compact

  • What moves: values in past blocks where live values take up at most half the block's data, which moves only what was left behind. Those blocks then hold nothing and are freed, with the logs that waited on them. Mostly live blocks stay.
  • No write pause: each value moves under its key's index lock, as a Put does. A put or delete of that key waits; other writes carry on. A key put again or deleted since it was listed doesn't move.
  • Opt-in: the engine's own memdb doesn't compact, because the engine finds entries by their block.
  • Counters: Varz reports compactions and values moved.

The server's message store compacts at open and every minute.

Tests

  • Pinned blocks: a key kept while others were rewritten and deleted held 152 logs. Compact moved 2 values and left 1 log; all values read correctly after a reopen.
  • Mostly live blocks: nothing moves.
  • Concurrent writes: about 95 compactions overlap writers putting and deleting their own keys. Each key keeps its last value, before and after a reopen.
  • Model test: Compact is now one of its operations.
  • Server adapter: four seconds of churn with a kept message leave 1 log with the compactor running, and 267 without it.
  • Results: memdb, its model test (100 seeds) and memdb under -race -tags lockcheck pass; so do the engine, the server under -race, and e2e.

🤖 Generated with Claude Code

mumtaz6 and others added 2 commits October 6, 2026 13:58
A block's logs stay in the WAL while it holds a value, and so do the logs
of every block that deletes from it, and of those that delete from them:
a delete must outlive the put it deletes. A value kept for long, a
session's row, a message waiting on a subscriber away, then keeps the logs
of most blocks written after it, in a store that rewrites and deletes its
keys. A server's message store on v0.7.0 kept every log since it started,
about 12 a minute idle; another, 1,379 logs for 1,644 keys, each block
held by one value or chained to one.

Compact moves the values left in blocks whose values take up at most half
their data to the current block, so that those blocks hold none and go,
with the logs that waited on them. A value moves under its key's index
lock, as a put does: a put or delete of the key waits, other writes go on,
and a key put again or deleted since it was listed doesn't move. The
engine's own memdb doesn't compact: it finds an entry by its block.

The server's message store compacts at open and every minute. A test
store that kept 152 logs for one value keeps 1 after compacting, moving 2.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The test failed now and then (1 in 30 runs, and in CI): "Compact moved 1;
want none". A write goes to the block of the log's last rotation, every
log interval, not to the block of the time it is made; a round of writes
made as a block began could put its first writes in the block before.
Split between two blocks, a round left one mostly dead, which Compact
rightly moved.

Each round now starts a few log intervals into a new block, with blocks
of 50 ms, far longer than a round. 200 runs pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@mumtaz6
mumtaz6 merged commit ebf8a9a into master Oct 6, 2026
11 checks passed
@mumtaz6
mumtaz6 deleted the selective-compaction branch October 6, 2026 09:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant