Skip to content

Backups off the cluster, reconciliation after a restore, and the runbooks - #49

Merged
mumtaz6 merged 3 commits into
masterfrom
backup-restore
Oct 4, 2026
Merged

mumtaz6 merged 3 commits into
masterfrom
backup-restore

Conversation

@mumtaz6

@mumtaz6 mumtaz6 commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #48 (checkpoints). Ported from unitdb_internal (its commits 84bae4d and 7f8a76f). New doc: docs/backup-restore.md.

1. Put a node that answers back in the ring (its own commit, a fix to the cluster).
A leader that shuts down takes itself out of the ring. The next leader inherited that ring, never saw the node fail, and so never set the ring again when the node came back. A rolling restart that restarted the leader could leave a node out for good. The leader now puts back a node that answers its ping, isn't leaving, and isn't in the ring. Test: TestClusterGracefulRestartRejoins, which fails every time without the fix.

2. Backups off the cluster, reconciliation after a restore, and the runbooks.

  • Checkpoints describe themselves. checkpoint.json records the node, run, ring and engine versions, key ids, and the copy's counts.
  • The start guard. A node won't start at a checkpoint into a running cluster unless its peers were restored from the same run (Cluster.StartedFrom), or it's started with -restored.
  • Backup runs. server/cmd/backup and deploy/kubernetes/backups.yaml checkpoint every node under one run id, write a canary first, and keep the manifest in each checkpoint.
  • Uploads. With -upload, each node streams its checkpoint to S3: tar, zstd, then age encryption to the backup key's public half. Each object is locked in compliance mode and tagged by tier (daily, weekly, monthly). deploy/aws holds the bucket, lifecycle rules, and the put-only writer, reader and escrow policies.
  • Reconciliation. -restored settles each topic with its other holders before the node takes clients. Messages are compared by digest and count, so two identical publishes stay two. Storing is idempotent, so two nodes reconciling each other at once settle on the larger count, not the sum.
  • The security journal. Each security change a node makes is fsynced on the node, uploaded every 10 s, and replayed by -journal on a restore, so nothing revoked after the run comes back.
  • Keyring escrow in Secrets Manager (backup escrow-keyring).
  • The weekly restore test. backup verify and restore-test.yaml open each node's checkpoint of the newest run on a scratch server, check it and its canary, then push success to the Pushgateway. backup-alerts.yaml alerts on that.
  • Runbooks A to D, a restore Job per node, backup runs and fetch -node.
  • TLS sender checks. The new cluster calls check their sender over TLS, as the others do.

Tests

  • Unit: archive, retention, keyring escrow, the journal surviving a restart, and the canary.
  • e2e: the run and its manifest, the guard, reconciliation (each holder of each topic checked alone), and a lost node.
  • e2e against MinIO (skipped without the minio binary): uploads that are locked and encrypted; a client id revoked after the run stays refused through a whole-cluster restore from the bucket, with a control showing the run alone accepts it; the restore test, including a corrupted archive and an escrow missing the run's key; and runbook B step by step.
  • Full suite on this stack: vet, unit and e2e (706 s) all pass.

Notes

  • Uploads use aws-sdk-go-v2/feature/s3/transfermanager, the replacement for the deprecated manager. It's still pre-1.0 (v0.4.13).
  • A run id from a request is parsed as a time and formatted anew before it goes into a path (parseRun).
  • The new dependencies (AWS SDK, age, zstd) go into the module's go.mod. A program that imports only the engine doesn't download them.

🤖 Generated with Claude Code

Comment thread server/internal/db/checkpoint_info.go Fixed
Comment thread server/internal/db/checkpoint_info.go Fixed
Comment thread server/internal/db/checkpoint_info.go Fixed
Comment thread server/internal/db/checkpoint_info.go Fixed
mumtaz6 and others added 3 commits October 4, 2026 12:12
A leader that shuts down takes itself out of the ring on its way out. The
next leader inherits that ring, and never saw the node fail, so when the
node came back and answered its pings, nothing changed from its point of
view: it never set the ring again, and the node stayed out for good. A
rolling restart that restarted the leader could leave it there.

The leader now puts back a node that answers its ping, isn't leaving and
isn't in the ring. Only one that answers: a leader that left and is gone
hasn't failed yet in the new leader's count, and putting it back would
send requests to it until it did.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ooks

docs/backup-restore.md:

- A checkpoint describes itself (checkpoint.json: node, run, ring and
  engine versions, key ids, the copy's counts), and a node won't start at
  one into a running cluster unless its peers were restored from the same
  run, or it is started with -restored.
- The backup run (server/cmd/backup, deploy/kubernetes/backups.yaml)
  checkpoints every node under one run id, keeps the manifest in each
  checkpoint, and has each node upload its checkpoint, archived,
  compressed and encrypted with age, to S3 with Object Lock, locked and
  tagged by tier (deploy/aws: bucket, lifecycle, put-only writer, reader
  and escrow policies).
- -restored reconciles each topic with its other holders by message
  digests and counts, idempotently, before the node takes clients.
- The security journal: each security change a node makes is fsynced on
  the node, uploaded every 10 s, and replayed by -journal on a restore,
  so nothing revoked after the run comes back.
- The keyring is escrowed in Secrets Manager (backup escrow-keyring).
- The weekly restore test (backup verify, restore-test.yaml) opens each
  node's checkpoint of the newest run on a scratch server, checks it and
  a canary written before the run, and pushes its success;
  backup-alerts.yaml alerts on it.
- Runbooks A to D, a restore Job per node, backup runs and fetch -node.

The new cluster calls check their sender over TLS, as the others do. A
checkpoint counts memdb's records as it copies them: memdb's Size counts a
key's versions. storedOn, an e2e helper, reads until the relay is quiet:
it stopped at n messages, and missed some when a topic held more.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The e2e tests took 599 s of go test's default 10 minutes with the
backup and restore tests in: give them room, within the job's 30.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@mumtaz6
mumtaz6 changed the base branch from checkpoints to master October 4, 2026 16:05
@mumtaz6
mumtaz6 merged commit 4bc997e into master Oct 4, 2026
6 checks passed
@mumtaz6
mumtaz6 deleted the backup-restore branch October 6, 2026 09:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants