Repository navigation
Backups off the cluster, reconciliation after a restore, and the runbooks - #49
Merged
Merged
Conversation
mumtaz6
force-pushed
the
backup-restore
branch
from
October 4, 2026 06:27
84ed329 to
57d21f2
Compare
A leader that shuts down takes itself out of the ring on its way out. The next leader inherits that ring, and never saw the node fail, so when the node came back and answered its pings, nothing changed from its point of view: it never set the ring again, and the node stayed out for good. A rolling restart that restarted the leader could leave it there. The leader now puts back a node that answers its ping, isn't leaving and isn't in the ring. Only one that answers: a leader that left and is gone hasn't failed yet in the new leader's count, and putting it back would send requests to it until it did. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ooks docs/backup-restore.md: - A checkpoint describes itself (checkpoint.json: node, run, ring and engine versions, key ids, the copy's counts), and a node won't start at one into a running cluster unless its peers were restored from the same run, or it is started with -restored. - The backup run (server/cmd/backup, deploy/kubernetes/backups.yaml) checkpoints every node under one run id, keeps the manifest in each checkpoint, and has each node upload its checkpoint, archived, compressed and encrypted with age, to S3 with Object Lock, locked and tagged by tier (deploy/aws: bucket, lifecycle, put-only writer, reader and escrow policies). - -restored reconciles each topic with its other holders by message digests and counts, idempotently, before the node takes clients. - The security journal: each security change a node makes is fsynced on the node, uploaded every 10 s, and replayed by -journal on a restore, so nothing revoked after the run comes back. - The keyring is escrowed in Secrets Manager (backup escrow-keyring). - The weekly restore test (backup verify, restore-test.yaml) opens each node's checkpoint of the newest run on a scratch server, checks it and a canary written before the run, and pushes its success; backup-alerts.yaml alerts on it. - Runbooks A to D, a restore Job per node, backup runs and fetch -node. The new cluster calls check their sender over TLS, as the others do. A checkpoint counts memdb's records as it copies them: memdb's Size counts a key's versions. storedOn, an e2e helper, reads until the relay is quiet: it stopped at n messages, and missed some when a topic held more. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The e2e tests took 599 s of go test's default 10 minutes with the backup and restore tests in: give them room, within the job's 30. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
mumtaz6
force-pushed
the
backup-restore
branch
from
October 4, 2026 06:42
57d21f2 to
ee7b158
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #48 (checkpoints). Ported from
unitdb_internal(its commits 84bae4d and 7f8a76f). New doc:docs/backup-restore.md.1. Put a node that answers back in the ring (its own commit, a fix to the cluster).
A leader that shuts down takes itself out of the ring. The next leader inherited that ring, never saw the node fail, and so never set the ring again when the node came back. A rolling restart that restarted the leader could leave a node out for good. The leader now puts back a node that answers its ping, isn't leaving, and isn't in the ring. Test:
TestClusterGracefulRestartRejoins, which fails every time without the fix.2. Backups off the cluster, reconciliation after a restore, and the runbooks.
checkpoint.jsonrecords the node, run, ring and engine versions, key ids, and the copy's counts.Cluster.StartedFrom), or it's started with-restored.server/cmd/backupanddeploy/kubernetes/backups.yamlcheckpoint every node under one run id, write a canary first, and keep the manifest in each checkpoint.-upload, each node streams its checkpoint to S3: tar, zstd, then age encryption to the backup key's public half. Each object is locked in compliance mode and tagged by tier (daily, weekly, monthly).deploy/awsholds the bucket, lifecycle rules, and the put-only writer, reader and escrow policies.-restoredsettles each topic with its other holders before the node takes clients. Messages are compared by digest and count, so two identical publishes stay two. Storing is idempotent, so two nodes reconciling each other at once settle on the larger count, not the sum.-journalon a restore, so nothing revoked after the run comes back.backup escrow-keyring).backup verifyandrestore-test.yamlopen each node's checkpoint of the newest run on a scratch server, check it and its canary, then push success to the Pushgateway.backup-alerts.yamlalerts on that.backup runsandfetch -node.Tests
miniobinary): uploads that are locked and encrypted; a client id revoked after the run stays refused through a whole-cluster restore from the bucket, with a control showing the run alone accepts it; the restore test, including a corrupted archive and an escrow missing the run's key; and runbook B step by step.Notes
aws-sdk-go-v2/feature/s3/transfermanager, the replacement for the deprecatedmanager. It's still pre-1.0 (v0.4.13).parseRun).go.mod. A program that imports only the engine doesn't download them.🤖 Generated with Claude Code