diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 5acc3ff..d83433f 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -25,7 +25,7 @@ jobs: go-version: ${{ matrix.go != 'go.mod' && matrix.go || '' }} go-version-file: ${{ matrix.go == 'go.mod' && 'go.mod' || '' }} - run: go vet ./... - - run: go test ./... + - run: go test -timeout 25m ./... staticcheck: runs-on: ubuntu-latest diff --git a/deploy/aws/README.md b/deploy/aws/README.md new file mode 100644 index 0000000..e23826b --- /dev/null +++ b/deploy/aws/README.md @@ -0,0 +1,105 @@ +# unitdb's backups on AWS + +What a cluster's off-site backups need on AWS (docs/backup-restore.md), +made once per cluster by an AWS admin: a bucket with Object Lock, a +put-only identity for the cluster (an access key in a Kubernetes Secret +here; any of the AWS SDK's credential sources will do), and the backup key +and the escrowed keyring in Secrets Manager, readable only by the weekly +restore test's identity and named admins. A cluster per AWS account keeps +staging and production apart. + +Below, `CLUSTER` is the cluster's name (`staging`, `production`): its +prefix in the bucket and in the secrets' names. `BUCKET` is the bucket, +`REGION` its region, `ACCOUNT` the account id. Replace them in the JSON +files here before using them. + +## 1. The bucket: Object Lock, and a lifecycle by tier + +Object Lock can only be turned on when the bucket is made. There is no +default retention: each object is locked in compliance mode by the node +that uploads it, for as long as its tier is kept, so that daily runs can +expire after a week while monthly ones are kept 13 months (server/internal/backup: +daily 8 days, weekly 29, monthly and the journal 396). Compliance mode: not +even the account's root user deletes a locked version before its date. + +```sh +aws s3api create-bucket --bucket BUCKET --region REGION \ + --create-bucket-configuration LocationConstraint=REGION \ + --object-lock-enabled-for-bucket +aws s3api put-public-access-block --bucket BUCKET \ + --public-access-block-configuration BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true +aws s3api put-bucket-encryption --bucket BUCKET \ + --server-side-encryption-configuration '{"Rules":[{"ApplyServerSideEncryptionByDefault":{"SSEAlgorithm":"AES256"}}]}' +aws s3api put-bucket-lifecycle-configuration --bucket BUCKET --lifecycle-configuration file://lifecycle.json +``` + +The lifecycle rules (`lifecycle.json`) expire each object by its `tier` +tag once its lock has passed; the old versions go a day later. + +## 2. The cluster's identity: put only + +An IAM user, `unitdb-backup-writer-CLUSTER`, with `writer-policy.json` +only: it puts objects, with their lock and tag, under `CLUSTER/`, and can +neither read, list nor delete them. A compromised cluster can't read or +erase its backups. + +```sh +aws iam create-user --user-name unitdb-backup-writer-CLUSTER +aws iam put-user-policy --user-name unitdb-backup-writer-CLUSTER \ + --policy-name unitdb-backup-put-only --policy-document file://writer-policy.json +aws iam create-access-key --user-name unitdb-backup-writer-CLUSTER +kubectl -n unitdb create secret generic unitdb-backup-writer \ + --from-literal=AWS_ACCESS_KEY_ID= \ + --from-literal=AWS_SECRET_ACCESS_KEY= +``` + +Check the policy before the first run, with the IAM policy simulator or a +test upload: the nodes' uploads use exactly `s3:PutObject` (multipart +included), `s3:PutObjectRetention`, `s3:PutObjectTagging` and, when an +upload fails, `s3:AbortMultipartUpload`. The e2e tests run against MinIO as +its root user, so they don't check the policy. + +Rotate the key by making a second one, updating the Secret, restarting the +nodes, and deleting the first. + +## 3. The backup key + +Archives are encrypted with [age](https://age-encryption.org) to the backup +key's public half; the private half is only in Secrets Manager. + +```sh +age-keygen -o backup-key.txt # prints the public key: age1... +aws secretsmanager create-secret --region REGION \ + --name unitdb/CLUSTER/backup-key --secret-string file://backup-key.txt +shred -u backup-key.txt +``` + +The public half goes into the nodes' environment as `BACKUP_AGE_RECIPIENT` +(docs/backup-restore.md lists the nodes' settings). + +## 4. The keyring's escrow + +A backup restores only with the keyring of its time. Put the keyring in +escrow before every change of the cluster's `UNITDB_KEYRING`, with an +admin's identity (`escrow-policy.json`): + +```sh +BACKUP_CLUSTER=CLUSTER backup escrow-keyring -keyring keyring.json +``` + +It keeps every key it ever held: a key gone from the keyring stays as a +`read` key, and a key id with another key is refused. Only then update +the cluster's Secret. + +## 5. Readers + +`reader-policy.json` reads the bucket under `CLUSTER/` and the two secrets: +for the weekly restore test's identity and the named admins' roles, +nothing else. A restore (docs/backup-restore.md): + +```sh +export BACKUP_S3_BUCKET=BUCKET BACKUP_CLUSTER=CLUSTER +backup fetch -run -out restore/run # each node's checkpoint +backup fetch-journal -since -out restore/journal +# then each node: -db_path=restore/run/ -restored -journal=restore/journal +``` diff --git a/deploy/aws/escrow-policy.json b/deploy/aws/escrow-policy.json new file mode 100644 index 0000000..bb15f13 --- /dev/null +++ b/deploy/aws/escrow-policy.json @@ -0,0 +1,15 @@ +{ + "Version": "2012-10-17", + "Statement": [ + { + "Sid": "UnitdbKeyringEscrow", + "Effect": "Allow", + "Action": [ + "secretsmanager:GetSecretValue", + "secretsmanager:PutSecretValue", + "secretsmanager:CreateSecret" + ], + "Resource": "arn:aws:secretsmanager:REGION:ACCOUNT:secret:unitdb/CLUSTER/keyring-*" + } + ] +} diff --git a/deploy/aws/lifecycle.json b/deploy/aws/lifecycle.json new file mode 100644 index 0000000..d48a2cf --- /dev/null +++ b/deploy/aws/lifecycle.json @@ -0,0 +1,15 @@ +{ + "Rules": [ + {"ID": "daily", "Status": "Enabled", "Filter": {"Tag": {"Key": "tier", "Value": "daily"}}, + "Expiration": {"Days": 8}, "NoncurrentVersionExpiration": {"NoncurrentDays": 1}}, + {"ID": "weekly", "Status": "Enabled", "Filter": {"Tag": {"Key": "tier", "Value": "weekly"}}, + "Expiration": {"Days": 29}, "NoncurrentVersionExpiration": {"NoncurrentDays": 1}}, + {"ID": "monthly", "Status": "Enabled", "Filter": {"Tag": {"Key": "tier", "Value": "monthly"}}, + "Expiration": {"Days": 396}, "NoncurrentVersionExpiration": {"NoncurrentDays": 1}}, + {"ID": "journal", "Status": "Enabled", "Filter": {"Tag": {"Key": "tier", "Value": "journal"}}, + "Expiration": {"Days": 396}, "NoncurrentVersionExpiration": {"NoncurrentDays": 1}}, + {"ID": "delete-markers", "Status": "Enabled", "Filter": {}, + "Expiration": {"ExpiredObjectDeleteMarker": true}, + "AbortIncompleteMultipartUpload": {"DaysAfterInitiation": 2}} + ] +} diff --git a/deploy/aws/reader-policy.json b/deploy/aws/reader-policy.json new file mode 100644 index 0000000..892cade --- /dev/null +++ b/deploy/aws/reader-policy.json @@ -0,0 +1,27 @@ +{ + "Version": "2012-10-17", + "Statement": [ + { + "Sid": "UnitdbBackupRead", + "Effect": "Allow", + "Action": ["s3:GetObject", "s3:GetObjectVersion"], + "Resource": "arn:aws:s3:::BUCKET/CLUSTER/*" + }, + { + "Sid": "UnitdbBackupList", + "Effect": "Allow", + "Action": "s3:ListBucket", + "Resource": "arn:aws:s3:::BUCKET", + "Condition": {"StringLike": {"s3:prefix": ["CLUSTER/*"]}} + }, + { + "Sid": "UnitdbBackupKeys", + "Effect": "Allow", + "Action": "secretsmanager:GetSecretValue", + "Resource": [ + "arn:aws:secretsmanager:REGION:ACCOUNT:secret:unitdb/CLUSTER/backup-key-*", + "arn:aws:secretsmanager:REGION:ACCOUNT:secret:unitdb/CLUSTER/keyring-*" + ] + } + ] +} diff --git a/deploy/aws/writer-policy.json b/deploy/aws/writer-policy.json new file mode 100644 index 0000000..d0bc8c6 --- /dev/null +++ b/deploy/aws/writer-policy.json @@ -0,0 +1,16 @@ +{ + "Version": "2012-10-17", + "Statement": [ + { + "Sid": "UnitdbBackupPutOnly", + "Effect": "Allow", + "Action": [ + "s3:PutObject", + "s3:PutObjectRetention", + "s3:PutObjectTagging", + "s3:AbortMultipartUpload" + ], + "Resource": "arn:aws:s3:::BUCKET/CLUSTER/*" + } + ] +} diff --git a/deploy/kubernetes/backup-alerts.yaml b/deploy/kubernetes/backup-alerts.yaml new file mode 100644 index 0000000..db04733 --- /dev/null +++ b/deploy/kubernetes/backup-alerts.yaml @@ -0,0 +1,68 @@ +# Alerts on unitdb's backups (docs/backup-restore.md), for the Prometheus +# Operator. The metrics: each node's /_metrics on its monitor +# port (checkpoints, uploads, the security journal), the Pushgateway (the +# weekly restore test), and kube-state-metrics (the backup job). +apiVersion: monitoring.coreos.com/v1 +kind: PrometheusRule +metadata: + name: unitdb-backups + namespace: unitdb +spec: + groups: + - name: unitdb-backups + rules: + # A node's last checkpoint is more than a day old: the daily run + # (backups.yaml) didn't reach it. + - alert: UnitdbCheckpointStale + expr: time() - min by (pod) (unitdb_checkpoint_last_success_timestamp_seconds) > 26 * 3600 + for: 10m + labels: + severity: critical + annotations: + summary: "{{ $labels.pod }}: no checkpoint for over 26 h" + - alert: UnitdbCheckpointFailing + expr: increase(unitdb_checkpoint_failures_total[2h]) > 0 + labels: + severity: warning + annotations: + summary: "{{ $labels.pod }}: a checkpoint failed (see its logs, context checkpoint)" + # A node's last upload to the bucket is more than a day old. + - alert: UnitdbBackupUploadStale + expr: time() - min by (pod) (unitdb_backup_upload_last_success_timestamp_seconds) > 26 * 3600 + for: 10m + labels: + severity: critical + annotations: + summary: "{{ $labels.pod }}: no backup uploaded for over 26 h" + - alert: UnitdbBackupJobFailed + expr: kube_job_status_failed{namespace="unitdb", job_name=~"unitdb-backup-.*"} > 0 + labels: + severity: critical + annotations: + summary: "the backup run {{ $labels.job_name }} failed: a node's checkpoint or upload (see the Job's output, its manifest)" + # Security changes not in the bucket: a whole-cluster restore now + # would bring back what they revoked. + - alert: UnitdbSecurityJournalLagging + expr: max by (pod) (unitdb_security_journal_lag_seconds) > 300 + for: 5m + labels: + severity: critical + annotations: + summary: "{{ $labels.pod }}: security changes over 5 min old aren't in the bucket" + - alert: UnitdbSecurityJournalNotRecording + expr: increase(unitdb_security_journal_record_errors_total[10m]) > 0 + labels: + severity: critical + annotations: + summary: "{{ $labels.pod }}: a security change couldn't be written to the journal" + # The weekly restore test hasn't passed for over 8 days, or never + # has: this fires from deployment until its first pass. + - alert: UnitdbRestoreTestStale + expr: | + (time() - max by (cluster) (unitdb_backup_restore_verified_timestamp_seconds) > 8 * 86400) + or absent(unitdb_backup_restore_verified_timestamp_seconds) + for: 1h + labels: + severity: critical + annotations: + summary: "no backup has been proven to restore for over 8 days (restore-test.yaml)" diff --git a/deploy/kubernetes/backups.yaml b/deploy/kubernetes/backups.yaml new file mode 100644 index 0000000..3be809f --- /dev/null +++ b/deploy/kubernetes/backups.yaml @@ -0,0 +1,57 @@ +# unitdb's daily backup run (docs/backup-restore.md): at 21:30 UTC, the +# backup command (server/cmd/backup, in the unitdb image) has every +# node take a checkpoint of one run, one node after another, each with up to +# three tries, into checkpoints/ckpt-- on its volume, and keeps +# the run's manifest.json (each node's checkpoint.json, or why it failed) in +# every node's checkpoint. Then each node uploads its checkpoint, encrypted, +# and the manifest, to the bucket (-upload; deploy/aws/README.md). A node +# that failed either fails the job: a restore needs every node of one run. +# A volume snapshot of the nodes' claims, if you take one, goes after it: +# only the checkpoints in it restore, not the live store beside them. +# +# The job's exit status is its metric: alert on a failed Job, or on +# kube_cronjob_status_last_successful_time older than 26 h. +# +# Restoring: docs/backup-restore.md. One node lost: an empty claim, not a +# checkpoint. The whole cluster: every node at its checkpoint of one run. +apiVersion: batch/v1 +kind: CronJob +metadata: + name: unitdb-backup + namespace: unitdb +spec: + schedule: "30 21 * * *" + concurrencyPolicy: Forbid + successfulJobsHistoryLimit: 3 + failedJobsHistoryLimit: 3 + jobTemplate: + spec: + # The command retries each node itself; a second Job would start a new + # run. + backoffLimit: 0 + template: + spec: + restartPolicy: Never + # No Service links in the environment: a Service named web would set + # WEB_PORT=tcp://..., which unitdb reads as its own port. + enableServiceLinks: false + containers: + - name: backup + image: unitdb:latest + command: ["/backup"] + args: + - -upload + - -nodes + - http://unitdb-0.unitdb.unitdb.svc.cluster.local:7374,http://unitdb-1.unitdb.unitdb.svc.cluster.local:7374,http://unitdb-2.unitdb.unitdb.svc.cluster.local:7374 + env: + - name: CHECKPOINT_TOKEN + valueFrom: + secretKeyRef: + name: unitdb + key: CHECKPOINT_TOKEN + resources: + requests: + cpu: 10m + memory: 16Mi + limits: + memory: 64Mi diff --git a/deploy/kubernetes/restore-node-job.yaml b/deploy/kubernetes/restore-node-job.yaml new file mode 100644 index 0000000..7257c03 --- /dev/null +++ b/deploy/kubernetes/restore-node-job.yaml @@ -0,0 +1,60 @@ +# One node's part of a whole-cluster restore (docs/backup-restore.md, +# runbook B): on the node's new, empty claim, its checkpoint of the run, +# decrypted, where the node's store is (/var/udb/data/unitdb), and the +# security journal since the run (/var/udb/data/restore-journal). Then the +# node starts with -restored and -journal (runbook B, step 6). +# +# Per node: replace NODE (unitdb-0, unitdb-1, unitdb-2) and RUN (the run's +# id, from `backup` runs listed in the bucket), and apply: +# sed -e s/NODE/unitdb-0/g -e s/RUN/20261004T213000Z/g restore-node-job.yaml | kubectl apply -f - +# +# It runs as the user unitdb runs as: the files it writes are the node's +# store. If unitdb runs as another user than root, give the Job the same +# securityContext. +# +# It reads the bucket and the backup key, so it needs the reader's identity +# in the unitdb namespace for the restore only: the unitdb-backup-reader +# Secret, made at runbook B's step 3 and deleted at its end. +apiVersion: batch/v1 +kind: Job +metadata: + name: unitdb-restore-NODE + namespace: unitdb +spec: + backoffLimit: 0 + template: + spec: + restartPolicy: Never + enableServiceLinks: false + initContainers: + - name: checkpoint + image: unitdb:latest + command: ["/backup"] + args: ["fetch", "-run=RUN", "-node=NODE", "-out=/var/udb/data/unitdb"] + envFrom: + - secretRef: + name: unitdb-backup-reader + env: &restoreEnv + - name: BACKUP_S3_BUCKET + value: unitdb-backups-production + - name: BACKUP_S3_REGION + value: + - name: BACKUP_CLUSTER + value: production + volumeMounts: &restoreMounts + - name: data + mountPath: /var/udb/data + containers: + - name: journal + image: unitdb:latest + command: ["/backup"] + args: ["fetch-journal", "-since=RUN", "-out=/var/udb/data/restore-journal"] + envFrom: + - secretRef: + name: unitdb-backup-reader + env: *restoreEnv + volumeMounts: *restoreMounts + volumes: + - name: data + persistentVolumeClaim: + claimName: data-NODE diff --git a/deploy/kubernetes/restore-test.yaml b/deploy/kubernetes/restore-test.yaml new file mode 100644 index 0000000..c122b96 --- /dev/null +++ b/deploy/kubernetes/restore-test.yaml @@ -0,0 +1,80 @@ +# unitdb's weekly restore test (docs/backup-restore.md): the backup +# command's verify downloads the newest complete run from the bucket, +# decrypts it, and opens each node's checkpoint on a scratch server in this +# pod (no cluster, no clients), with the escrowed keyring: it must open and +# be ready, hold what the manifest says, and hold the run's canary. It +# pushes unitdb_backup_restore_verified_timestamp_seconds to the Pushgateway +# only if every node passes; backup-alerts.yaml fires when that is more than +# 8 days old. +# +# Its own namespace and identity: the reader's (deploy/aws/README.md, step +# 5), the only one besides named admins that may read the bucket, the +# backup key and the escrowed keyring. Nothing in the unitdb namespace can. +# +# Before applying: +# kubectl create namespace unitdb-restore-test +# kubectl -n unitdb-restore-test create secret generic unitdb-backup-reader \ +# --from-literal=AWS_ACCESS_KEY_ID= \ +# --from-literal=AWS_SECRET_ACCESS_KEY= +apiVersion: batch/v1 +kind: CronJob +metadata: + name: unitdb-restore-test + namespace: unitdb-restore-test +spec: + schedule: "30 23 * * 6" # Saturdays 23:30 UTC, after that day's run + concurrencyPolicy: Forbid + successfulJobsHistoryLimit: 3 + failedJobsHistoryLimit: 3 + jobTemplate: + spec: + backoffLimit: 0 + activeDeadlineSeconds: 21600 + template: + spec: + restartPolicy: Never + enableServiceLinks: false + containers: + - name: verify + image: unitdb:latest + command: ["/backup"] + args: + - verify + - -run=latest + - -server=/unitdb + - -work=/scratch + - -pushgateway=http://pushgateway.monitoring.svc.cluster.local:9091 + env: + - name: BACKUP_S3_BUCKET + value: unitdb-backups-production + - name: BACKUP_S3_REGION + value: + - name: BACKUP_CLUSTER + value: production + - name: AWS_ACCESS_KEY_ID + valueFrom: + secretKeyRef: + name: unitdb-backup-reader + key: AWS_ACCESS_KEY_ID + - name: AWS_SECRET_ACCESS_KEY + valueFrom: + secretKeyRef: + name: unitdb-backup-reader + key: AWS_SECRET_ACCESS_KEY + # One scratch server at a time; its memory store is the + # nodes' (mem_size, 500 MB). + resources: + requests: + cpu: 500m + memory: 1Gi + limits: + memory: 2Gi + volumeMounts: + - name: scratch + mountPath: /scratch + volumes: + # A run, decrypted: every node's checkpoint. Size it to three + # nodes' stores. + - name: scratch + emptyDir: + sizeLimit: 150Gi diff --git a/docs/README.md b/docs/README.md index 7d04fa4..bf8c035 100644 --- a/docs/README.md +++ b/docs/README.md @@ -8,5 +8,8 @@ changes. - [Message log replication](message-log-replication.md): replicas of stored messages and session logs, hints, and rebuilding a node that lost its disk. +- [Backup and restore](backup-restore.md): checkpoints, the backup run, copies + off the cluster, the security journal, reconciliation after a restore, the + weekly restore test, and the runbooks. - [Rolling deploys](rolling-deploys.md): upgrading a cluster node by node, and the maintenance-window upgrade from v0.3.0. diff --git a/docs/backup-restore.md b/docs/backup-restore.md new file mode 100644 index 0000000..76442ea --- /dev/null +++ b/docs/backup-restore.md @@ -0,0 +1,198 @@ +# Backup and restore + +How a unitdb cluster's data is backed up and restored: checkpoints, the +backup run and its manifest, copies off the cluster, the security journal, +reconciliation after a restore, the weekly restore test, and the runbooks. +Everything here is built; the deploy examples are in `deploy/kubernetes` and +`deploy/aws`. + +## What a backup is + +- **A checkpoint** is a copy of a node's store that opens as the store was + at one moment (`server/internal/db/unitdb/checkpoint.go`). A copy of a + running store's files isn't one: they are copied at different moments, + and the WAL isn't fsynced, so a volume snapshot is like a power cut. The + adapter holds writes back, for up to about 1.5 s, until the DB has written + out the latest; then it syncs, copies the DB's files, and copies memdb's + records (not its files) into a fresh memdb. Reads go on. Last it writes + `checkpoint.json`: the node, the backup run, the time, the ring and + engine versions, the keyring's key ids (never keys), and the copy's + counts. A checkpoint without it is incomplete. +- `POST /_checkpoint` on the monitor port (`monitor_listen`), with + `Authorization: Bearer `, takes one into `CHECKPOINT_DIR`; + off unless both are set. One runs at a time, one per + `CHECKPOINT_MIN_MINUTES` (10); the newest `CHECKPOINT_KEEP` (3) are kept. +- **A backup run** (`server/cmd/backup`, `deploy/kubernetes/backups.yaml`) + checkpoints every node under one run id (`POST /_checkpoint?run=`, + into `ckpt--`), one node after another, each tried three + times, and keeps the run's `manifest.json` (each node's + `checkpoint.json`, or why it failed) in every node's checkpoint + (`PUT /_checkpoint/manifest`). A node that failed fails the run: a + restore needs every node of one run. +- Just before its checkpoint of a run, a node writes **a canary**, the run's + id, into its store as a message and as a memdb record; the restore test + reads it back. + +## Copies off the cluster + +With `-upload`, the backup run has each node upload its own checkpoint (the +checkpoints are on each node's volume): `POST /_checkpoint/upload?run=` +streams it, archived (tar), compressed (zstd) and encrypted with +[age](https://age-encryption.org) to the backup key's public half, to S3, +with the manifest (`server/internal/backup`): + +``` +//manifest.json +//.tar.zst.age +/journal//-