Skip to content

Latest commit

 

History

History
15 lines (8 loc) · 3.3 KB

File metadata and controls

15 lines (8 loc) · 3.3 KB

Replica snapshot installation and recovery

KeyLoad keeps the provider's native Raft log and snapshot lifecycle. Snapshot payloads use the existing verified KeyLoad checkpoint format; published native files are named by their included index and term. Every image must pass frame checksums, the complete-image checksum/count, incarnation and applied-position validation before it can replace canonical storage.

An incoming installation first flushes a bounded, versioned incoming intent to disk and atomically publishes it in the node's private snapshot directory. It records the intended incarnation, index and term. This happens before the provider can create or publish its snapshot file. The .NEXT 6.8.1 receiver can rename a partial file even when writing throws; KeyLoad therefore guards that lifecycle rather than treating the existence of a native snapshot filename as successful acceptance.

If transfer or validation fails before the image is verified, KeyLoad removes that attempt's published file and clears its intent. The previous native snapshot, committed log tail and canonical materialization remain available. A complete verified image survives a failure after canonical installation begins, so reopening can finish the same installation. The intent is cleared after a successful installation. The native append acknowledgement still waits for DurableRaftLog's explicit foreground checkpoint.

Startup recovers the intent before restoring the native state machine or constructing its WAL. An incomplete, corrupt or wrong-scope attempt is discarded only while canonical apply remains below its intended cut. A complete verified image is retained and restored. If canonical apply has reached that cut but the corresponding image is missing or invalid, recovery fails with Corruption and preserves the files. Corrupt intent metadata also fails closed. Established snapshot corruption without a pending intent is never treated as a torn upload. Unpublished .tmp files in this exclusively owned directory are reclaimed before the raft host starts.

The resulting recovery cut is either the prior committed tail or the complete snapshot's verified committed cut. A completed image is not installed a second time on reopening, preserving its read-generation fence. An interrupted response still says nothing about whether a complete image was accepted; the leader may retry through the normal Raft protocol.

Qualification uses the real persistent Raft interface for interrupted/cancelled writes, invalid checksums and foreign incarnations, and proves reopening followed by a valid retry. Five process-kill cases stop at intent publication, during body writing, after a rejected file has been published, after image verification and after intent removal. They verify the recovered canonical/native positions, dispatch state, read generation and temporary-file cleanup. Additional regressions preserve corrupted materialized images and invalid intents for diagnosis. The RF3 suite separately exercises signed HTTP snapshot catch-up on an empty follower.

These are process recovery guarantees. Directory-entry persistence under power loss, disk/controller failures, large snapshots under transport deadlines, broader network fault schedules and endurance remain qualification gates. This contract does not advertise power-loss durability or production readiness.