diff --git a/CHANGELOG.md b/CHANGELOG.md index 2070512ff..a59524295 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,6 +7,15 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +- **R-15 Gate 4D:** the conversion session renews its lease (part b-1 of the owner's steps through the + gate). `ConversionSession::renew_lease` builds the renewal from the session's own snapshot (`renewed_lease`), + takes the root event through the held admission, commits it with the key route and active epoch the + committed root names, reads the committed journal back under that root event and returns the commit; it is + the first caller of the D3b renewal. The journal directory must lie below the admitted installation and is + pinned for the session, and every file operation goes through canonical paths; the root settles what a + step did, a failed step keeps the snapshot, a retry adopts what was written, a differing candidate of a + crashed attempt is moved aside, and an unsettled step spends the session. Entering + `CONVERT` and moving the cursor are the next part. PR #1014. - **R-15 Gate 4D:** the entry gate of the exclusive conversion driver (part a). `begin_conversion` takes the installation's exclusive admission by mutable borrow, derives the root layout from the scope that admission was acquired for, re-reads the committed binding, the registry-routed journal key and the root-named diff --git a/crates/worldscript-secure-storage/src/authority.rs b/crates/worldscript-secure-storage/src/authority.rs index c411defe8..75481f8e2 100644 --- a/crates/worldscript-secure-storage/src/authority.rs +++ b/crates/worldscript-secure-storage/src/authority.rs @@ -555,7 +555,7 @@ pub fn begin_streamed_capture( layout: RootLayout<'_>, begin: SessionBegin<'_>, ) -> Result { - let live = committed_live_migration(fs, provider, layout)?; + let live = committed_route(fs, provider, layout)?.live; let key = route_journal_key(fs, provider, layout, begin.journal)?; let capture = { let mut ctx = journal_context(&mut *fs, begin.journal, &key, CandidateConflict::Refuse); @@ -577,26 +577,42 @@ pub fn begin_streamed_capture( }) } -/// The binding the committed root names, or [`AuthorityError::NoLiveMigration`]. +/// What the committed root names for a bound migration: the binding, and the key route and active +/// epoch the journal-owner operations must carry unchanged. +struct CommittedRoute { + live: LiveMigration, + root_key_ref: RootKeyRefV1, + active_key_epoch: u64, +} + +/// The route the committed root names, or [`AuthorityError::NoLiveMigration`]. /// /// The root alone names the binding: the catalog pages are not read, so starting a capture of a very /// large inventory does not first materialise a very large catalog. -fn committed_live_migration( +fn committed_route( fs: &mut F, provider: &P, layout: RootLayout<'_>, -) -> Result { - load_committed_root(fs, provider, layout) - .map_err(AuthorityError::Root)? - .and_then(|view| view.root.live_migration) - .ok_or(AuthorityError::NoLiveMigration) +) -> Result { + let view = load_committed_root(fs, provider, layout).map_err(AuthorityError::Root)?; + view.and_then(|view| { + Some(CommittedRoute { + live: view.root.live_migration?, + root_key_ref: view.root_key_ref, + active_key_epoch: view.root.active_key_epoch, + }) + }) + .ok_or(AuthorityError::NoLiveMigration) } -/// The journal state the committed root vouches for: its binding and the manifest it names. +/// The journal state the committed root vouches for: its binding, the manifest it names, and the key +/// route and active epoch the root itself carries. #[derive(Debug, Clone, PartialEq, Eq)] pub(crate) struct CommittedJournal { pub(crate) live: LiveMigration, pub(crate) manifest: JournalManifest, + pub(crate) root_key_ref: RootKeyRefV1, + pub(crate) active_key_epoch: u64, } /// Reads the committed journal state without a key in the caller's hands. @@ -612,13 +628,18 @@ pub(crate) fn read_committed_journal( layout: RootLayout<'_>, journal: JournalSource<'_>, ) -> Result { - let live = committed_live_migration(fs, provider, layout)?; + let route = committed_route(fs, provider, layout)?; let key = route_journal_key(fs, provider, layout, journal)?; let manifest = { let mut ctx = journal_context(&mut *fs, journal, &key, CandidateConflict::Refuse); - load_authoritative_manifest(&mut ctx, &live).map_err(AuthorityError::Journal)? + load_authoritative_manifest(&mut ctx, &route.live).map_err(AuthorityError::Journal)? }; - Ok(CommittedJournal { live, manifest }) + Ok(CommittedJournal { + live: route.live, + manifest, + root_key_ref: route.root_key_ref, + active_key_epoch: route.active_key_epoch, + }) } impl StagingSession { @@ -822,6 +843,21 @@ pub fn commit_lease_renewal( check_operation_id(&renewal.manifest.operation_id) .map_err(|_| AuthorityError::InvalidOperationId)?; let held = acquire_root_commit(layout)?; + commit_lease_renewal_held(fs, provider, layout, renewal, &held) +} + +/// As [`commit_lease_renewal`], under a root event the caller already holds: the exclusive conversion +/// session takes it through its admission, so the identity check and the root lock are coupled. `held` +/// must guard the root of `layout`. +pub(crate) fn commit_lease_renewal_held( + fs: &mut F, + provider: &mut P, + layout: RootLayout<'_>, + renewal: JournalCheckpoint<'_>, + held: &RootCommitGuard, +) -> Result { + check_operation_id(&renewal.manifest.operation_id) + .map_err(|_| AuthorityError::InvalidOperationId)?; let commit = journal_catalog_commit(&renewal); let committed = committed_binding(fs, provider, layout, commit)?; let key = route_journal_key(fs, provider, layout, renewal.journal)?; @@ -836,7 +872,7 @@ pub fn commit_lease_renewal( binding: BindingStep::Renewal(&advance), key: &key, }; - commit_planned(fs, provider, layout, commit, &held, Some(step)) + commit_planned(fs, provider, layout, commit, held, Some(step)) } fn acquire_root_commit(layout: RootLayout<'_>) -> Result { diff --git a/crates/worldscript-secure-storage/src/conversion.rs b/crates/worldscript-secure-storage/src/conversion.rs index 86fea7101..6b73ba71f 100644 --- a/crates/worldscript-secure-storage/src/conversion.rs +++ b/crates/worldscript-secure-storage/src/conversion.rs @@ -3,9 +3,9 @@ //! Conversion rewrites every record of the final inventory, so it must run with no ordinary operation //! in flight. [`begin_conversion`] makes that structural: it takes the installation's //! [`ExclusiveAdmissionGuard`] by mutable borrow, derives the root layout from the scope that guard -//! was acquired for (the root cannot differ from the one admitted), and returns a -//! [`ConversionSession`] that keeps the borrow, so one admission carries one session and the -//! admission outlives it. A shared guard cannot be passed: +//! was acquired for (the root cannot differ from the one admitted), pins the journal directory below +//! the admitted installation, and returns a [`ConversionSession`] that keeps the borrow, so one +//! admission carries one session and the admission outlives it. A shared guard cannot be passed: //! //! ```compile_fail //! use worldscript_secure_storage::{ @@ -30,18 +30,48 @@ //! the existing takeover followed by a fresh begin. No clock is read, so a lease past its expiry is //! still the caller's while nobody has taken over (the fence arbitrates). //! -//! This slice writes nothing and holds no key; the steps that move the journal are the next slice. +//! Beginning writes nothing and holds no key. The session then moves the journal only through the +//! fenced journal-owner operations, one step at a time ([`ConversionSession::renew_lease`] is the +//! first). A step checks the admission and the journal directory, builds the successor from the +//! session's own snapshot, takes the root event through the held admission (so the identity check and +//! the root lock are coupled), commits with the key route and active epoch the committed root itself +//! names, reads the committed journal back under that same root event and checks the admission again. +//! +//! What a step leaves behind. All file operations of a session go through the canonical paths of the +//! root, the installation and the journal directory that were validated and pinned at begin, never +//! through the caller's spelling (a symlink retargeted later cannot redirect them). Before the commit +//! the tree and the snapshot are unchanged, and a retry rebuilds the same successor, which the journal +//! adopts if it is identical to a candidate already written; a differing candidate of a crashed attempt +//! is moved aside with its bytes preserved, so it never blocks the next step. Whatever the commit +//! reports, the journal is then read back under the same root event and the root settles the outcome: +//! it names the successor (committed, even if the commit reported an error), or it does not (not +//! committed, snapshot kept). If the read-back fails, or names a journal that is neither of the two +//! where the commit reported success, the session is spent: every later step is refused and the caller +//! begins again to learn the state. A lost admission is reported even when it was lost after the +//! commit; the snapshot then is the committed journal. + +use std::fs::File; +use std::path::{Path, PathBuf}; -use crate::admission::{AdmissionScope, ExclusiveAdmissionGuard}; -use crate::authority::{read_committed_journal, AuthorityError, CommittedJournal, JournalSource}; -use crate::durable::DurableFs; -use crate::journal::{phase_code, JournalManifest, MigrationFence}; +use crate::admission::{AdmissionError, AdmissionScope, ExclusiveAdmissionGuard}; +use crate::authority::{ + commit_lease_renewal_held, read_committed_journal, AuthorityError, CommittedJournal, + JournalCheckpoint, JournalSource, +}; +use crate::durable::{DurableFs, WriteOperationId}; +use crate::error::SealError; +use crate::journal::{ + phase_code, renewed_lease, CandidateConflict, JournalManifest, MigrationExecutionError, + MigrationFence, +}; use crate::provider::KeyProvider; use crate::root::LiveMigration; -use crate::root_store::RootLayout; +use crate::root_lock::sys; +use crate::root_store::{RootCommitted, RootLayout}; /// What the gate needs from its caller: the scope the admission was acquired for, where the journal -/// lives, and the identity the committed lease must carry. +/// lives (below the installation directory of that scope), and the identity the committed lease must +/// carry. #[derive(Clone, Copy)] pub struct ConversionBegin<'a> { pub scope: AdmissionScope<'a>, @@ -49,12 +79,22 @@ pub struct ConversionBegin<'a> { pub owner_id: &'a str, } -/// Why the gate refused. Every refusal is returned before anything is written. +/// Why the gate or a step refused. A refusal by the gate, and by a step before its commit, leaves the +/// tree unchanged. #[derive(Debug, Clone, PartialEq, Eq)] pub enum ConversionError { - /// The guard does not hold the scope: other directories, or directories replaced since acquisition. + /// The guard does not hold the scope, or the journal directory is no longer the one pinned: + /// other directories, or directories replaced since acquisition. NotAdmitted, - /// Reading the committed journal failed (no bound migration, key route, authentication, I/O). + /// The admission could not take the root event, or could not pin a directory. + Admission(AdmissionError), + /// Another root commit holds the root lock; nothing was written, try again. + RootBusy, + /// The journal directory is not below the admitted installation directory, or lies within the + /// authority root (a journal generation could then collide with a root slot). + JournalMisplaced, + /// Reading the committed journal, or committing a step, failed (no bound migration, key route, + /// authentication, I/O, a stale token). Authority(AuthorityError), /// The operation finished (`DONE`). TerminalPhase, @@ -66,6 +106,25 @@ pub enum ConversionError { FinalInventoryNotCaptured, /// The committed lease is owned by another owner, or by none. ForeignLeaseOwner, + /// A step's successor is not a valid one (a renewal with no lease to renew, or an expiry that does + /// not move strictly forward). + Migration(MigrationExecutionError), + /// No write-operation identifier could be drawn from the operating system. + OperationId(SealError), + /// The step committed, but the journal read back is not the one it committed: the session is + /// spent, begin again. + Superseded, + /// The step committed, but the committed journal could not be read back: the session is spent, + /// begin again. + Unreadable(AuthorityError), + /// The commit reported this error, yet the root names the successor: the step is committed (an + /// ambiguous durability outcome of the anchor, say) and the session follows the root. + Committed(AuthorityError), + /// The commit reported this error and the root could not be read back to settle whether it + /// committed: the session is spent, begin again. + Unsettled(AuthorityError), + /// An earlier step spent this session; begin again. + Spent, } impl From for ConversionError { @@ -74,13 +133,52 @@ impl From for ConversionError { } } +fn io_error(error: std::io::Error) -> ConversionError { + ConversionError::Admission(AdmissionError::from(error)) +} + +/// A directory held open, so that a later look can tell whether its path still names it. +#[derive(Debug)] +struct PinnedDirectory { + pin: File, + canonical: PathBuf, +} + +impl PinnedDirectory { + /// Pins `dir` (before canonicalising it, as the admission does) and requires it to lie strictly + /// below the canonical installation directory and outside the canonical authority root. + fn placed(installation: &Path, root: &Path, dir: &Path) -> Result { + let pin = sys::open_directory(dir).map_err(io_error)?; + let canonical = std::fs::canonicalize(dir).map_err(io_error)?; + let below = canonical != installation && canonical.starts_with(installation); + if !below || canonical.starts_with(root) { + return Err(ConversionError::JournalMisplaced); + } + let pinned = Self { pin, canonical }; + if pinned.is_current() { + Ok(pinned) + } else { + Err(ConversionError::NotAdmitted) + } + } + + fn is_current(&self) -> bool { + sys::still_named(&self.pin, &self.canonical).unwrap_or(false) + } +} + /// The exclusive conversion entry: the committed journal state, read under an admission that no /// ordinary operation can share. #[derive(Debug)] pub struct ConversionSession<'a> { held: &'a mut ExclusiveAdmissionGuard, - scope: AdmissionScope<'a>, + /// The canonical installation and root directories the admission was checked against; every file + /// operation of the session goes through these, not through the caller's spelling. + installation: PathBuf, + root: PathBuf, + journal_pin: PinnedDirectory, journal: CommittedJournal, + spent: bool, } /// Opens the conversion gate over the committed journal; see the module documentation. @@ -88,23 +186,41 @@ pub fn begin_conversion<'a, F: DurableFs, P: KeyProvider>( held: &'a mut ExclusiveAdmissionGuard, fs: &mut F, provider: &P, - begin: ConversionBegin<'a>, + begin: ConversionBegin<'_>, ) -> Result, ConversionError> { - ensure_admitted(held, begin.scope)?; - let layout = RootLayout { - root_dir: begin.scope.root_dir, + let installation = std::fs::canonicalize(begin.scope.installation_dir).map_err(io_error)?; + let root = std::fs::canonicalize(begin.scope.root_dir).map_err(io_error)?; + ensure_admitted(held, canonical(&installation, &root))?; + let journal_pin = PinnedDirectory::placed(&installation, &root, begin.journal.dir)?; + let layout = RootLayout { root_dir: &root }; + let source = JournalSource { + dir: &journal_pin.canonical, + operation: begin.journal.operation, }; - let journal = read_committed_journal(fs, provider, layout, begin.journal)?; + let journal = read_committed_journal(fs, provider, layout, source)?; // The directories must still be the ones admitted after the reads, not only before them. - ensure_admitted(held, begin.scope)?; + ensure_admitted(held, canonical(&installation, &root))?; + if !journal_pin.is_current() { + return Err(ConversionError::NotAdmitted); + } check_convertible(&journal.manifest, begin.owner_id)?; Ok(ConversionSession { held, - scope: begin.scope, + installation, + root, + journal_pin, journal, + spent: false, }) } +fn canonical<'p>(installation: &'p Path, root: &'p Path) -> AdmissionScope<'p> { + AdmissionScope { + installation_dir: installation, + root_dir: root, + } +} + fn ensure_admitted( held: &ExclusiveAdmissionGuard, scope: AdmissionScope<'_>, @@ -132,8 +248,8 @@ fn check_convertible(manifest: &JournalManifest, owner_id: &str) -> Result<(), C Ok(()) } -impl ConversionSession<'_> { - /// The manifest the committed root names. +impl<'a> ConversionSession<'a> { + /// The manifest the committed root names, as of the last step this session could read back. pub fn manifest(&self) -> &JournalManifest { &self.journal.manifest } @@ -148,8 +264,108 @@ impl ConversionSession<'_> { MigrationFence::from_manifest(&self.journal.manifest) } - /// Whether the admission still holds the scope the session began under. + /// Whether the admission still holds the scope the session began under and the journal directory + /// is still the one pinned. pub fn is_admitted(&self) -> bool { - self.held.guards(self.scope) + self.check_admitted().is_ok() + } + + /// Renews the owner's lease to `expires_unix_ms`, a time the caller chose (Core reads no clock), + /// and returns the root commit, whose `directories` says whether the directory entries are + /// confirmed durable. + /// + /// The expiry must move strictly forward (`Migration(InvalidLeaseRenewal)`); a lease long past its + /// expiry is still renewed while nobody has taken over, because a takeover advances the fence and + /// the committed manifest then names another owner. The renewal is committed by + /// [`commit_lease_renewal_held`] under the root event the admission hands out, so a snapshot that + /// went stale is refused before any write. See the module documentation for what a failed step + /// leaves behind. + pub fn renew_lease( + &mut self, + fs: &mut F, + provider: &mut P, + expires_unix_ms: u64, + ) -> Result { + if self.spent { + return Err(ConversionError::Spent); + } + self.check_admitted()?; + let renewed = renewed_lease(&self.journal.manifest, &self.fence(), expires_unix_ms) + .map_err(ConversionError::Migration)?; + let operation = WriteOperationId::generate().map_err(ConversionError::OperationId)?; + let layout = RootLayout { + root_dir: &self.root, + }; + let journal = JournalSource { + dir: &self.journal_pin.canonical, + operation: &operation, + }; + let fence = MigrationFence::from_manifest(&renewed); + let step = JournalCheckpoint { + manifest: &renewed, + fence: &fence, + journal, + root_key_ref: &self.journal.root_key_ref, + active_key_epoch: self.journal.active_key_epoch, + conflict: CandidateConflict::Quarantine, + }; + let (committed, read_back) = { + let event = self + .held + .try_root_commit() + .map_err(ConversionError::Admission)? + .ok_or(ConversionError::RootBusy)?; + let root = event.root_guard().map_err(ConversionError::Admission)?; + let committed = commit_lease_renewal_held(fs, provider, layout, step, root); + // Whatever the commit reported, read the root back while the event is still held: no other + // root commit comes between, and the root says whether the step committed. + ( + committed, + read_committed_journal(fs, &*provider, layout, journal), + ) + }; + let settled = self.settle(&renewed, committed, read_back); + // A step that landed, whatever the commit reported, is followed by the admission check. + if matches!(settled, Ok(_) | Err(ConversionError::Committed(_))) { + self.check_admitted()?; + } + settled + } + + /// Settles what a step did from what the commit reported and what the root names afterwards. + fn settle( + &mut self, + step: &JournalManifest, + committed: Result, + read_back: Result, + ) -> Result { + match (committed, read_back) { + (Ok(commit), Ok(journal)) if journal.manifest == *step => { + self.journal = journal; + Ok(commit) + } + (Ok(_), Ok(_)) => self.spend(ConversionError::Superseded), + (Ok(_), Err(error)) => self.spend(ConversionError::Unreadable(error)), + (Err(error), Ok(journal)) if journal.manifest == *step => { + self.journal = journal; + Err(ConversionError::Committed(error)) + } + (Err(error), Ok(_)) => Err(ConversionError::Authority(error)), + (Err(error), Err(_)) => self.spend(ConversionError::Unsettled(error)), + } + } + + fn spend(&mut self, error: ConversionError) -> Result { + self.spent = true; + Err(error) + } + + fn check_admitted(&self) -> Result<(), ConversionError> { + ensure_admitted(self.held, canonical(&self.installation, &self.root))?; + if self.journal_pin.is_current() { + Ok(()) + } else { + Err(ConversionError::NotAdmitted) + } } } diff --git a/crates/worldscript-secure-storage/src/journal/mod.rs b/crates/worldscript-secure-storage/src/journal/mod.rs index 1a9e38cae..876a644d5 100644 --- a/crates/worldscript-secure-storage/src/journal/mod.rs +++ b/crates/worldscript-secure-storage/src/journal/mod.rs @@ -144,9 +144,9 @@ pub use state::{ allows_phase_transition, assert_binding_successor, assert_fence, assert_live_binding, assert_manifest_promote_authority, assert_page_promote_authority, authoritative_manifest_revision, checkpoint_progress, is_terminal_phase, mark_done, - mark_recovery, ordinary_mutating_writes_admitted, transition_phase, JournalCheckpointCursor, - JournalInventoryExtent, JournalRevision, ManifestEnvelopeDigest, MigrationExecutionError, - MigrationFence, MigrationPhase, RecoveryReasonCode, + mark_recovery, ordinary_mutating_writes_admitted, renewed_lease, transition_phase, + JournalCheckpointCursor, JournalInventoryExtent, JournalRevision, ManifestEnvelopeDigest, + MigrationExecutionError, MigrationFence, MigrationPhase, RecoveryReasonCode, }; pub use stream_capture::{ load_staged_page, CaptureStart, StagedCapture, StagedPage, StreamedCapture, diff --git a/crates/worldscript-secure-storage/src/journal/state.rs b/crates/worldscript-secure-storage/src/journal/state.rs index 78f182e09..03cf5c61e 100644 --- a/crates/worldscript-secure-storage/src/journal/state.rs +++ b/crates/worldscript-secure-storage/src/journal/state.rs @@ -5,6 +5,7 @@ use crate::root::LiveMigration; use super::manifest::JournalManifest; +use super::renewal::assert_renewal_successor; use super::wire::final_inventory_required; use super::{phase_code, JournalError}; @@ -455,6 +456,25 @@ pub fn checkpoint_progress( Ok(next) } +/// The owner's renewal of its lease: the next revision with the expiry moved to `expires_unix_ms`. +/// +/// Only the revision and the expiry change, and the result is proved a valid renewal of `manifest` +/// ([`assert_renewal_successor`]) before it is returned: a named lease to renew, an expiry strictly +/// after the held one, a journal that is not terminal. The expiry is the caller's; Core reads no clock. +pub fn renewed_lease( + manifest: &JournalManifest, + fence: &MigrationFence, + expires_unix_ms: u64, +) -> Result { + assert_fence(manifest, fence)?; + let mut next = manifest.clone(); + next.journal_revision = bump_revision(manifest)?.wire(); + next.lease_expires_unix_ms = Some(expires_unix_ms); + assert_renewal_successor(manifest, &next)?; + next.encode()?; + Ok(next) +} + /// Moves the operation into durable terminal refusal with an explicit reason code. pub fn mark_recovery( manifest: &JournalManifest, diff --git a/crates/worldscript-secure-storage/tests/gate4d_conversion_entry_test.rs b/crates/worldscript-secure-storage/tests/gate4d_conversion_entry_test.rs index 100b65cd2..d0717b535 100644 --- a/crates/worldscript-secure-storage/tests/gate4d_conversion_entry_test.rs +++ b/crates/worldscript-secure-storage/tests/gate4d_conversion_entry_test.rs @@ -1,7 +1,8 @@ //! Gate 4D C2a: the entry gate of the exclusive conversion driver (§10.3 `ADMIT`, `CONVERT`). The gate //! needs the installation's exclusive admission, re-reads the journal the committed root vouches for //! and refuses anything but the owner's conversion; it writes nothing. Every scenario therefore -//! commits a real root that binds a real journal. +//! commits a real root that binds a real journal. C2b-1: the session renews its lease through the +//! fenced journal-owner operation and follows the root. use std::collections::BTreeMap; use std::ffi::OsString; @@ -11,17 +12,20 @@ use std::io; use std::path::{Path, PathBuf}; use std::sync::atomic::{AtomicU32, Ordering}; -use worldscript_secure_storage::memory_provider::MemoryKeyProvider; +use worldscript_secure_storage::journal::renewed_lease; +use worldscript_secure_storage::memory_provider::{AnchorOp, Fault, MemoryKeyProvider}; use worldscript_secure_storage::{ - begin_conversion, commit_catalog_change, commit_root, content_digest, empty_inventory_digest, - empty_journal_page_set_digest, generation_path, load_catalog, operation_type, phase_code, - write_key_epoch, AdmissionScope, AuthorityError, CatalogChange, CatalogCommit, ConversionBegin, - ConversionError, DirectoryDurability, DurableFs, ExclusiveAdmissionGuard, InstallationScopeId, - JournalDurableError, JournalManifest, JournalRouteError, JournalSource, Key, KeyEpochCommit, - KeyEpochRecord, KeyEpochStatus, KeyProvider, LiveMigration, MigrationExecutionError, - MigrationFence, RecordClass, RecordIdentity, RecordMeta, RootBody, RootCommitEvidence, - RootCommitGuard, RootCommitRequest, RootCommitState, RootKeyRefV1, RootLayout, StdFs, - WriteOperationId, JOURNAL_MANIFEST_RECORD_SCHEMA, OPERATION_ADMISSION_LOCK_FILE, + begin_conversion, commit_catalog_change, commit_journal_takeover, commit_root, content_digest, + empty_inventory_digest, empty_journal_page_set_digest, generation_path, load_catalog, + operation_type, phase_code, write_key_epoch, AdmissionScope, AuthorityError, CandidateConflict, + CatalogChange, CatalogCommit, ConversionBegin, ConversionError, ConversionSession, + DirectoryDurability, DurableFs, ExclusiveAdmissionGuard, InstallationScopeId, + JournalCheckpoint, JournalDurableError, JournalManifest, JournalRouteError, JournalSource, + JournalTakeoverCommit, Key, KeyEpochCommit, KeyEpochRecord, KeyEpochStatus, KeyProvider, + LiveMigration, MigrationExecutionError, MigrationFence, RecordClass, RecordIdentity, + RecordMeta, RootBody, RootCommitEvidence, RootCommitGuard, RootCommitRequest, RootCommitState, + RootKeyRefV1, RootLayout, StdFs, WriteOperationId, JOURNAL_MANIFEST_RECORD_SCHEMA, + OPERATION_ADMISSION_LOCK_FILE, }; const OPERATION: &str = "conversion-op"; @@ -30,6 +34,8 @@ const REVISION: u64 = 3; const FENCE: u64 = 4; /// A lease long past its expiry: the gate reads no clock, so it is still the owner's. const LEASE_EXPIRY: u64 = 1_000; +/// The expiry a renewal moves the lease to. +const RENEWED_EXPIRY: u64 = 5_000; const MATERIAL: [u8; 32] = [9; 32]; const EPOCH: u64 = 1; @@ -239,11 +245,61 @@ impl Fixture { } } - /// The gate as `owner` under this installation's own admission, over the real file system. - fn begin(&self, owner: &str) -> Result<(), ConversionError> { + /// The manifest the gate returns to `owner` under this installation's own admission, which is + /// released again: no other admission may be held. + fn committed(&self, owner: &str) -> Result { let paths = self.paths(); let mut held = paths.admission(); - begin_conversion(&mut held, &mut StdFs, &self.provider, paths.begin(owner)).map(|_| ()) + begin_conversion(&mut held, &mut StdFs, &self.provider, paths.begin(owner)) + .map(|session| session.manifest().clone()) + } + + /// The gate as `owner` under this installation's own admission, over the real file system. + fn begin(&self, owner: &str) -> Result<(), ConversionError> { + self.committed(owner).map(|_| ()) + } + + /// Another owner takes the journal over from `committed`, whose lease has lapsed: the next + /// revision at the next fence, with a lease of its own. + fn take_over(&mut self, committed: &JournalManifest) { + let mut claim = committed.clone(); + claim.journal_revision += 1; + claim.fencing_generation += 1; + claim.lease_owner_id = Some("new-owner".into()); + claim.lease_expires_unix_ms = Some(LEASE_EXPIRY + 10_000); + let (journal_dir, operation) = (self.journal_dir(), WriteOperationId::generate().unwrap()); + let fence = MigrationFence::from_manifest(&claim); + let takeover = JournalTakeoverCommit { + claim: JournalCheckpoint { + manifest: &claim, + fence: &fence, + journal: JournalSource { + dir: &journal_dir, + operation: &operation, + }, + root_key_ref: &self.key_ref, + active_key_epoch: EPOCH, + conflict: CandidateConflict::Refuse, + }, + now_unix_ms: LEASE_EXPIRY + 1, + }; + let root_dir = self.root_dir(); + let layout = RootLayout { + root_dir: &root_dir, + }; + commit_journal_takeover(&mut StdFs, &mut self.provider, layout, takeover).unwrap(); + } + + /// The generation of the committed root. + fn root_generation(&self) -> u64 { + let root_dir = self.root_dir(); + let layout = RootLayout { + root_dir: &root_dir, + }; + let catalog = load_catalog(&mut StdFs, &self.provider, layout) + .unwrap() + .unwrap(); + catalog.root.root_generation } /// Every file under the fixture with a digest of its content. @@ -289,12 +345,13 @@ impl Paths { } } -/// The real file system, reporting to `on_read` after every read has happened. -struct WatchedFs { +/// The real file system, reporting to `on_read` after every read has happened; the hook may turn the +/// read into a failure. +struct WatchedFs io::Result<()>> { on_read: H, } -impl DurableFs for WatchedFs { +impl io::Result<()>> DurableFs for WatchedFs { type File = File; fn create_new(&mut self, path: &Path) -> io::Result { @@ -307,13 +364,13 @@ impl DurableFs for WatchedFs { fn read(&mut self, path: &Path) -> io::Result> { let bytes = StdFs.read(path); - (self.on_read)(path); + (self.on_read)(path)?; bytes } fn read_at_most(&mut self, path: &Path, limit: usize) -> io::Result>> { let bytes = StdFs.read_at_most(path, limit); - (self.on_read)(path); + (self.on_read)(path)?; bytes } @@ -433,7 +490,7 @@ fn a_lease_of_another_owner_or_of_none_is_refused() { } /// The gate over `fixture` as its owner under `held`, reporting every read to `on_read`. -fn begin_watched( +fn begin_watched io::Result<()>>( fixture: &Fixture, held: &mut ExclusiveAdmissionGuard, on_read: H, @@ -449,7 +506,10 @@ fn a_guard_admitted_for_another_installation_is_refused_before_anything_is_read( let (fixture, other) = (Fixture::bound(&committed), Fixture::bound(&committed)); let before = fixture.snapshot(); let mut reads = 0; - let outcome = begin_watched(&fixture, &mut other.paths().admission(), |_| reads += 1); + let outcome = begin_watched(&fixture, &mut other.paths().admission(), |_| { + reads += 1; + Ok(()) + }); let unchanged = fixture.snapshot() == before; assert_eq!( (outcome, reads, unchanged), @@ -465,7 +525,11 @@ fn an_installation_replaced_while_the_journal_is_read_is_refused() { let committed = manifest(phase_code::ADMIT); let counted = Fixture::bound(&committed); let mut total = 0; - begin_watched(&counted, &mut counted.paths().admission(), |_| total += 1).unwrap(); + begin_watched(&counted, &mut counted.paths().admission(), |_| { + total += 1; + Ok(()) + }) + .unwrap(); let fixture = Fixture::bound(&committed); let installation = fixture.base.0.clone(); let moved = installation.with_extension("moved"); @@ -475,6 +539,7 @@ fn an_installation_replaced_while_the_journal_is_read_is_refused() { if seen == total { fs::rename(&installation, &moved).unwrap(); } + Ok(()) }); fs::rename(&moved, &installation).unwrap(); assert_eq!(outcome, Err(ConversionError::NotAdmitted)); @@ -532,3 +597,497 @@ fn a_sibling_generation_is_never_read_instead_of_the_root_named_one() { Err(ConversionError::ForeignLeaseOwner) ); } + +/// A session over `fixture`'s installation as its owner, under `held`. +fn open<'a>( + fixture: &Fixture, + paths: &'a Paths, + held: &'a mut ExclusiveAdmissionGuard, +) -> ConversionSession<'a> { + begin_conversion(held, &mut StdFs, &fixture.provider, paths.begin(OWNER)).unwrap() +} + +/// The digest of the manifest generation `revision` in `dir`. +fn generation_digest(dir: &Path, revision: u64) -> [u8; 32] { + content_digest(&fs::read(generation_path(dir, revision)).unwrap()) +} + +#[test] +fn the_owner_renews_its_lease_in_admit_and_in_convert_and_the_session_follows_the_root() { + let mut seen = Vec::new(); + for phase in [phase_code::ADMIT, phase_code::CONVERT] { + let committed = manifest(phase); + let mut fixture = Fixture::bound(&committed); + let mut expected = committed.clone(); + expected.journal_revision += 1; + expected.lease_expires_unix_ms = Some(RENEWED_EXPIRY); + let generation = fixture.root_generation(); + let followed = { + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + let renewed = session.renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY); + let token = MigrationFence::from_manifest(&expected); + ( + // The commit is returned: one root generation further. + renewed.map(|committed| committed.root_generation == generation + 1), + session.manifest() == &expected, + session.fence() == token && session.is_admitted(), + ) + }; + // A fresh gate reads what the root now names: the renewed manifest and nothing else moved. + seen.push((followed, fixture.committed(OWNER) == Ok(expected))); + } + assert_eq!(seen, vec![((Ok(true), true, true), true); 2]); +} + +#[test] +fn an_expiry_that_does_not_move_forward_is_refused_and_writes_nothing() { + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let before = fixture.snapshot(); + let outcomes: Vec<_> = { + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + [LEASE_EXPIRY, LEASE_EXPIRY - 1] + .into_iter() + .map(|expiry| { + session + .renew_lease(&mut StdFs, &mut fixture.provider, expiry) + .map(|_| ()) + }) + .collect() + }; + let refused = ConversionError::Migration(MigrationExecutionError::InvalidLeaseRenewal); + assert_eq!(outcomes, vec![Err(refused); 2]); + assert_eq!(fixture.snapshot(), before); +} + +#[test] +fn a_session_that_another_owner_took_over_from_is_refused_before_any_write() { + let committed = manifest(phase_code::ADMIT); + let mut fixture = Fixture::bound(&committed); + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + fixture.take_over(&committed); + let before = fixture.snapshot(); + let outcome = session + .renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY) + .map(|_| ()); + let stale = AuthorityError::Journal(JournalDurableError::Authority( + MigrationExecutionError::StaleMigrationOwner, + )); + assert_eq!(outcome, Err(ConversionError::Authority(stale))); + assert_eq!(fixture.snapshot(), before); +} + +#[test] +fn a_failed_root_commit_leaves_the_session_where_it_was_and_the_retry_adopts_the_candidate() { + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let journal_dir = fixture.journal_dir(); + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + // The anchor refuses the root's preparation: the renewal is already in the journal as revision + // `r + 1`, the root still names `r`, and nothing is left pending. + fixture + .provider + .inject(Fault::BeforePersist(AnchorOp::Prepare)); + let failed = session.renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY); + let candidate = generation_digest(&journal_dir, REVISION + 1); + let kept = session.manifest().journal_revision == REVISION; + let retried = session + .renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY) + .map(|_| ()); + let adopted = generation_digest(&journal_dir, REVISION + 1) == candidate; + let revision = session.manifest().journal_revision; + assert_eq!( + (failed.is_err(), kept, retried, adopted, revision), + (true, true, Ok(()), true, REVISION + 1) + ); +} + +#[cfg(unix)] +#[test] +fn a_step_after_the_installation_moved_is_refused_before_any_write() { + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let before = fixture.snapshot(); + let installation = fixture.base.0.clone(); + let moved = installation.with_extension("moved"); + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + fs::rename(&installation, &moved).unwrap(); + let outcome = session + .renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY) + .map(|_| ()); + fs::rename(&moved, &installation).unwrap(); + assert_eq!(outcome, Err(ConversionError::NotAdmitted)); + assert_eq!(fixture.snapshot(), before); +} + +/// A renewal over a fixture, with `on_read` told of every read it makes. +fn renewal_watched io::Result<()>>( + fixture: &mut Fixture, + on_read: H, +) -> Result<(), ConversionError> { + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(fixture, &paths, &mut held); + let mut watched = WatchedFs { on_read }; + session + .renew_lease(&mut watched, &mut fixture.provider, RENEWED_EXPIRY) + .map(|_| ()) +} + +#[cfg(unix)] +#[test] +fn an_admission_lost_after_a_step_is_reported_though_the_step_committed() { + // The step is deterministic, so a first run counts its reads and a second run moves the + // installation away right after the last one: only the check after the step sees it. + let committed = manifest(phase_code::ADMIT); + let mut total = 0; + renewal_watched(&mut Fixture::bound(&committed), |_| { + total += 1; + Ok(()) + }) + .unwrap(); + let mut fixture = Fixture::bound(&committed); + let installation = fixture.base.0.clone(); + let moved = installation.with_extension("moved"); + let mut seen = 0; + let outcome = renewal_watched(&mut fixture, |_| { + seen += 1; + if seen == total { + fs::rename(&installation, &moved).unwrap(); + } + Ok(()) + }); + fs::rename(&moved, &installation).unwrap(); + let renewed = fixture.committed(OWNER).map(|m| m.lease_expires_unix_ms); + assert_eq!( + (outcome, renewed), + (Err(ConversionError::NotAdmitted), Ok(Some(RENEWED_EXPIRY))) + ); +} + +#[test] +fn the_renewal_builder_refuses_another_token_a_missing_lease_and_a_terminal_journal() { + let committed = manifest(phase_code::ADMIT); + let token = MigrationFence::from_manifest(&committed); + let stale = MigrationFence { + journal_revision: REVISION - 1, + ..token + }; + let mut unowned = committed.clone(); + (unowned.has_lease_owner, unowned.lease_owner_id) = (false, None); + unowned.lease_expires_unix_ms = None; + let done = manifest(phase_code::DONE); + let outcomes = vec![ + renewed_lease(&committed, &stale, RENEWED_EXPIRY), + renewed_lease(&unowned, &token, RENEWED_EXPIRY), + renewed_lease(&done, &MigrationFence::from_manifest(&done), RENEWED_EXPIRY), + ]; + let refusals: Vec<_> = outcomes.into_iter().map(Result::unwrap_err).collect(); + assert_eq!( + refusals, + vec![ + MigrationExecutionError::StaleMigrationOwner, + MigrationExecutionError::InvalidLeaseRenewal, + MigrationExecutionError::TerminalPhase, + ] + ); +} + +/// The names in `dir` of generations moved aside with their bytes preserved. +fn quarantined(dir: &Path) -> Vec { + let mut names: Vec = fs::read_dir(dir) + .unwrap() + .map(|entry| entry.unwrap().path()) + .filter(|path| path.to_string_lossy().contains(".rejected-")) + .collect(); + names.sort(); + names +} + +#[test] +fn a_differing_candidate_of_a_crashed_attempt_is_quarantined_and_does_not_block_the_next_step() { + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let journal_dir = fixture.journal_dir(); + { + // The first attempt dies after writing its candidate: the anchor refuses the root's preparation. + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + fixture + .provider + .inject(Fault::BeforePersist(AnchorOp::Prepare)); + session + .renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY) + .unwrap_err(); + } + let candidate = generation_digest(&journal_dir, REVISION + 1); + // The restarted owner derives another expiry from its clock. + let other = RENEWED_EXPIRY + 1; + let outcome = { + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + session + .renew_lease(&mut StdFs, &mut fixture.provider, other) + .map(|_| ()) + }; + let kept: Vec<_> = quarantined(&journal_dir) + .iter() + .map(|path| content_digest(&fs::read(path).unwrap())) + .collect(); + let renamed = generation_digest(&journal_dir, REVISION + 1) != candidate; + let expiry = fixture.committed(OWNER).map(|m| m.lease_expires_unix_ms); + assert_eq!( + (outcome, kept, renamed, expiry), + (Ok(()), vec![candidate], true, Ok(Some(other))) + ); +} + +#[test] +fn a_renewal_is_busy_while_another_root_commit_holds_the_lock_and_writes_nothing() { + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + let lock = RootCommitGuard::acquire(&fixture.root_dir()).unwrap(); + let before = fixture.snapshot(); + let busy = session + .renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY) + .map(|_| ()); + let unchanged = fixture.snapshot() == before; + drop(lock); + let retried = session + .renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY) + .map(|_| ()); + assert_eq!( + (busy, unchanged, retried), + (Err(ConversionError::RootBusy), true, Ok(())) + ); +} + +/// A hook that fails the `total`-th read it is told of and lets every other one through. +fn failing_nth_read(total: usize) -> impl FnMut(&Path) -> io::Result<()> { + let mut seen = 0; + move |_: &Path| { + seen += 1; + if seen == total { + Err(io::Error::other("injected")) + } else { + Ok(()) + } + } +} + +/// How many reads a renewal over a fresh fixture makes, with `fault` injected into the anchor if given. +/// The step is deterministic, so the last of them is the read-back of the committed journal. +fn renewal_reads(fault: Option) -> usize { + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + if let Some(fault) = fault { + fixture.provider.inject(fault); + } + let mut total = 0; + let _ = renewal_watched(&mut fixture, |_| { + total += 1; + Ok(()) + }); + total +} + +#[test] +fn a_read_back_that_fails_spends_the_session_though_the_step_committed() { + let total = renewal_reads(None); + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let mut watched = WatchedFs { + on_read: failing_nth_read(total), + }; + let (first, later, snapshot) = { + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + let first = session.renew_lease(&mut watched, &mut fixture.provider, RENEWED_EXPIRY); + let later = session.renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY + 1); + ( + first, + later.map(|_| ()), + session.manifest().journal_revision, + ) + }; + let first_is_unreadable = matches!(first, Err(ConversionError::Unreadable(_))); + assert_eq!( + (first_is_unreadable, later, snapshot), + (true, Err(ConversionError::Spent), REVISION) + ); + // The renewal itself was committed: a new gate reads it. + assert_eq!( + fixture.committed(OWNER).map(|m| m.lease_expires_unix_ms), + Ok(Some(RENEWED_EXPIRY)) + ); +} + +#[test] +fn a_commit_that_reports_an_error_after_it_landed_is_settled_by_the_root() { + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let first = { + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + // The anchor commits the root but reports the outcome as unavailable. + fixture + .provider + .inject(Fault::AfterPersist(AnchorOp::Commit)); + let first = session.renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY); + (first, session.manifest().journal_revision) + }; + let is_committed = matches!(first.0, Err(ConversionError::Committed(_))); + assert_eq!((is_committed, first.1), (true, REVISION + 1)); + // The root names the renewal, which the session followed. + assert_eq!( + fixture.committed(OWNER).map(|m| m.lease_expires_unix_ms), + Ok(Some(RENEWED_EXPIRY)) + ); +} + +#[test] +fn a_commit_error_whose_outcome_cannot_be_read_back_spends_the_session() { + let anchor_fault = Fault::AfterPersist(AnchorOp::Commit); + let total = renewal_reads(Some(anchor_fault)); + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let mut watched = WatchedFs { + on_read: failing_nth_read(total), + }; + let (first, later) = { + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + fixture.provider.inject(anchor_fault); + let first = session.renew_lease(&mut watched, &mut fixture.provider, RENEWED_EXPIRY); + let later = session.renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY + 1); + (first, later.map(|_| ())) + }; + let first_is_unsettled = matches!(first, Err(ConversionError::Unsettled(_))); + assert_eq!( + (first_is_unsettled, later), + (true, Err(ConversionError::Spent)) + ); +} + +#[test] +fn a_journal_directory_outside_the_installation_or_inside_the_root_is_refused_before_anything_is_read( +) { + let fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let outside = Dir::new(); + let mut outcomes = Vec::new(); + let root = fixture.root_dir(); + let slot = root.join("slot-a"); + for journal in [ + outside.0.clone(), + fixture.base.0.clone(), + root.clone(), + slot, + ] { + let mut paths = fixture.paths(); + paths.journal = journal; + let mut held = paths.admission(); + let mut reads = 0; + let mut watched = WatchedFs { + on_read: |_: &Path| { + reads += 1; + Ok(()) + }, + }; + let outcome = begin_conversion( + &mut held, + &mut watched, + &fixture.provider, + paths.begin(OWNER), + ); + outcomes.push((outcome.map(|_| ()), reads)); + } + let refused = (Err(ConversionError::JournalMisplaced), 0); + assert_eq!(outcomes, vec![refused; 4]); +} + +#[cfg(unix)] +#[test] +fn a_journal_directory_replaced_after_begin_is_refused_before_any_write() { + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let before = fixture.snapshot(); + let journal = fixture.journal_dir(); + let moved = journal.with_extension("moved"); + let (paths, mut held) = (fixture.paths(), fixture.paths().admission()); + let mut session = open(&fixture, &paths, &mut held); + // A copy of the journal takes the place of the directory the session pinned. + fs::rename(&journal, &moved).unwrap(); + fs::create_dir(&journal).unwrap(); + fs::copy( + generation_path(&moved, REVISION), + generation_path(&journal, REVISION), + ) + .unwrap(); + let outcome = session + .renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY) + .map(|_| ()); + let admitted = session.is_admitted(); + fs::remove_dir_all(&journal).unwrap(); + fs::rename(&moved, &journal).unwrap(); + assert_eq!( + (outcome, admitted, fixture.snapshot() == before), + (Err(ConversionError::NotAdmitted), false, true) + ); +} + +#[cfg(unix)] +#[test] +fn a_symlinked_journal_directory_retargeted_after_begin_does_not_redirect_the_step() { + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + let (journal, link) = (fixture.journal_dir(), fixture.base.0.join("journal-link")); + let other = fixture.base.0.join("journal-other"); + std::os::unix::fs::symlink(&journal, &link).unwrap(); + // A copy of the journal sits where the link will point later. + fs::create_dir(&other).unwrap(); + fs::copy( + generation_path(&journal, REVISION), + generation_path(&other, REVISION), + ) + .unwrap(); + let mut paths = fixture.paths(); + paths.journal = link.clone(); + let mut held = paths.admission(); + let mut session = + begin_conversion(&mut held, &mut StdFs, &fixture.provider, paths.begin(OWNER)).unwrap(); + fs::remove_file(&link).unwrap(); + std::os::unix::fs::symlink(&other, &link).unwrap(); + let outcome = session + .renew_lease(&mut StdFs, &mut fixture.provider, RENEWED_EXPIRY) + .map(|_| ()); + let (renewed_in_pinned, untouched_copy) = ( + generation_path(&journal, REVISION + 1).exists(), + !generation_path(&other, REVISION + 1).exists(), + ); + assert_eq!( + (outcome, renewed_in_pinned, untouched_copy), + (Ok(()), true, true) + ); +} + +#[cfg(unix)] +#[test] +fn an_admission_lost_after_a_commit_that_reported_an_error_is_still_reported() { + // The anchor commits the root but reports the outcome as unavailable, and the installation is + // moved away right after the last read of the step: the step is committed and the loss is reported. + let anchor_fault = Fault::AfterPersist(AnchorOp::Commit); + let total = renewal_reads(Some(anchor_fault)); + let mut fixture = Fixture::bound(&manifest(phase_code::ADMIT)); + fixture.provider.inject(anchor_fault); + let installation = fixture.base.0.clone(); + let moved = installation.with_extension("moved"); + let mut seen = 0; + let outcome = renewal_watched(&mut fixture, |_| { + seen += 1; + if seen == total { + fs::rename(&installation, &moved).unwrap(); + } + Ok(()) + }); + fs::rename(&moved, &installation).unwrap(); + let renewed = fixture.committed(OWNER).map(|m| m.lease_expires_unix_ms); + assert_eq!( + (outcome, renewed), + (Err(ConversionError::NotAdmitted), Ok(Some(RENEWED_EXPIRY))) + ); +} diff --git a/docs/native/R15-SECURE-STORAGE-CONTRACT.md b/docs/native/R15-SECURE-STORAGE-CONTRACT.md index cc0874f27..86dbe323a 100644 --- a/docs/native/R15-SECURE-STORAGE-CONTRACT.md +++ b/docs/native/R15-SECURE-STORAGE-CONTRACT.md @@ -2388,6 +2388,26 @@ can be entered) or `CONVERT` (a resumed conversion), the final inventory was cap lease belongs to the caller. A lease owned by another owner, or by none, is refused: the gate takes nothing over, and a restart is the takeover followed by a fresh begin. The gate reads no clock, writes nothing and holds no key, and the identity of the admitted directories is checked before the reads and again after them. +The session then moves the journal only through the fenced journal-owner operations, one step at a time, +each built from the session's own snapshot. The journal directory must lie strictly below the admitted +installation directory, outside the authority root, and is pinned for the session, and every file operation goes through the canonical +paths of the root, the installation and the journal directory validated at begin, never through the +caller's spelling, so a symlink retargeted later cannot redirect a step. A step checks the admission and +that pin, builds the successor and proves it against the snapshot, takes the root event through the held +admission (so the identity check and the root lock are coupled), commits with the key route and active +epoch the committed root itself names (never a caller's), and then, whatever the commit reported, reads +the committed journal back under that same root event, so no other root commit can come between, and lets +the root settle the outcome: if it names the successor the step is committed (even when the commit +reported an error, such as an ambiguous anchor outcome) and the session follows it; if it names anything +else the step did not commit and the snapshot is kept. Before the commit a failed step leaves the tree and +the snapshot as they were; a retry rebuilds the same successor, which the journal adopts if it is +identical to a candidate already written, and a differing candidate of a crashed attempt is moved aside +with its bytes preserved, so it never blocks the next step. If the read-back fails, or names a journal +that is neither of the two where the commit reported success, the session is spent and every later step +is refused until the caller begins again. The first step is the owner's lease renewal at +an expiry the caller chooses: strictly after the held one, accepted after the held one has lapsed while +nobody has taken over, and refused as a stale owner once another owner has; the root commit, with its +durability result, is returned. A lost admission is reported even when it was lost after the commit. A token check performed as a separate preflight is insufficient. Every mutation-capable adapter operation therefore exposes the semantic equivalent of `with_fence(fencing_generation, mutation_and_durability)`: it acquires the cross-process migration diff --git a/docs/native/r15/GATE4D-JOURNAL-DURABLE-EVIDENCE.md b/docs/native/r15/GATE4D-JOURNAL-DURABLE-EVIDENCE.md index 2bd45a118..cf1dc8832 100644 --- a/docs/native/r15/GATE4D-JOURNAL-DURABLE-EVIDENCE.md +++ b/docs/native/r15/GATE4D-JOURNAL-DURABLE-EVIDENCE.md @@ -665,6 +665,59 @@ after authentication (D2b-2), and resolving the journal key through the authenti readability proof for a root already at the target epoch (D2b-3a), used by every journal-owner operation (D2b-3b). Each is described in its own section below. +## Slice C2b-1 — the session renews its lease + +The conversion session of C2a could only be read, and the D3b renewal had no caller. `ConversionSession::renew_lease` +is the first step that moves the journal through the gate, and the first real caller of `commit_lease_renewal`. + +| Step | Rule | +|---|---| +| builder | `renewed_lease(manifest, fence, expires)`: the fence checked, the next revision, the expiry replaced, the renewal relation proved (`assert_renewal_successor`: a named lease, an expiry strictly after the held one, not terminal) and the manifest encodable; the expiry is the caller's, Core reads no clock | +| key route | `read_committed_journal` also returns the root key reference and active epoch the committed root names; the step carries those, so no key, route or epoch comes from the caller and the commit cannot meet `KeyRotationNotAdmitted` | +| journal directory | pinned at begin (opened before it is canonicalised, as the admission does) and required to lie strictly below the admitted installation directory and outside the authority root (`JournalMisplaced` otherwise, before any read; inside the root a journal generation could collide with a root slot); every step and `is_admitted` look at the pin again, so a directory replaced after begin is `NotAdmitted` before any write | +| canonical paths | the session canonicalises the installation, the root and the journal directory once at begin, checks the admission against the canonical scope and does every read and write through those paths, so a symlink retargeted after begin cannot redirect a step to a directory that was never pinned | +| root event | `try_root_commit` on the held admission (`RootBusy` when another root commit holds the lock, nothing written), `root_guard` while the lock is held, and `commit_lease_renewal_held`, the commit of `commit_lease_renewal` under a root event the caller holds, so the identity check and the root lock are coupled | +| settle | whatever the commit reported, the committed journal is read back under that same root event, where no other root commit can come between, and the root settles the outcome: success and the renewal named (installed); success and another journal (`Superseded`, spent) or an unreadable root (`Unreadable`, spent); an error and the renewal named (`Committed`: the step landed although the commit reported an error, such as `AfterPersist(Commit)` at the anchor, and the session follows the root); an error and another journal (`Authority`: not committed, snapshot kept, a stale session is refused as before); an error and an unreadable root (`Unsettled`, spent). A spent session answers every later step with `Spent` | +| conflict | a differing candidate at the next revision (a crashed attempt that used another expiry) is moved aside with its bytes preserved (`CandidateConflict::Quarantine`, the policy of the composed journal commits, which hold the root lock and have read the committed binding), so it never blocks the next step; an identical candidate is adopted | +| result | the `RootCommitted` of the commit, whose `directories` says whether the directory entries are confirmed durable | + +Decisions, disclosed: (a) renewal alone first, the smallest slice that gives D3b a caller; entering `CONVERT` and +the cursor follow; (b) a lost admission after a step is reported as `NotAdmitted` even though the step may have +committed, because a check after the commit can report but not undo it, and the snapshot then is the committed +journal; (c) no internal retry, and a session spent by a failed read-back is replaced by a new begin rather than +resynchronised, while a commit error that the root shows to have landed is reported as `Committed` and followed; (d) the key reference and epoch come from the committed root; (e) the journal directory must lie +below the admitted installation directory and outside the authority root: the contract does not place the journal, but an exclusive admission +guards nothing outside the installation, so a journal outside it would be written without the admission's +protection; (f) the path-based file operations cannot be made atomic with an identity check, so the checks bracket +every step (before it, under the root lock at the root event, and after it) and fail closed. + +Proof (`gate4d_conversion_entry_test`): the owner renews in `ADMIT` and in `CONVERT`, the commit is returned (one root +generation further), the session equals the expected renewal (revision plus one, expiry replaced, nothing else) and a +fresh gate reads the same from the root, with the original lease long past its expiry; an expiry equal to or before +the held one is refused with nothing written; a session whose journal another owner took over is refused as a stale +owner before any write; a busy root is `RootBusy` with nothing written and the retry succeeds; with the anchor +refusing the root's preparation the session keeps its snapshot, the journal holds the renewal as an unadopted +candidate and the retry adopts it without rewriting the generation; a restarted owner with another expiry +quarantines the candidate (bytes preserved) and renews; an installation moved away before the step is refused before +any write and one moved away right after the last read of the step is reported while the step stays committed, also when the commit reported an error that the root shows to have landed; a +failed read-back is `Unreadable`, the renewal is committed and the next step is `Spent`; with the anchor reporting +an error after it committed the root the step is `Committed` and the session follows it, and when the read-back also +fails it is `Unsettled` and the next step is `Spent`; a symlink to the journal directory retargeted after begin does +not redirect the step; a journal directory outside the installation, the installation itself, the root or a slot inside it is refused before any read, and one replaced after begin is refused +before any write; the builder refuses another token, a missing lease and a terminal journal. Mutation-checked: the +admission check before and after the step, the snapshot refresh, the carried epoch and fence, the conflict policy, the +spent states, the settle arms for a commit error, the root-busy mapping, the journal pin, the canonical journal path, +the containment rule and its order, and each of the builder's three checks, removed one at a time, fail the test that +owns them. Two checks have no failing test: the comparison of the read-back with the committed renewal where the +commit reported success (`Superseded`, and installing the read-back unchecked) is defence in depth, because the read +happens under the root event and is authenticated against the digest the commit just bound, so it can only differ +through a defect. Residual: the root directory is canonicalised once at begin like the journal directory, but only +the journal directory's retargeting is tested; the file operations remain path-based, so no check is atomic with them: a local +attacker who can rename directories inside the installation can still race a step, at worst misplacing a generation that the +root then does not find (`RECOVERY_REQUIRED`, never a forged one: the bytes are sealed and bound by the root's digest). +Handle-relative (openat-style) journal and root I/O is an adapter capability and an acceptance criterion on #359 for the +platform adapter work, shared by every journal-owner operation. + ## Slice C2a — the entry gate of the exclusive conversion driver The journal-owner operations take no admission guard, and nothing read the committed journal state without a @@ -1020,8 +1073,7 @@ this API's reach; the journal key route of D2b-3 fails closed on a `Revoked` or (`DurableFs::read_at_most`) and the page-directory listing is bounded (`DurableFs::list_dir_at_most`; both defaults must be overridden by an adapter over real files, which `StdFs` does), but the Gate 3 post-promotion verify and the page, marker and root reads still use the whole-file `DurableFs::read`. Applying the same size limits to them is a separate slice, recorded as an acceptance criterion on #359. -- A caller of the lease renewal: `commit_lease_renewal` exists (D3b) but no orchestrator renews a lease yet. -- Conversion over the verified page set: the entry gate exists (C2a: exclusive admission by construction, the root-bound manifest re-read, `final_inventory_captured = 1`; maintainer decision D), but nothing yet moves the journal under it: entering `CONVERT`, advancing the cursor and renewing the lease through the gate (C2b), the cursor-driven iteration with a per-entry step and the crash/resume evidence (C2c). +- Conversion over the verified page set: the entry gate (C2a: exclusive admission by construction, the root-bound manifest re-read, `final_inventory_captured = 1`; maintainer decision D) and the lease renewal through it (C2b-1: `ConversionSession::renew_lease`, the first caller of the D3b renewal) exist, but nothing yet enters `CONVERT` or moves the cursor under the gate: entering `CONVERT` and advancing the cursor (C2b-2), the cursor-driven iteration with a per-entry step and the crash/resume evidence (C2c). - Write barrier of the final capture: `commit_inventory_capture` takes no admission guard; the barrier is the durable `ADMIT` phase the orchestrator establishes by draining writers, and the write path must refuse ordinary mutating writes by that phase (`ordinary_mutating_writes_admitted`) before the final capture has a caller (Gate 4E/5; acceptance criterion on #359). - Inheriting unchanged pages: the C1b-2 reader now returns the authenticated page references, so the store may accept a page that keeps an earlier generation if those references name exactly its bytes (acceptance criterion on #359, a follow-up slice). Until then every page of a capture is rewritten at the new revision. - Reclaiming abandoned pending directories: an attempt that was killed, or a finished capture dropped without `discard`, leaves inert files under `inventory/pending-*`, and every attempt leaves its empty page directories; none is authority or ever read. Reclaiming them needs a directory-removal primitive and a sweep that knows no live attempt owns them, as for the orphaned digest directories of a discarded capture (acceptance criterion on #359).