Skip to content

Commit 2bc7ccc

Browse files
committed
Add verified checkpoints and native RF3 snapshot catch-up
1 parent 516e769 commit 2bc7ccc

32 files changed

Lines changed: 1267 additions & 119 deletions

‎README.md‎

Lines changed: 9 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -8,9 +8,9 @@ The original load-testing prototype has been replaced. This repository implement
88

99
This is an early implementation of the clustered kernel. The default server topology has three persistent voting nodes and requires a majority for writes and strong reads. It needs no external database, Redis or message broker.
1010

11-
Implemented surfaces include document CRUD and field patches, partition-scoped unique/composite equality indexes, command deduplication, stream expected-revision append/read, scheduled work queues with fenced leases and retry/DLQ, atomic inbox completion, bounded property-graph traversal, ordered samples, exact vector search, BM25 and weighted reciprocal-rank fusion, a bounded read-only Q1 SQL dialect, API keys, field omission and field-use policies, verified backups, a .NET SDK and CLI.
11+
Implemented surfaces include document CRUD and field patches, partition-scoped unique/composite equality indexes, command deduplication, stream expected-revision append/read, scheduled work queues with fenced leases and retry/DLQ, atomic inbox completion, bounded property-graph traversal, ordered samples, exact vector search, BM25 and weighted reciprocal-rank fusion, a bounded read-only Q1 SQL dialect, API keys, field omission and field-use policies, verified backups, native Raft snapshots and empty-replica catch-up, offline journal compaction, a .NET SDK and CLI.
1212

13-
The [104-task implementation tracker](docs/implementation/status.json) records the remaining work. Raft snapshot compaction, shard movement, distributed multi-shard query planning, managed HNSW, topics and subscription groups, CDC, retention, schema migrations, external backup stores and full release qualification remain under development. The current server uses one replicated physical shard which contains many independent atomic partitions.
13+
The [104-task implementation tracker](docs/implementation/status.json) records the remaining work. Automatic canonical journal maintenance, large and interrupted snapshot transfer qualification, shard movement, distributed multi-shard query planning, managed HNSW, topics and subscription groups, CDC, retention, schema migrations, external backup stores and full release qualification remain under development. The current server uses one replicated physical shard which contains many independent atomic partitions.
1414

1515
The advertised profiles are `ProcessDurable` for embedded storage and `QuorumProcessDurable` for the cluster. The kernel has process-kill recovery tests and the cluster has real leader-loss and minority tests. Power-loss qualification, broader platform qualification and the 72-hour endurance gate are still required before advertising `LocalDurable`, `QuorumDurable` or production readiness.
1616

@@ -31,7 +31,9 @@ dotnet run --project src/KeyLoad.Cli -- status \
3131
http://localhost:5101 data/cluster/local-profile.json
3232
```
3333

34-
The standalone server requires explicit cluster identity, voter endpoints, shared peer authentication secrets, an administrator credential and a silo port through `KeyLoad` configuration. Set `KeyLoad:SiloAddress` to each node's advertised IP address for a cluster across hosts. HTTPS is required for cluster endpoints; the AppHost explicitly enables loopback HTTP for development. Peer RPCs and forwarding requests are signed for their recipient, path, query and body, timestamped and replay checked. Public HTTP requests authenticate against the replicated credential catalog. Client-supplied roles are never trusted.
34+
The standalone server requires explicit cluster identity, voter endpoints, shared peer authentication secrets, an administrator credential and a silo port through `KeyLoad` configuration. Set `KeyLoad:SiloAddress` to each node's advertised IP address for a cluster across hosts. HTTPS is required for cluster endpoints; the AppHost explicitly enables loopback HTTP for development. Peer RPCs and forwarding requests are signed for their recipient, path, query, body and Raft protocol headers, timestamped and replay checked. Large peer bodies use bounded temporary files. Public HTTP requests authenticate against the replicated credential catalog. Client-supplied roles are never trusted.
35+
36+
The default peer connection/RPC/request timeouts are 500/1500/2000 milliseconds, below the 4000–8000 millisecond election range. These values are configurable through the corresponding `KeyLoad` options; validation requires that ordering. Native snapshots are produced every `KeyLoad:SnapshotThreshold` committed entries (default 1024). Small snapshot catch-up is tested; large transfers still need qualification against the configured RPC deadline.
3537

3638
## .NET client
3739

@@ -93,16 +95,17 @@ page.ThrowIfFail();
9395

9496
Q1 supports projections and aliases, scalar parameters, comparisons, `AND`/`OR`/`NOT`, `IN`, `IS NULL`, `IS MISSING`, `ORDER BY`, `LIMIT` and `EXPLAIN`. Identifiers containing dots need double quotes. `id` and `revision` refer to canonical document identity/revision. Parameters use `@name`. JSON numbers use the decimal scalar policy. SQL is read-only; unsupported statements fail explicitly.
9597

96-
Queries require a matching point/equality index or explicit `AllowFullScan`. Scans, parser depth, tokens, rows, bytes and execution time have budgets. Cursor tokens bind the principal, policy epoch, schema, query and current storage cut; writes or policy changes can expire a cursor. Sensitive predicates and sorting require a field-use grant. Returned documents omit protected paths, including classified values nested in arrays.
98+
Queries require a matching point/equality index or explicit `AllowFullScan`. Scans, parser depth, tokens, rows, bytes and execution time have budgets. Cursor tokens bind the principal, policy epoch, schema, query, node identity, read generation and current storage cut. Writes, policy changes or snapshot installation can expire a cursor; compaction preserves its cut. Sensitive predicates and sorting require a field-use grant. Returned documents omit protected paths, including classified values nested in arrays.
9799

98100
Search accepts typed vector spaces and explicit text/vector fields. Both branches use one authorized read cut. Exact vector scores and BM25 ranks are combined with weighted RRF using one-based ranks. The managed ANN and graph retrieval extensions are tracked separately.
99101

100102
## Backups and Cartograph
101103

102-
The canonical backup includes the checksummed redo journal, database identity and a SHA-256 manifest. Domain data, schemas, credentials, outcomes and inbox receipts are journaled together. ZoneTree files can be rebuilt from that canonical journal. Restore validates every manifest file, creates a new incarnation, resets consensus routing metadata and leaves queue dispatch paused.
104+
The canonical backup includes the checksummed redo journal, database identity and a SHA-256 manifest. Domain data, schemas, credentials, outcomes and inbox receipts are journaled together. The journal can begin with a verified checkpoint followed by newer transaction frames. ZoneTree files can be rebuilt from that canonical history. Restore validates every manifest file, creates a new incarnation, resets consensus routing metadata and leaves queue dispatch paused.
103105

104106
```sh
105107
# Stop the node before using the offline CLI.
108+
dotnet run --project src/KeyLoad.Cli -- compact data/cluster/node1/database
106109
dotnet run --project src/KeyLoad.Cli -- backup data/cluster/node1/database backups/snapshot
107110
dotnet run --project src/KeyLoad.Cli -- restore backups/snapshot data/restored
108111

@@ -140,7 +143,7 @@ dotnet test --project tests/KeyLoad.RecoveryTests --no-build --no-restore
140143
dotnet test --project tests/KeyLoad.IntegrationTests --no-build --no-restore
141144
```
142145

143-
Tests use xUnit and Microsoft.Testing.Platform. Recovery qualification runs 1000 seeded real-process kills, checks complete-frame corruption and verifies a clean backup restore. Integration tests own the Aspire lifecycle and use three independent server processes with isolated persistent directories. They kill the elected leader, retry the same command, verify atomic effects on surviving voters and reject a minority write. No manually running AppHost is needed for tests.
146+
Tests use xUnit and Microsoft.Testing.Platform. Recovery qualification runs 1000 seeded real-process kills, checks complete-frame corruption and verifies a clean backup restore. Additional process kills cover checkpoint publication and all native Raft append paths immediately after acknowledgement. Integration tests own the Aspire lifecycle and use three independent server processes with isolated persistent directories. They kill the elected leader, retry the same command, verify atomic effects on surviving voters, reject a minority write and erase/restart one replica to require snapshot catch-up. No manually running AppHost is needed for tests.
144147

145148
Benchmarks can be started with `dotnet run -c Release --project benchmarks/KeyLoad.Benchmarks`. Comparative PostgreSQL/Marten/Wolverine and multi-node scaling qualification are part of the development plan; no comparative performance claims are made yet.
146149

‎docs/implementation/durability-audit.md‎

Lines changed: 6 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -10,9 +10,13 @@ The adapter compiles each operation under a write gate, writes a bounded frame,
1010

1111
The separate node ownership lock prevents concurrent writers. Returned bytes and staged mutations are copied. Identity and backup files are versioned; the backup manifest verifies every canonical file with SHA-256. A backup is made under the apply gate and reports that exact cut. Restore changes the incarnation, invalidates old signed tokens, resets the Raft/Orleans watermark and membership, and explicitly persists paused dispatch even if the old journal had unpaused it.
1212

13+
Format 2 adds a checkpoint prefix with ordered live records, per-frame SHA-256 checksums and a complete-image checksum/count. It records both the local journal position and the Raft apply position. Installation verifies the entire incoming image before closing the old generation, flushes the staged file and atomically replaces the canonical journal. The derived ZoneTree generation is rebuilt. Installation keeps the node identity, increments the read generation and invalidates existing query cursors; offline compaction preserves the logical cut and cursor generation. Process-kill regressions cover snapshot write/flush and the two publication boundaries, including reclamation after reopening.
14+
1315
## Replicated acknowledgement
1416

15-
The .NEXT 6.8.1 persistent Raft log owns terms, votes, log matching and consensus commit. `DurableRaftLog` reimplements all public append paths on the native persistent log and awaits `FlushAsync` before an AppendEntries acknowledgement can be returned. `FlushOnCommit` alone would not establish this follower append boundary. A regression invokes append through the actual `IPersistentState` interface and reopens an uncommitted tail.
17+
The .NEXT 6.8.1 persistent Raft log owns terms, votes, log matching and consensus commit. `DurableRaftLog` requires native manual checkpoints (`FlushInterval = Timeout.InfiniteTimeSpan`), serializes its public append/commit paths and awaits the native foreground `FlushAsync` before acknowledgement. Background flush notifications do not establish this boundary for an isolated first append or snapshot installation. Four regressions invoke append through the actual `IPersistentState` interface, kill the process immediately after ACK and reopen the tail using private page memory; they do not rely on graceful disposal or shared memory pages surviving exit.
18+
19+
The state machine produces native Raft snapshots from a verified canonical store cut. A snapshot-only regression verifies that installation acknowledges without a subsequent append and reopens at that cut. The RF3 test erases one stopped follower's data directory, requires native snapshot installation, checks the fresh node identity and read generation and verifies documents and command outcomes through the HTTP API. Incoming bodies are authenticated and rewound for both `Body` and `BodyReader`; outgoing one-shot payloads are serialized once into a bounded disk spool. Raft term, snapshot metadata, request ID and content type are included in the peer signature.
1620

1721
The bounded leader writer replicates a trusted operation with one leader-chosen evaluation time. It returns success after consensus commit and local state-machine apply, then resolves the persisted outcome under current authorization. Cancellation or response loss after admission has an unknown write outcome; retries must retain the command ID. A minority cannot acknowledge writes.
1822

@@ -26,4 +30,4 @@ Local tests run on macOS arm64 with .NET SDK 10.0.401. The recovery suite execut
2630

2731
The local result and crash-stage distribution are recorded in `kernel-qualification.json`. Each trial writes its seed, fault stage, mutation index, recovered values and platform to `artifacts/qualification/crash-trials-*.jsonl`; CI retains these as run artifacts.
2832

29-
The advertised profiles remain `ProcessDurable` and `QuorumProcessDurable`. These tests do not establish filesystem directory-entry persistence, storage-controller guarantees, real power-cut behavior or every operating system/filesystem combination. Linux/macOS/Windows CI, network faults and subsequent snapshot/compaction qualification are tracked separately. Power-loss qualification, network-partition coverage, long histories, shard movement and the 72-hour endurance gate remain required for the broader architecture's durable release profile.
33+
The advertised profiles remain `ProcessDurable` and `QuorumProcessDurable`. These tests do not establish filesystem directory-entry persistence, storage-controller guarantees, real power-cut behavior or every operating system/filesystem combination. Linux/macOS/Windows CI and network faults are tracked separately. Native interrupted transfer/corruption recovery, large snapshots under transport deadlines, automatic canonical compaction, power-loss qualification, network partitions, long histories, shard movement and the 72-hour endurance gate remain required for the broader architecture's durable release profile.

‎docs/implementation/kernel-qualification.json‎

Lines changed: 10 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -7,13 +7,17 @@
77
"restore": "locked-mode passed",
88
"build": "passed with zero warnings",
99
"unitTests": {
10-
"passed": 22,
10+
"passed": 26,
1111
"failed": 0
1212
},
1313
"recoveryTests": {
14-
"passed": 25,
14+
"passed": 43,
1515
"failed": 0,
1616
"seededProcessCrashes": 1000,
17+
"checkpointPublicationCrashes": 8,
18+
"nativeRaftAcknowledgementCrashes": 4,
19+
"nativeSnapshotTests": 2,
20+
"peerSecurityTests": 6,
1721
"stages": {
1822
"PayloadWritten": 201,
1923
"MutationApplied": 200,
@@ -24,15 +28,17 @@
2428
"evidence": "CI retains artifacts/qualification/crash-trials-*.jsonl"
2529
},
2630
"rf3Integration": {
27-
"passed": 2,
31+
"passed": 3,
2832
"failed": 0,
2933
"scenarios": [
3034
"leader process kill",
3135
"same-command outcome retry",
3236
"atomic inbox/effects/ACK",
3337
"minority read/write rejection",
3438
"voter restart",
35-
"unsigned peer RPC rejection"
39+
"unsigned peer RPC rejection",
40+
"empty-replica native snapshot catch-up",
41+
"command outcome preserved across snapshot installation"
3642
]
3743
},
3844
"durability": [

‎docs/implementation/status.json‎

Lines changed: 9 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -262,8 +262,15 @@
262262
},
263263
"KL-035": {
264264
"title": "Replica snapshot і installation",
265-
"status": "pending",
266-
"evidence": []
265+
"status": "in_progress",
266+
"evidence": [
267+
"src/KeyLoad.Storage.ZoneTree/Checkpoints.cs",
268+
"src/KeyLoad.Replication/ReplicatedStateMachine.cs",
269+
"tests/KeyLoad.UnitTests/CheckpointTests.cs",
270+
"tests/KeyLoad.RecoveryTests/RaftSnapshotTests.cs",
271+
"tests/KeyLoad.RecoveryTests/RecoveryTests.cs",
272+
"tests/KeyLoad.IntegrationTests/ClusterTests.cs"
273+
]
267274
},
268275
"KL-036": {
269276
"title": "Controlled partition movement",

‎src/KeyLoad.Abstractions/Queries.cs‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -17,4 +17,5 @@ public sealed record SearchRequest(PartitionRef Partition, string Collection, st
1717
string? VectorField = null, float[]? Vector = null, VectorSpace? Space = null, int Limit = 10,
1818
double TextWeight = 1, double VectorWeight = 1, int FusionConstant = 60);
1919
public sealed record BackupReceipt(string Id, long Position);
20-
public sealed record NodeStatus(string NodeId, Guid Incarnation, long Applied, string? Leader, int Voters, DurabilityProfile Durability, bool RoutingReady, int ProcessId);
20+
public sealed record NodeStatus(string NodeId, Guid Incarnation, long Applied, string? Leader, int Voters, DurabilityProfile Durability, bool RoutingReady, int ProcessId,
21+
long ReadGeneration = 0);

‎src/KeyLoad.Abstractions/Storage/StorageContracts.cs‎

Lines changed: 6 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -3,6 +3,7 @@ namespace KeyLoad.Storage;
33
public sealed record KeyValueRecord(byte[] Key, byte[] Value);
44
public sealed record StorageMutation(byte[] Key, byte[]? Value);
55
public sealed record ScanPage(KeyValueRecord[] Records, bool HasMore);
6+
public sealed record StorageSnapshot(Guid Incarnation, long Position, long AppliedPosition, long RecordCount);
67
public interface IKeyValueView
78
{
89
byte[]? Get(byte[] key);
@@ -15,7 +16,7 @@ public interface IAtomicTransaction : IKeyValueView
1516
void Reset();
1617
}
1718
public sealed record StoreIdentity(int FormatVersion, int KeyCodecVersion, Guid NodeId, Guid Incarnation,
18-
byte[] SigningKey, DurabilityProfile Durability, bool DispatchPaused = false);
19+
byte[] SigningKey, DurabilityProfile Durability, bool DispatchPaused = false, long ReadGeneration = 0);
1920
public interface IAtomicStore : IDisposable
2021
{
2122
StoreIdentity Identity { get; }
@@ -24,6 +25,10 @@ public interface IAtomicStore : IDisposable
2425
T Commit<T>(Func<IAtomicTransaction, long, T> compile);
2526
void SetDispatchPaused(bool paused);
2627
long CreateBackup(string directory);
28+
StorageSnapshot CreateSnapshot(string path, long? expectedAppliedPosition = null);
29+
StorageSnapshot VerifySnapshot(string path);
30+
StorageSnapshot InstallSnapshot(string path, long expectedAppliedPosition);
31+
StorageSnapshot Compact();
2732
}
2833
public static class StorageRecords
2934
{

‎src/KeyLoad.AppHost/Program.cs‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -29,8 +29,10 @@
2929
.WithEnvironment("KeyLoad__Incarnation", profile.Incarnation.ToString())
3030
.WithEnvironment("KeyLoad__SigningKey", signing).WithEnvironment("KeyLoad__PeerSecret", peerSecret)
3131
.WithEnvironment("KeyLoad__AdminKey", admin).WithEnvironment("KeyLoad__AllowLoopbackHttp", "true")
32+
.WithEnvironment("KeyLoad__SnapshotThreshold", builder.Configuration["KeyLoad:SnapshotThreshold"] ?? "1024")
3233
.WithEnvironment("Logging__LogLevel__Default", "Warning")
3334
.WithEnvironment("Logging__LogLevel__DotNext.Net.Cluster.Consensus.Raft", "Information")
35+
.WithEnvironment("Logging__LogLevel__KeyLoad.Replication.PeerSecurity", "Information")
3436
.WithHttpHealthCheck("/health/ready", endpointName: "http")).ToArray();
3537
foreach (var resource in nodes)
3638
{

‎src/KeyLoad.Cli/Program.cs‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -12,6 +12,10 @@
1212
case "backup" when args.Length == 3:
1313
using (var store = new ZoneTreeStore(new(Path.GetFullPath(args[1])))) store.CreateBackup(Path.GetFullPath(args[2]));
1414
Console.WriteLine("Backup verified at creation."); break;
15+
case "compact" when args.Length == 2:
16+
using (var store = new ZoneTreeStore(new(Path.GetFullPath(args[1]))))
17+
Console.WriteLine(JsonSerializer.Serialize(store.Compact(), JsonDefaults.Options));
18+
break;
1519
case "restore" when args.Length == 3:
1620
var identity = ZoneTreeStore.Restore(Path.GetFullPath(args[1]), Path.GetFullPath(args[2]));
1721
Console.WriteLine(JsonSerializer.Serialize(new { identity.NodeId, identity.Incarnation, identity.DispatchPaused }, JsonDefaults.Options)); break;
@@ -46,6 +50,7 @@ static void Help() => Console.WriteLine("""
4650
KeyLoad CLI
4751
status <node-url> [local-profile.json]
4852
backup <offline-database-directory> <empty-backup-directory>
53+
compact <offline-database-directory>
4954
restore <backup-directory> <empty-database-directory>
5055
pack-backup <backup-directory> <new-artifact-file>
5156
inspect-artifact <artifact-file>

0 commit comments

Comments
 (0)