You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 2bc7ccc
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: README.md
+9-6Lines changed: 9 additions & 6 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -8,9 +8,9 @@ The original load-testing prototype has been replaced. This repository implement
8
8
9
9
This is an early implementation of the clustered kernel. The default server topology has three persistent voting nodes and requires a majority for writes and strong reads. It needs no external database, Redis or message broker.
10
10
11
-
Implemented surfaces include document CRUD and field patches, partition-scoped unique/composite equality indexes, command deduplication, stream expected-revision append/read, scheduled work queues with fenced leases and retry/DLQ, atomic inbox completion, bounded property-graph traversal, ordered samples, exact vector search, BM25 and weighted reciprocal-rank fusion, a bounded read-only Q1 SQL dialect, API keys, field omission and field-use policies, verified backups, a .NET SDK and CLI.
11
+
Implemented surfaces include document CRUD and field patches, partition-scoped unique/composite equality indexes, command deduplication, stream expected-revision append/read, scheduled work queues with fenced leases and retry/DLQ, atomic inbox completion, bounded property-graph traversal, ordered samples, exact vector search, BM25 and weighted reciprocal-rank fusion, a bounded read-only Q1 SQL dialect, API keys, field omission and field-use policies, verified backups, native Raft snapshots and empty-replica catch-up, offline journal compaction, a .NET SDK and CLI.
12
12
13
-
The [104-task implementation tracker](docs/implementation/status.json) records the remaining work. Raft snapshot compaction, shard movement, distributed multi-shard query planning, managed HNSW, topics and subscription groups, CDC, retention, schema migrations, external backup stores and full release qualification remain under development. The current server uses one replicated physical shard which contains many independent atomic partitions.
13
+
The [104-task implementation tracker](docs/implementation/status.json) records the remaining work. Automatic canonical journal maintenance, large and interrupted snapshot transfer qualification, shard movement, distributed multi-shard query planning, managed HNSW, topics and subscription groups, CDC, retention, schema migrations, external backup stores and full release qualification remain under development. The current server uses one replicated physical shard which contains many independent atomic partitions.
14
14
15
15
The advertised profiles are `ProcessDurable` for embedded storage and `QuorumProcessDurable` for the cluster. The kernel has process-kill recovery tests and the cluster has real leader-loss and minority tests. Power-loss qualification, broader platform qualification and the 72-hour endurance gate are still required before advertising `LocalDurable`, `QuorumDurable` or production readiness.
16
16
@@ -31,7 +31,9 @@ dotnet run --project src/KeyLoad.Cli -- status \
The standalone server requires explicit cluster identity, voter endpoints, shared peer authentication secrets, an administrator credential and a silo port through `KeyLoad` configuration. Set `KeyLoad:SiloAddress` to each node's advertised IP address for a cluster across hosts. HTTPS is required for cluster endpoints; the AppHost explicitly enables loopback HTTP for development. Peer RPCs and forwarding requests are signed for their recipient, path, query and body, timestamped and replay checked. Public HTTP requests authenticate against the replicated credential catalog. Client-supplied roles are never trusted.
34
+
The standalone server requires explicit cluster identity, voter endpoints, shared peer authentication secrets, an administrator credential and a silo port through `KeyLoad` configuration. Set `KeyLoad:SiloAddress` to each node's advertised IP address for a cluster across hosts. HTTPS is required for cluster endpoints; the AppHost explicitly enables loopback HTTP for development. Peer RPCs and forwarding requests are signed for their recipient, path, query, body and Raft protocol headers, timestamped and replay checked. Large peer bodies use bounded temporary files. Public HTTP requests authenticate against the replicated credential catalog. Client-supplied roles are never trusted.
35
+
36
+
The default peer connection/RPC/request timeouts are 500/1500/2000 milliseconds, below the 4000–8000 millisecond election range. These values are configurable through the corresponding `KeyLoad` options; validation requires that ordering. Native snapshots are produced every `KeyLoad:SnapshotThreshold` committed entries (default 1024). Small snapshot catch-up is tested; large transfers still need qualification against the configured RPC deadline.
35
37
36
38
## .NET client
37
39
@@ -93,16 +95,17 @@ page.ThrowIfFail();
93
95
94
96
Q1 supports projections and aliases, scalar parameters, comparisons, `AND`/`OR`/`NOT`, `IN`, `IS NULL`, `IS MISSING`, `ORDER BY`, `LIMIT` and `EXPLAIN`. Identifiers containing dots need double quotes. `id` and `revision` refer to canonical document identity/revision. Parameters use `@name`. JSON numbers use the decimal scalar policy. SQL is read-only; unsupported statements fail explicitly.
95
97
96
-
Queries require a matching point/equality index or explicit `AllowFullScan`. Scans, parser depth, tokens, rows, bytes and execution time have budgets. Cursor tokens bind the principal, policy epoch, schema, queryand current storage cut; writes or policy changes can expire a cursor. Sensitive predicates and sorting require a field-use grant. Returned documents omit protected paths, including classified values nested in arrays.
98
+
Queries require a matching point/equality index or explicit `AllowFullScan`. Scans, parser depth, tokens, rows, bytes and execution time have budgets. Cursor tokens bind the principal, policy epoch, schema, query, node identity, read generation and current storage cut. Writes, policy changes or snapshot installation can expire a cursor; compaction preserves its cut. Sensitive predicates and sorting require a field-use grant. Returned documents omit protected paths, including classified values nested in arrays.
97
99
98
100
Search accepts typed vector spaces and explicit text/vector fields. Both branches use one authorized read cut. Exact vector scores and BM25 ranks are combined with weighted RRF using one-based ranks. The managed ANN and graph retrieval extensions are tracked separately.
99
101
100
102
## Backups and Cartograph
101
103
102
-
The canonical backup includes the checksummed redo journal, database identity and a SHA-256 manifest. Domain data, schemas, credentials, outcomes and inbox receipts are journaled together. ZoneTree files can be rebuilt from that canonical journal. Restore validates every manifest file, creates a new incarnation, resets consensus routing metadata and leaves queue dispatch paused.
104
+
The canonical backup includes the checksummed redo journal, database identity and a SHA-256 manifest. Domain data, schemas, credentials, outcomes and inbox receipts are journaled together. The journal can begin with a verified checkpoint followed by newer transaction frames. ZoneTree files can be rebuilt from that canonical history. Restore validates every manifest file, creates a new incarnation, resets consensus routing metadata and leaves queue dispatch paused.
103
105
104
106
```sh
105
107
# Stop the node before using the offline CLI.
108
+
dotnet run --project src/KeyLoad.Cli -- compact data/cluster/node1/database
106
109
dotnet run --project src/KeyLoad.Cli -- backup data/cluster/node1/database backups/snapshot
107
110
dotnet run --project src/KeyLoad.Cli -- restore backups/snapshot data/restored
108
111
@@ -140,7 +143,7 @@ dotnet test --project tests/KeyLoad.RecoveryTests --no-build --no-restore
140
143
dotnet test --project tests/KeyLoad.IntegrationTests --no-build --no-restore
141
144
```
142
145
143
-
Tests use xUnit and Microsoft.Testing.Platform. Recovery qualification runs 1000 seeded real-process kills, checks complete-frame corruption and verifies a clean backup restore. Integration tests own the Aspire lifecycle and use three independent server processes with isolated persistent directories. They kill the elected leader, retry the same command, verify atomic effects on surviving voters and reject a minority write. No manually running AppHost is needed for tests.
146
+
Tests use xUnit and Microsoft.Testing.Platform. Recovery qualification runs 1000 seeded real-process kills, checks complete-frame corruption and verifies a clean backup restore. Additional process kills cover checkpoint publication and all native Raft append paths immediately after acknowledgement. Integration tests own the Aspire lifecycle and use three independent server processes with isolated persistent directories. They kill the elected leader, retry the same command, verify atomic effects on surviving voters, reject a minority write and erase/restart one replica to require snapshot catch-up. No manually running AppHost is needed for tests.
144
147
145
148
Benchmarks can be started with `dotnet run -c Release --project benchmarks/KeyLoad.Benchmarks`. Comparative PostgreSQL/Marten/Wolverine and multi-node scaling qualification are part of the development plan; no comparative performance claims are made yet.
Copy file name to clipboardExpand all lines: docs/implementation/durability-audit.md
+6-2Lines changed: 6 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -10,9 +10,13 @@ The adapter compiles each operation under a write gate, writes a bounded frame,
10
10
11
11
The separate node ownership lock prevents concurrent writers. Returned bytes and staged mutations are copied. Identity and backup files are versioned; the backup manifest verifies every canonical file with SHA-256. A backup is made under the apply gate and reports that exact cut. Restore changes the incarnation, invalidates old signed tokens, resets the Raft/Orleans watermark and membership, and explicitly persists paused dispatch even if the old journal had unpaused it.
12
12
13
+
Format 2 adds a checkpoint prefix with ordered live records, per-frame SHA-256 checksums and a complete-image checksum/count. It records both the local journal position and the Raft apply position. Installation verifies the entire incoming image before closing the old generation, flushes the staged file and atomically replaces the canonical journal. The derived ZoneTree generation is rebuilt. Installation keeps the node identity, increments the read generation and invalidates existing query cursors; offline compaction preserves the logical cut and cursor generation. Process-kill regressions cover snapshot write/flush and the two publication boundaries, including reclamation after reopening.
14
+
13
15
## Replicated acknowledgement
14
16
15
-
The .NEXT 6.8.1 persistent Raft log owns terms, votes, log matching and consensus commit. `DurableRaftLog` reimplements all public append paths on the native persistent log and awaits `FlushAsync` before an AppendEntries acknowledgement can be returned. `FlushOnCommit` alone would not establish this follower append boundary. A regression invokes append through the actual `IPersistentState` interface and reopens an uncommitted tail.
17
+
The .NEXT 6.8.1 persistent Raft log owns terms, votes, log matching and consensus commit. `DurableRaftLog` requires native manual checkpoints (`FlushInterval = Timeout.InfiniteTimeSpan`), serializes its public append/commit paths and awaits the native foreground `FlushAsync` before acknowledgement. Background flush notifications do not establish this boundary for an isolated first append or snapshot installation. Four regressions invoke append through the actual `IPersistentState` interface, kill the process immediately after ACK and reopen the tail using private page memory; they do not rely on graceful disposal or shared memory pages surviving exit.
18
+
19
+
The state machine produces native Raft snapshots from a verified canonical store cut. A snapshot-only regression verifies that installation acknowledges without a subsequent append and reopens at that cut. The RF3 test erases one stopped follower's data directory, requires native snapshot installation, checks the fresh node identity and read generation and verifies documents and command outcomes through the HTTP API. Incoming bodies are authenticated and rewound for both `Body` and `BodyReader`; outgoing one-shot payloads are serialized once into a bounded disk spool. Raft term, snapshot metadata, request ID and content type are included in the peer signature.
16
20
17
21
The bounded leader writer replicates a trusted operation with one leader-chosen evaluation time. It returns success after consensus commit and local state-machine apply, then resolves the persisted outcome under current authorization. Cancellation or response loss after admission has an unknown write outcome; retries must retain the command ID. A minority cannot acknowledge writes.
18
22
@@ -26,4 +30,4 @@ Local tests run on macOS arm64 with .NET SDK 10.0.401. The recovery suite execut
26
30
27
31
The local result and crash-stage distribution are recorded in `kernel-qualification.json`. Each trial writes its seed, fault stage, mutation index, recovered values and platform to `artifacts/qualification/crash-trials-*.jsonl`; CI retains these as run artifacts.
28
32
29
-
The advertised profiles remain `ProcessDurable` and `QuorumProcessDurable`. These tests do not establish filesystem directory-entry persistence, storage-controller guarantees, real power-cut behavior or every operating system/filesystem combination. Linux/macOS/Windows CI, network faults and subsequent snapshot/compaction qualification are tracked separately. Power-loss qualification, network-partition coverage, long histories, shard movement and the 72-hour endurance gate remain required for the broader architecture's durable release profile.
33
+
The advertised profiles remain `ProcessDurable` and `QuorumProcessDurable`. These tests do not establish filesystem directory-entry persistence, storage-controller guarantees, real power-cut behavior or every operating system/filesystem combination. Linux/macOS/Windows CI and network faults are tracked separately. Native interrupted transfer/corruption recovery, large snapshots under transport deadlines, automatic canonical compaction, power-loss qualification, network partitions, long histories, shard movement and the 72-hour endurance gate remain required for the broader architecture's durable release profile.
0 commit comments