Skip to content

Commit 15ea503

Browse files
committed
Group benchmark matrices by database
Keep nine independent database matrices, sequential isolated-cell steps, complete aggregation and authenticated historical result names. Static actionlint, YAML/matrix inventory, JavaScript syntax and governance checks passed. Tests, benchmarks and workflow dispatch were not run at the owner request.
1 parent eca5132 commit 15ea503

27 files changed

Lines changed: 479 additions & 191 deletions

‎.github/workflows/AGENTS.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -63,3 +63,6 @@
6363

6464
## Independent website queue, 2026-10-04
6565
- CI website generation MUST NOT wait for unrelated ordinary CI/RF3 execution from another run. Push/manual CI runs use their own run identity; PR keeps its existing ref-based cancellation. Website qualification and deployment use separate bounded job concurrency groups with cancel-in-progress=false, preserving source/latest-evidence freshness and all gates. This implements the owner-authorized independent website action without canceling database or test work.
66+
67+
## Separate database matrices, 2026-10-04
68+
- Owner correction requires one named job matrix per database in Benchmarks. Each target has its own three checks and thirty canonical node/scenario workloads with readable database/node/scenario names. All nine groups depend only on actual plan/image inputs and run independently without a max-parallel cap; steps within each isolated cell remain sequential. Aggregate joins every group and preserves authenticated failed/null results. This supersedes the shared preflight/CRUD/specialized grouping above without changing native topology, cell isolation, required suites or immutable evidence.

‎.github/workflows/benchmarks.yml‎

Lines changed: 86 additions & 122 deletions
Original file line numberDiff line numberDiff line change
@@ -40,9 +40,7 @@ jobs:
4040
runs-on: ubuntu-latest
4141
timeout-minutes: 10
4242
outputs:
43-
crud: ${{ steps.plan.outputs.crud_matrix }}
44-
specialized: ${{ steps.plan.outputs.specialized_matrix }}
45-
preflight: ${{ steps.plan.outputs.preflight_matrix }}
43+
databases: ${{ steps.plan.outputs.database_matrices }}
4644
steps:
4745
- name: Download source code
4846
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
@@ -184,24 +182,24 @@ jobs:
184182
artifacts/code-quality/**
185183
if-no-files-found: error
186184
retention-days: 90
187-
comparison-preflight:
188-
name: Check database / ${{ matrix.id }}
185+
comparison-keyload:
186+
name: KeyLoad / ${{ matrix.label }}
189187
needs: [comparison-plan, comparison-images]
190188
runs-on: ubuntu-latest
191189
timeout-minutes: 150
192190
permissions: {contents: read, actions: read}
193191
strategy:
194192
fail-fast: false
195-
matrix: ${{ fromJSON(needs.comparison-plan.outputs.preflight) }}
196-
env:
193+
matrix: ${{ fromJSON(needs.comparison-plan.outputs.databases).keyload }}
194+
env: &database-environment
197195
GH_TOKEN: ${{ github.token }}
198196
KEYLOAD_COMPARISON_CELL_ID: ${{ matrix.id }}
199-
KEYLOAD_COMPARISON_JOB_NAME: Check database / ${{ matrix.id }}
197+
KEYLOAD_COMPARISON_JOB_NAME: ${{ matrix.jobName }}
200198
Benchmarks__Target: ${{ matrix.target }}
201199
Benchmarks__NodeCount: ${{ matrix.nodeCount }}
202200
Benchmarks__Scenario: ${{ matrix.scenario }}
203201
Benchmarks__EvidenceProfile: ${{ matrix.profile }}
204-
steps:
202+
steps: &database-steps
205203
- name: Download source code
206204
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
207205
- name: Prepare database containers
@@ -235,7 +233,7 @@ jobs:
235233
if: ${{ !cancelled() }}
236234
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
237235
with:
238-
name: comparison-preflight-${{ matrix.id }}
236+
name: ${{ matrix.artifactPrefix }}${{ matrix.id }}
239237
path: artifacts/comparisons/isolated/workers/${{ matrix.id }}/*
240238
if-no-files-found: error
241239
retention-days: 90
@@ -244,132 +242,98 @@ jobs:
244242
uses: ./.github/workflows/Features/BenchmarkComparisons/IsolatedCellTeardown
245243
with:
246244
setup-outcome: ${{ steps.images.outcome }}
247-
artifact-name: comparison-preflight-qualification-${{ matrix.id }}
248-
comparison-crud:
249-
name: Benchmark / ${{ matrix.id }}
245+
artifact-name: ${{ matrix.qualificationPrefix }}${{ matrix.id }}
246+
comparison-postgresql:
247+
name: PostgreSQL + pgvector / ${{ matrix.label }}
250248
needs: [comparison-plan, comparison-images]
251249
runs-on: ubuntu-latest
252250
timeout-minutes: 150
253251
permissions: {contents: read, actions: read}
254252
strategy:
255253
fail-fast: false
256-
matrix: ${{ fromJSON(needs.comparison-plan.outputs.crud) }}
257-
env:
258-
GH_TOKEN: ${{ github.token }}
259-
KEYLOAD_COMPARISON_CELL_ID: ${{ matrix.id }}
260-
KEYLOAD_COMPARISON_JOB_NAME: Benchmark / ${{ matrix.id }}
261-
Benchmarks__Target: ${{ matrix.target }}
262-
Benchmarks__NodeCount: ${{ matrix.nodeCount }}
263-
Benchmarks__Scenario: ${{ matrix.scenario }}
264-
Benchmarks__EvidenceProfile: ${{ matrix.profile }}
265-
steps:
266-
- name: Download source code
267-
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
268-
- name: Prepare database containers
269-
id: images
270-
uses: ./.github/workflows/Features/BenchmarkComparisons/IsolatedCellSetup
271-
with:
272-
bundle-artifact-id: ${{ needs.comparison-images.outputs.bundle-artifact-id }}
273-
- name: Restore .NET packages
274-
run: dotnet restore KeyLoad.slnx
275-
- name: Build benchmark tests
276-
id: benchmark-build
277-
run: dotnet build tests/KeyLoad.ComparisonTests --no-restore --configuration Release
278-
- name: Run database workload
279-
id: workload
280-
if: ${{ !cancelled() }}
281-
env:
282-
PREPARATION_OUTCOME: ${{ steps.images.outcome }}
283-
BUILD_OUTCOME: ${{ steps.benchmark-build.outcome }}
284-
run: |
285-
if [ "$PREPARATION_OUTCOME" != success ] || [ "$BUILD_OUTCOME" != success ]; then
286-
echo '::error::Benchmark preparation failed; no measurement data is available.'
287-
exit 1
288-
fi
289-
dotnet run --project src/KeyLoad.AppHost --no-build --no-restore --configuration Release -- --KeyLoadTests:Suite=comparison '--KeyLoadTests:Filter=/*/*/IsolatedNativeComparisonTests/*' --KeyLoadTests:TimeoutMinutes=140
290-
- name: Record benchmark availability
291-
if: ${{ !cancelled() }}
292-
env:
293-
KEYLOAD_WORKLOAD_OUTCOME: ${{ steps.workload.outcome }}
294-
run: node scripts/Features/BenchmarkComparisons/finalize-worker.mjs
295-
- name: Save benchmark results
296-
if: ${{ !cancelled() }}
297-
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
298-
with:
299-
name: comparison-worker-${{ matrix.id }}
300-
path: artifacts/comparisons/isolated/workers/${{ matrix.id }}/*
301-
if-no-files-found: error
302-
retention-days: 90
303-
- name: Clean up containers and save logs
304-
if: always()
305-
uses: ./.github/workflows/Features/BenchmarkComparisons/IsolatedCellTeardown
306-
with:
307-
setup-outcome: ${{ steps.images.outcome }}
308-
artifact-name: comparison-case-qualification-${{ matrix.id }}
309-
comparison-specialized:
310-
name: Benchmark / ${{ matrix.id }}
254+
matrix: ${{ fromJSON(needs.comparison-plan.outputs.databases).postgresql }}
255+
env: *database-environment
256+
steps: *database-steps
257+
comparison-qdrant:
258+
name: Qdrant / ${{ matrix.label }}
311259
needs: [comparison-plan, comparison-images]
312260
runs-on: ubuntu-latest
313261
timeout-minutes: 150
314262
permissions: {contents: read, actions: read}
315263
strategy:
316264
fail-fast: false
317-
matrix: ${{ fromJSON(needs.comparison-plan.outputs.specialized) }}
318-
env:
319-
GH_TOKEN: ${{ github.token }}
320-
KEYLOAD_COMPARISON_CELL_ID: ${{ matrix.id }}
321-
KEYLOAD_COMPARISON_JOB_NAME: Benchmark / ${{ matrix.id }}
322-
Benchmarks__Target: ${{ matrix.target }}
323-
Benchmarks__NodeCount: ${{ matrix.nodeCount }}
324-
Benchmarks__Scenario: ${{ matrix.scenario }}
325-
Benchmarks__EvidenceProfile: ${{ matrix.profile }}
326-
steps:
327-
- name: Download source code
328-
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1
329-
- name: Prepare database containers
330-
id: images
331-
uses: ./.github/workflows/Features/BenchmarkComparisons/IsolatedCellSetup
332-
with:
333-
bundle-artifact-id: ${{ needs.comparison-images.outputs.bundle-artifact-id }}
334-
- name: Restore .NET packages
335-
run: dotnet restore KeyLoad.slnx
336-
- name: Build benchmark tests
337-
id: benchmark-build
338-
run: dotnet build tests/KeyLoad.ComparisonTests --no-restore --configuration Release
339-
- name: Run database workload
340-
id: workload
341-
if: ${{ !cancelled() }}
342-
env:
343-
PREPARATION_OUTCOME: ${{ steps.images.outcome }}
344-
BUILD_OUTCOME: ${{ steps.benchmark-build.outcome }}
345-
run: |
346-
if [ "$PREPARATION_OUTCOME" != success ] || [ "$BUILD_OUTCOME" != success ]; then
347-
echo '::error::Benchmark preparation failed; no measurement data is available.'
348-
exit 1
349-
fi
350-
dotnet run --project src/KeyLoad.AppHost --no-build --no-restore --configuration Release -- --KeyLoadTests:Suite=comparison '--KeyLoadTests:Filter=/*/*/IsolatedNativeComparisonTests/*' --KeyLoadTests:TimeoutMinutes=140
351-
- name: Record benchmark availability
352-
if: ${{ !cancelled() }}
353-
env:
354-
KEYLOAD_WORKLOAD_OUTCOME: ${{ steps.workload.outcome }}
355-
run: node scripts/Features/BenchmarkComparisons/finalize-worker.mjs
356-
- name: Save benchmark results
357-
if: ${{ !cancelled() }}
358-
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
359-
with:
360-
name: comparison-worker-${{ matrix.id }}
361-
path: artifacts/comparisons/isolated/workers/${{ matrix.id }}/*
362-
if-no-files-found: error
363-
retention-days: 90
364-
- name: Clean up containers and save logs
365-
if: always()
366-
uses: ./.github/workflows/Features/BenchmarkComparisons/IsolatedCellTeardown
367-
with:
368-
setup-outcome: ${{ steps.images.outcome }}
369-
artifact-name: comparison-case-qualification-${{ matrix.id }}
265+
matrix: ${{ fromJSON(needs.comparison-plan.outputs.databases).qdrant }}
266+
env: *database-environment
267+
steps: *database-steps
268+
comparison-rabbitmq:
269+
name: RabbitMQ / ${{ matrix.label }}
270+
needs: [comparison-plan, comparison-images]
271+
runs-on: ubuntu-latest
272+
timeout-minutes: 150
273+
permissions: {contents: read, actions: read}
274+
strategy:
275+
fail-fast: false
276+
matrix: ${{ fromJSON(needs.comparison-plan.outputs.databases).rabbitmq }}
277+
env: *database-environment
278+
steps: *database-steps
279+
comparison-redis:
280+
name: Redis / ${{ matrix.label }}
281+
needs: [comparison-plan, comparison-images]
282+
runs-on: ubuntu-latest
283+
timeout-minutes: 150
284+
permissions: {contents: read, actions: read}
285+
strategy:
286+
fail-fast: false
287+
matrix: ${{ fromJSON(needs.comparison-plan.outputs.databases).redis }}
288+
env: *database-environment
289+
steps: *database-steps
290+
comparison-neo4j:
291+
name: Neo4j / ${{ matrix.label }}
292+
needs: [comparison-plan, comparison-images]
293+
runs-on: ubuntu-latest
294+
timeout-minutes: 150
295+
permissions: {contents: read, actions: read}
296+
strategy:
297+
fail-fast: false
298+
matrix: ${{ fromJSON(needs.comparison-plan.outputs.databases).neo4j }}
299+
env: *database-environment
300+
steps: *database-steps
301+
comparison-mongodb:
302+
name: MongoDB / ${{ matrix.label }}
303+
needs: [comparison-plan, comparison-images]
304+
runs-on: ubuntu-latest
305+
timeout-minutes: 150
306+
permissions: {contents: read, actions: read}
307+
strategy:
308+
fail-fast: false
309+
matrix: ${{ fromJSON(needs.comparison-plan.outputs.databases).mongodb }}
310+
env: *database-environment
311+
steps: *database-steps
312+
comparison-opensearch:
313+
name: OpenSearch / ${{ matrix.label }}
314+
needs: [comparison-plan, comparison-images]
315+
runs-on: ubuntu-latest
316+
timeout-minutes: 150
317+
permissions: {contents: read, actions: read}
318+
strategy:
319+
fail-fast: false
320+
matrix: ${{ fromJSON(needs.comparison-plan.outputs.databases).opensearch }}
321+
env: *database-environment
322+
steps: *database-steps
323+
comparison-kurrentdb:
324+
name: KurrentDB / ${{ matrix.label }}
325+
needs: [comparison-plan, comparison-images]
326+
runs-on: ubuntu-latest
327+
timeout-minutes: 150
328+
permissions: {contents: read, actions: read}
329+
strategy:
330+
fail-fast: false
331+
matrix: ${{ fromJSON(needs.comparison-plan.outputs.databases).kurrentdb }}
332+
env: *database-environment
333+
steps: *database-steps
370334
comparison-aggregate:
371335
name: Combine benchmark results
372-
needs: [comparison-build, comparison-plan, comparison-images, comparison-preflight, comparison-crud, comparison-specialized]
336+
needs: [comparison-build, comparison-plan, comparison-images, comparison-keyload, comparison-postgresql, comparison-qdrant, comparison-rabbitmq, comparison-redis, comparison-neo4j, comparison-mongodb, comparison-opensearch, comparison-kurrentdb]
373337
if: ${{ always() && !cancelled() && needs.comparison-plan.result == 'success' && needs.comparison-images.result == 'success' }}
374338
runs-on: ubuntu-latest
375339
timeout-minutes: 150

‎AGENTS.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -43,6 +43,7 @@ ManagedCode packages are our projects. Fix dependency defects in their owning si
4343
- Performance qualification MUST cover actual datasets of 100,000, 1,000,000 and 5,000,000 records, with at least 100,000 measured operations per applicable workload cell, sequential and deterministic random reads, ordered/range search, indexing and complex queries. Record dataset size separately from BDN iteration/sample counts, validate actual loaded records and caller-visible results, and retain equal topology, durability, resource and workload contracts across supported competitors. Tiny hot-key fixtures remain explicitly labelled microbenchmark controls and MUST NOT substitute for this scale evidence or justify selecting a performance winner. Repeating more reads over the same tiny corpus does not establish performance at the required record counts. When the owner requests serious performance qualification, complete the representative long workloads rather than presenting short controls as the result; local experiments remain development-only and website figures require authenticated original GitHub artifacts (owner repeated correction 2026-10-03).
4444
- Qualification MUST use Linux runners only; the owner explicitly replaced the three-OS qualification matrix on 2026-10-03. Preserve every required suite, scalar portability check, recovery and RF3 gate while removing redundant macOS/Windows execution.
4545
- Each comparison database MUST run the same complete workload suite in its own isolated GitHub Actions runner/job agent, with only that database's native topology and load generator. KeyLoad RF3 and competitor clusters remain genuine clusters; different databases MUST NOT share a runner, process, containers, volumes or measurement session (owner direction 2026-10-03).
46+
- Owner correction 2026-10-04 requires a separate named GitHub Actions job group and matrix for each comparison database. Keep each database's checks and workload cells together under its readable database name; do not flatten all databases into shared CRUD/specialized matrix groups. Preserve isolated runners per cell, the complete canonical inventory and one authenticated aggregation after every database group finishes.
4647
- Each isolated comparison agent MUST retain its own source/run/attempt/target/topology/profile-bound JSON. After all target agents finish, one aggregation job MUST validate completeness, comparable settings and provenance, collect their results and generate the website metrics from those JSON files. Failed, missing, skipped or mixed-cohort measurements MUST NOT refresh published performance evidence (owner direction 2026-10-03).
4748
- The comparison matrix MUST include actual native one-node, two-node and three-node configurations and intensive read/create/update/delete workloads. Each engine/node-count/scenario measurement MUST run on a separate isolated runner agent; record real membership, acknowledgement, correctness, latency, throughput and resource use. Do not relabel client counts or independent standalone databases as cluster node counts. Unsupported native/community topology remains explicitly unavailable, and benchmark topology changes require a fault/consistency ADR before implementation; the initial production RF3 requirement remains mandatory (owner direction 2026-10-03).
4849

‎docs/ADR/ADR-080-benchmark-failure-isolation.md‎

Lines changed: 29 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -294,3 +294,32 @@ serialization, and adds separate qualification/deployment job groups. Website
294294
work starts independently while all ordinary suites remain mandatory. Root owns
295295
this CI policy/YAML/regression join; unchanged source/latest-evidence freshness
296296
and false cancellation protect publication, without canceling existing tests.
297+
298+
## Database groups, 2026-10-04
299+
300+
REQ/AC-BC-GROUP-001/002 replace the shared preflight/CRUD/specialized display groups
301+
with nine database matrices derived from the unchanged canonical plan. Each matrix
302+
has 3 checks and 30 workloads; all 297 isolated Linux jobs retain sequential
303+
prepare/build/run/finalize/upload/cleanup steps. All database groups fan out from
304+
plan/images concurrently, and aggregate joins every database group even on workload
305+
failure. YAML anchors share step definitions while leaving native top-level GitHub
306+
step conclusions visible. Exactly three workflows remain.
307+
308+
Ordered delivery: root updates requirements/policy; root edits planner, workflow
309+
and exact name validators; a worker updates disjoint source-contract TUnit
310+
assertions; root joins and statically checks YAML/syntax/full inventory/diff, then
311+
commits the scoped stage on current main. The owner explicitly excludes executing
312+
tests, benchmarks or workflow dispatch for this Actions-only task. This evidence
313+
exception does not assert runtime qualification. Original authenticated historical
314+
job names remain accepted as exact cell contracts, preserving already produced JSON.
315+
No data/dependency/topology migration; rollback restores workflow/planner/name
316+
contracts together without modifying immutable evidence. Agent coordination is
317+
limited to static contract review and disjoint test-source updates.
318+
319+
```mermaid
320+
flowchart LR
321+
Inputs[Plan and native images] --> KeyLoad[KeyLoad matrix]
322+
Inputs --> Databases[Eight separate competitor matrices]
323+
KeyLoad --> Aggregate[Complete authenticated aggregation]
324+
Databases --> Aggregate
325+
```

‎docs/Features/BenchmarkComparisons.md‎

Lines changed: 24 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -974,3 +974,27 @@ flowchart LR
974974
Failure --> Aggregate
975975
Aggregate --> Site[Qualified website]
976976
```
977+
978+
## Database job groups, 2026-10-04
979+
980+
REQ-BC-GROUP-001 maps to AC-BC-GROUP-001: Benchmarks exposes nine separately
981+
named database job matrices. Each contains only its own target, three preflight
982+
checks and all 30 canonical node/scenario workloads. Every cell retains its own
983+
Linux runner. Database groups run concurrently with only plan/image dependencies;
984+
steps inside each cell run in order. No shared cross-database matrix or parallelism
985+
cap. Aggregate waits for every group and retains failed/null results.
986+
987+
REQ-BC-GROUP-002 maps to AC-BC-GROUP-002: readable database/node/scenario job names
988+
must agree with authenticated job discovery, finalization, aggregation and website
989+
receipt validation. Canonical plan schema, cell IDs, workload counts, artifacts,
990+
measurements and topology remain unchanged. Frozen historical names remain valid
991+
only as exact cell-name contracts in original authenticated receipts; no result
992+
bytes are rewritten. ADR-080 owns the implementation contract.
993+
994+
TASK-BC-GROUP-001: root owns workflow, planner/name validators and documentation;
995+
worker owns disjoint source-contract TUnit assertions. Update mapped
996+
WorkflowLayout/WorkflowBenchmarkFailure/NativeSerializationWorkflow/IsolatedPlanCli
997+
assertions before joining. Owner explicitly requests no tests or benchmarks for
998+
this Actions layout task: verify YAML, JavaScript syntax, matrix inventory and diff
999+
statically; do not dispatch workflows. Runtime/native qualification is unverified.
1000+
Backend/API/storage/transport and database tests are outside this change.

0 commit comments

Comments
 (0)