Skip to content

Commit 883a207

Browse files
committed
Isolate benchmark test execution and define shared load scenarios
1 parent 556c13a commit 883a207

20 files changed

Lines changed: 516 additions & 18 deletions

File tree

‎AGENTS.md‎

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -47,7 +47,7 @@ ManagedCode packages are our projects. Fix dependency defects in their owning si
4747
- Performance qualification MUST compare both single-node and native multi-node configurations against the complete competitor matrix with equivalent workload and acknowledgement/durability contracts. Retain exact-source measurements and regression evidence; the goal of leading competitors MUST NOT be reported as achieved without those results (owner direction 2026-10-02).
4848
- Performance qualification MUST cover actual datasets of 100,000 and 1,000,000 records, with at least 100,000 measured operations per applicable workload cell, sequential and deterministic random reads, ordered/range search, indexing and complex queries. Record dataset size separately from BDN iteration/sample counts, validate actual loaded records and caller-visible results, and retain equal topology, durability, resource and workload contracts across supported competitors. Tiny hot-key fixtures remain explicitly labelled microbenchmark controls and MUST NOT substitute for this scale evidence or justify selecting a performance winner. Repeating more reads over the same tiny corpus does not establish performance at the required record counts. When the owner requests serious performance qualification, complete the representative long workloads rather than presenting short controls as the result; local experiments remain development-only and website figures require authenticated original GitHub artifacts (owner repeated correction 2026-10-03).
4949
- Qualification MUST use Linux runners only; the owner explicitly replaced the three-OS qualification matrix on 2026-10-03. Preserve every required suite, scalar portability check, recovery and RF3 gate while removing redundant macOS/Windows execution.
50-
- Each comparison database MUST run the same complete workload suite in its own isolated GitHub Actions runner/job agent, with only that database's native topology and load generator. KeyLoad RF3 and competitor clusters remain genuine clusters; different databases MUST NOT share a runner, process, containers, volumes or measurement session (owner direction 2026-10-03).
50+
- Each comparison database MUST run the same complete workload suite in its own isolated GitHub Actions runner/job agent, with only that database's native topology and load generator. Owner reiteration2026-10-09 requires one canonical shared scenario inventory for all databases, including ingestion at1/10/500 clients: match actual record counts, corpus/seed/payload, operation order/count, concurrency, timing boundaries, correctness oracles and effective resource/acknowledgement/durability contracts. Target adapters MUST execute that shared scenario rather than substitute easier target-specific workloads; unsupported native capabilities remain explicit unavailable cells. KeyLoad RF3 and competitor clusters remain genuine clusters; different databases MUST NOT share a runner, process, containers, volumes or measurement session (owner direction 2026-10-03).
5151
- Owner correction 2026-10-04 requires a separate named GitHub Actions job group and matrix for each comparison database. Keep each database's checks and workload cells together under its readable database name; do not flatten all databases into shared CRUD/specialized matrix groups. Preserve isolated runners per cell, the complete canonical inventory and one authenticated aggregation after every database group finishes.
5252
- Each isolated comparison agent MUST retain its own source/run/attempt/target/topology/profile-bound JSON. After all target agents finish, one aggregation job MUST validate completeness, comparable settings and provenance, collect their results and generate the website metrics from those JSON files. Failed, missing, skipped or mixed-cohort measurements MUST NOT refresh published performance evidence (owner direction 2026-10-03).
5353
- The comparison matrix MUST include actual native one-node, two-node and three-node configurations and intensive read/create/update/delete workloads. Each engine/node-count/scenario measurement MUST run on a separate isolated runner agent; record real membership, acknowledgement, correctness, latency, throughput and resource use. Do not relabel client counts or independent standalone databases as cluster node counts. Unsupported native/community topology remains explicitly unavailable, and benchmark topology changes require a fault/consistency ADR before implementation; the initial production RF3 requirement remains mandatory (owner direction 2026-10-03).
@@ -574,3 +574,9 @@ A bounded website qualification candidate contains the20-project historical runt
574574
- Owner clarification 2026-10-09 permits task iterations and checkpoints to execute only the tests mapped to that task's acceptance scope, locally or through an explicitly selected GitHub task run. Preserve the full mandatory final qualification; a focused run's green result proves its selected task flows only and MUST NOT be reported as the complete solution or coverage gate.
575575
- Owner correction 2026-10-09 requires at least 20 native TUnit execution slots for independent functional tests, rather than substituting parallel CI jobs for test concurrency. Default the native selector and typed test execution options to 20 and remove forced one-test execution from task lanes. Preserve genuine shared-resource invariants with narrowly justified native scheduling; audit blanket serial attributes and prove actual concurrency from original native execution records rather than fabricated counters.
576576
- Owner authorization 2026-10-09 permits increasing native TUnit concurrency up to 50 during implementation when actual CPU/resource use and complete-flow outcomes support it. Start at 20, compare real execution duration and resource pressure, retain original failures, and choose the measured useful concurrency without changing acceptance, timeouts, shared-resource ownership or cleanup.
577+
578+
## Functional concurrency and isolated load measurements, owner clarification 2026-10-09
579+
580+
- The 20-slot default and measured increase up to 50 apply to independent ordinary functional tests. Benchmark measurements MUST execute one test/scenario at a time inside each job so unrelated tests do not contaminate timing, CPU, memory, storage or backlog measurements. Independent benchmark jobs MAY run concurrently only on genuinely isolated Linux runners with their own native database topology, client processes, containers, volumes and cleanup; preserve the complete comparison and provenance gates.
581+
- Heavy functional load/stress tests, including concurrent ingestion while other database operations are under load, MUST execute exclusively on their owned test resources without overlapping ordinary tests or other heavy cases in the same runner. Keep these correctness scenarios in KeyLoad functional acceptance, separately selected from ordinary parallel cases; they do not contribute performance measurements or coverage totals.
582+
- Ingestion load qualification MUST include a one-client baseline and concurrent-client cases at 10 and 500 clients. Each case adds exactly 1,000,000 total distinct records across its clients, verifies acknowledged writes and the final stored records through real clients, and records actual client concurrency separately from native TUnit test concurrency. Bound client admission, cancellation and cleanup; retain the existing 100,000/1,000,000 dataset inventory and matched comparison contracts.

‎README.md‎

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -371,7 +371,13 @@ Its [document contract](docs/Features/DocumentStorage.md) keeps the captured dat
371371
cut distinct from the current authorization cut. Complete current-source RF3
372372
qualification remains open.
373373

374-
Native TUnit defaults to 20 parallel tests. GitHub runs separate normal/scalar
374+
Ordinary native TUnit tests default to 20 parallel cases, with measured tuning up
375+
to 50. Benchmark tests run one at a time inside each isolated Linux job; separate
376+
jobs may run concurrently. Heavy ingestion and mixed-operation load tests require
377+
exclusive execution and stay outside coverage. The required million-record
378+
ingestion cases for 1, 10 and 500 clients are specified but remain unimplemented
379+
and unqualified. Every comparison target must execute the same canonical scenario
380+
set, datasets, client counts and measurement rules. GitHub runs separate normal/scalar
375381
acceptance lanes for storage ownership, strict indexes, the server/SDK,
376382
security/telemetry, strong/session reads, exact vectors and controlled partition
377383
movement. These seven tasks run in fourteen isolated Linux jobs. Focused task

‎benchmarks/KeyLoad.ComparisonHost/Features/BenchmarkComparisons/Dockerfile‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -13,4 +13,4 @@ FROM mcr.microsoft.com/dotnet/aspnet:10.0.12@sha256:222759b391a1aaf241166672c8f9
1313
WORKDIR /app
1414
COPY --from=build /app/publish/ ./
1515
USER $APP_UID
16-
ENTRYPOINT ["dotnet", "KeyLoad.ComparisonHost.dll", "--output", "Detailed", "--report-trx", "--results-directory", "/reports/TestResults"]
16+
ENTRYPOINT ["dotnet", "KeyLoad.ComparisonHost.dll", "--output", "Detailed", "--maximum-parallel-tests", "1", "--report-trx", "--results-directory", "/reports/TestResults"]

‎docs/ADR/ADR-056-isolated-linux-comparison-cells.md‎

Lines changed: 58 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -65,3 +65,61 @@ Local Aspire builds and test runs are permitted development evidence only. They
6565
This decision changes benchmark execution and evidence accounting, not database format or public operation contracts. Deploy producer, aggregate, Website consumer and workflow joins as one source-reviewed checkpoint. Publish metrics only from the complete eligible current cohort after all applicable gates and provider evidence pass. A source rollback restores the last qualified Website artifact and does not rewrite or relabel immutable historical measurements. Historical receipts keep their original source and settings outside the active plan; they are never fallback evidence for current metrics.
6666

6767
The ADR remains Accepted until the required implementation and exact-source Linux evidence are complete. The canonical feature documents and status records remain the authority for which gates have actually passed.
68+
69+
## Measurement scheduling and ingestion contract, 2026-10-09
70+
71+
Related requirements: REQ/AC-SCALE-024..028 and TASK-SCALE-MEASUREMENT-SCHEDULING-001,
72+
TASK-SCALE-INGESTION-001, TASK-SCALE-MIXED-LOAD-001 in
73+
[ScalingQualification](../Features/BenchmarkComparisons/ScalingQualification.md#functional-scheduling-and-ingestion-clarification-2026-10-09).
74+
[NativeTUnitEntry](../Features/TestInfrastructure/NativeTUnitEntry.md)
75+
REQ/AC-TUNIT-ENTRY-013/014 govern serial benchmark execution and exclusive heavy
76+
functional selection while preserving original outcomes and joined resource cleanup; [ADR-117](ADR-117-native-tunit-ci-entry.md)
77+
retains direct native TUnit invocation and fixture-owned Aspire lifetimes.
78+
79+
Independent functional tests use the ordinary 20-slot default, with a measured
80+
increase up to 50. Measurements execute one native test/scenario at a time inside
81+
each benchmark job. The measurement case may itself execute the declared number
82+
of workload clients concurrently: client concurrency never represents native
83+
TUnit scheduling. Independent benchmark jobs may execute concurrently only on
84+
separate Linux runners with their own native topology, clients, containers,
85+
volumes and joined cleanup. No arbitrary cross-job serialization is introduced.
86+
Heavy functional mixed-ingestion cases execute exclusively on owned RF3 resources,
87+
separately from ordinary cases and other heavy cases in the runner; they qualify
88+
correctness and contribute neither timing metrics nor coverage totals.
89+
90+
The new ingestion contract contains a one-client baseline and separate 10-client
91+
and 500-client cases. Each creates exactly 1,000,000 distinct total records, using
92+
real clients, deterministic disjoint identity allocation, bounded admission,
93+
original cancellation/drain/disposal and complete untimed independent stored-data
94+
validation. Retain genuine SDK receipts, official MCP interoperability, persisted
95+
authorization and RF3. Freeze actual startup/per-call/drain bounds and original
96+
failure handling in the typed policy before implementation; no million-task
97+
materialization, sampled final validation or unjoined shutdown is accepted.
98+
99+
Owner reiteration2026-10-09 requires one canonical shared scenario inventory for
100+
all comparison databases under REQ/AC-SCALE-029. Target adapters preserve identical
101+
datasets, seed/payload, operation schedules/counts, client concurrency, measured
102+
boundaries and correctness oracles with equivalent effective resources and
103+
acknowledgement/durability. Planner/aggregate joins reject missing, substituted or
104+
incomparable cells; unsupported native capabilities remain explicitly unavailable.
105+
The new ingestion inventory is shared across targets rather than a KeyLoad-only
106+
performance workload. It remains unimplemented/unqualified in this checkpoint.
107+
108+
Implementation order and ownership are the three task rows in ScalingQualification:
109+
root joins the native selector and existing typed execution options first;
110+
BenchmarkComparisons owners then freeze and implement feature-local ingestion
111+
contracts/execution/reporting with real ComparisonTests flows; DocumentStorage
112+
integration ownership specifies and implements bounded mixed-load cases through
113+
the existing ClusterFixture and SDK/MCP owners. Root integrates exclusive heavy
114+
selection, coverage exclusion and isolated Linux acceptance. No new resource
115+
harness, public API, dependency or database format is introduced.
116+
117+
Rollout changes scheduling only in this checkpoint. New load cases are **not
118+
implemented or qualified**; they enter dispatch only after their selector,
119+
report, inventory, correctness and provenance contracts are implemented and
120+
verified together. Preserve existing c16 profiles, 100K/1M inventory, native
121+
matrix identities and original immutable reports. Rollback of this scoped work
122+
may remove its scheduling/doc change or unqualified future routes; it never
123+
rewrites historical results or presents overlapping measurements as qualified.
124+
This ADR remains Accepted and adds no successful ingestion, load or publication
125+
evidence.

‎docs/ADR/ADR-117-native-tunit-ci-entry.md‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,12 @@
22

33
Status: Accepted; compilation and static verification passed, runtime qualification pending.
44

5+
## Isolated measurements and heavy load scheduling (2026-10-09)
6+
7+
Owner clarification binds REQ/AC-TUNIT-ENTRY-013/014 in [NativeTUnitEntry](../Features/TestInfrastructure/NativeTUnitEntry.md). Ordinary independent cases retain the native 20-slot default and measured tuning up to 50. Comparison selections require one native test at a time and reject greater concurrency before startup; concurrent workload clients remain independently configured. Independent benchmark jobs keep their genuinely isolated Linux native topology and resources. Heavy functional ingestion/mixed-operation scenarios require a separate exclusive selection and no coverage contribution.
8+
9+
TASK-TUNIT-ISOLATED-MEASUREMENTS-013 freezes this policy before source edits: scheduling worker updates the Node selector and typed AppHost selection with real selection regressions; root owns direct ComparisonHost container/process one-test arguments, docs, final review, canonical build/format and original native regression outcomes. Read-only inventory review verifies native measurement workers and workflow isolation. No client workload, deadline, database format, API or provider change belongs to this scheduling stage. Rollback removes the source delta only; prior qualification gates and this owner requirement remain. Native 1/10/500-client million-record and mixed-load complete-flow qualification remains open until implemented and executed.
10+
511
Owner correction 2026-10-07 supersedes ADR-074's outer AppHost test-runner entry. CI starts native TUnit after build. TUnit fixtures use DistributedApplicationTestingBuilder, await readiness, execute real C# SDK/official MCP/native comparison clients and dispose the owned Aspire applications. Unit and recovery suites do not acquire an unnecessary RF3 topology. Benchmarks remain in their separate pipeline. Native --output Detailed exposes original results during execution; no console-log file bridge or custom outcome counter is needed.
612

713
REQ-TUNIT-ENTRY-001: every CI suite and isolated benchmark workload invokes TUnit directly; infrastructure remains Aspire-owned within tests. AC-TUNIT-ENTRY-001: native command selections preserve project, filter, scalar environment, bounded parallelism, original TRX/coverage output and nonzero exit codes. AC-TUNIT-ENTRY-002: RF3 coverage preparation executes as a TUnit session lifecycle using the existing Aspire preparation resource, propagates its verified original manifest and policy, joins output and cleanup, and does not recursively launch a test runner. Existing real fixture and workload tests remain mandatory Linux acceptance evidence.

0 commit comments

Comments
 (0)