Skip to content

Commit 65e8bf5

Browse files
committed
Publish authenticated benchmark failures without blocking competitors
1 parent a7ca35d commit 65e8bf5

59 files changed

Lines changed: 1885 additions & 67 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.github/workflows/AGENTS.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -52,3 +52,6 @@
5252
- The latest owner correction places the actual pinned TimeSeries image test and retained facts in common image preparation, with no standalone feature branch. Keep all existing workload checks and isolated native measurements; staged intensive TimeSeries measurements remain explicitly pending.
5353
- Build/checks, plan and images run independently; preflight, CRUD and specialized matrices share only real plan/image inputs, never each other's completion or an arbitrary max-parallel cap. Aggregate explicitly waits for all six preparation/check/matrix jobs before website checks/publication.
5454
- Every authored job/step in the three workflows and used composites has a concrete readable name, including checkout, SDK setup, restore/build/test and uploads. Change displayed-name validators/oracles atomically; preserve stable internal job IDs, authenticated artifacts and all historical report bytes. Explicit owner direction supersedes earlier frozen displayed labels, without relaxing any identity, workload, topology, permission, coverage or publication gate.
55+
56+
## Owner-directed failed-cell publication, 2026-10-04
57+
- ADR-080 and the latest explicit owner instruction supersede the all-success cohort requirement: finish every independent workload and publish authenticated successful measurements while failed workloads retain null reports and actual failed job/workload conclusions. Keep complete planned-cell artifact accounting, successful image authority, full site qualification, exact source/run/attempt identity and least privileges. Unrelated KeyLoad build/runtime failures do not by themselves skip aggregation or website jobs; missing/corrupt evidence and failed site gates still block publication.

‎.github/workflows/Features/BenchmarkComparisons/IsolatedCellTeardown/AGENTS.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -17,3 +17,6 @@
1717

1818
## Owner-directed comparison pipeline, 2026-10-03
1919
- ADR-062 moves the exact-SHA native comparison cells from the historical ci.yml placement above to benchmarks.yml (`Benchmarks`). Every existing cleanup, retention, isolation and publication-failure requirement remains mandatory.
20+
21+
## Owner-directed failed-cell publication, 2026-10-04
22+
- Retain original failed runner reports separately from the null-report worker envelope, plus the finalization job capture. Workload failure may be published as unavailable under ADR-080; teardown failure itself remains a failed ownership/cleanup gate and cannot validate a measured cell.

‎.github/workflows/Features/BenchmarkComparisons/IsolatedCellTeardown/action.yml‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -24,5 +24,7 @@ runs:
2424
${{ runner.temp }}/keyload-cell-github/**
2525
**/TestResults/**
2626
artifacts/code-quality/**
27+
artifacts/comparisons/isolated/failures/**
28+
${{ runner.temp }}/keyload-cell-github-finalize/**
2729
if-no-files-found: error
2830
retention-days: 90

‎.github/workflows/Features/BenchmarkComparisons/QualifySite/AGENTS.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -13,3 +13,6 @@
1313

1414
## Applicable skills
1515
- No installed skill is required for this bounded workflow composition; install none.
16+
17+
## Owner-directed failed-cell publication, 2026-10-04
18+
- ADR-080 allows authenticated failed workload/null-report cells; keep every existing archive, source, TUnit, coverage, browser and freshness gate. The trusted producer/builder dependency closure now includes failed-cell finalization and its actual imports; source-bound counts change together with validators.

‎.github/workflows/Features/BenchmarkComparisons/QualifySite/action.yml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -60,7 +60,7 @@ runs:
6060
git -C measured-isolated rev-parse HEAD > "$EVIDENCE_DIR/isolated-measured-revision.txt"
6161
closure="scripts/Features/BenchmarkComparisons/site-isolated-dependencies.txt"
6262
cmp "control/$closure" "website/$closure"
63-
[[ "$(wc -l < "control/$closure")" == 65 ]]
63+
[[ "$(wc -l < "control/$closure")" == 70 ]]
6464
LC_ALL=C sort -u "control/$closure" > "$EVIDENCE_DIR/isolated-dependencies.txt"
6565
cmp "control/$closure" "$EVIDENCE_DIR/isolated-dependencies.txt"
6666
sha256sum "website/$closure" >> "$EVIDENCE_DIR/executed-evidence-tools.sha256"

‎.github/workflows/benchmarks.yml‎

Lines changed: 55 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -89,6 +89,7 @@ jobs:
8989
- name: Restore .NET packages
9090
run: dotnet restore KeyLoad.slnx
9191
- name: Build benchmark tests
92+
id: benchmark-build
9293
run: dotnet build tests/KeyLoad.ComparisonTests --no-restore --configuration Release
9394
- name: Check TimescaleDB Docker image and startup tools
9495
env:
@@ -201,9 +202,25 @@ jobs:
201202
- name: Restore .NET packages
202203
run: dotnet restore KeyLoad.slnx
203204
- name: Build benchmark tests
205+
id: benchmark-build
204206
run: dotnet build tests/KeyLoad.ComparisonTests --no-restore --configuration Release
205207
- name: Run database workload
206-
run: dotnet run --project src/KeyLoad.AppHost --no-build --no-restore --configuration Release -- --KeyLoadTests:Suite=comparison '--KeyLoadTests:Filter=/*/*/IsolatedNativeComparisonTests/*' --KeyLoadTests:TimeoutMinutes=145
208+
id: workload
209+
if: always()
210+
env:
211+
PREPARATION_OUTCOME: ${{ steps.images.outcome }}
212+
BUILD_OUTCOME: ${{ steps.benchmark-build.outcome }}
213+
run: |
214+
if [ "$PREPARATION_OUTCOME" != success ] || [ "$BUILD_OUTCOME" != success ]; then
215+
echo '::error::Benchmark preparation failed; no measurement data is available.'
216+
exit 1
217+
fi
218+
dotnet run --project src/KeyLoad.AppHost --no-build --no-restore --configuration Release -- --KeyLoadTests:Suite=comparison '--KeyLoadTests:Filter=/*/*/IsolatedNativeComparisonTests/*' --KeyLoadTests:TimeoutMinutes=140
219+
- name: Record benchmark availability
220+
if: always()
221+
env:
222+
KEYLOAD_WORKLOAD_OUTCOME: ${{ steps.workload.outcome }}
223+
run: node scripts/Features/BenchmarkComparisons/finalize-worker.mjs
207224
- name: Save benchmark results
208225
if: always()
209226
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
@@ -246,9 +263,25 @@ jobs:
246263
- name: Restore .NET packages
247264
run: dotnet restore KeyLoad.slnx
248265
- name: Build benchmark tests
266+
id: benchmark-build
249267
run: dotnet build tests/KeyLoad.ComparisonTests --no-restore --configuration Release
250268
- name: Run database workload
251-
run: dotnet run --project src/KeyLoad.AppHost --no-build --no-restore --configuration Release -- --KeyLoadTests:Suite=comparison '--KeyLoadTests:Filter=/*/*/IsolatedNativeComparisonTests/*' --KeyLoadTests:TimeoutMinutes=145
269+
id: workload
270+
if: always()
271+
env:
272+
PREPARATION_OUTCOME: ${{ steps.images.outcome }}
273+
BUILD_OUTCOME: ${{ steps.benchmark-build.outcome }}
274+
run: |
275+
if [ "$PREPARATION_OUTCOME" != success ] || [ "$BUILD_OUTCOME" != success ]; then
276+
echo '::error::Benchmark preparation failed; no measurement data is available.'
277+
exit 1
278+
fi
279+
dotnet run --project src/KeyLoad.AppHost --no-build --no-restore --configuration Release -- --KeyLoadTests:Suite=comparison '--KeyLoadTests:Filter=/*/*/IsolatedNativeComparisonTests/*' --KeyLoadTests:TimeoutMinutes=140
280+
- name: Record benchmark availability
281+
if: always()
282+
env:
283+
KEYLOAD_WORKLOAD_OUTCOME: ${{ steps.workload.outcome }}
284+
run: node scripts/Features/BenchmarkComparisons/finalize-worker.mjs
252285
- name: Save benchmark results
253286
if: always()
254287
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
@@ -291,9 +324,25 @@ jobs:
291324
- name: Restore .NET packages
292325
run: dotnet restore KeyLoad.slnx
293326
- name: Build benchmark tests
327+
id: benchmark-build
294328
run: dotnet build tests/KeyLoad.ComparisonTests --no-restore --configuration Release
295329
- name: Run database workload
296-
run: dotnet run --project src/KeyLoad.AppHost --no-build --no-restore --configuration Release -- --KeyLoadTests:Suite=comparison '--KeyLoadTests:Filter=/*/*/IsolatedNativeComparisonTests/*' --KeyLoadTests:TimeoutMinutes=145
330+
id: workload
331+
if: always()
332+
env:
333+
PREPARATION_OUTCOME: ${{ steps.images.outcome }}
334+
BUILD_OUTCOME: ${{ steps.benchmark-build.outcome }}
335+
run: |
336+
if [ "$PREPARATION_OUTCOME" != success ] || [ "$BUILD_OUTCOME" != success ]; then
337+
echo '::error::Benchmark preparation failed; no measurement data is available.'
338+
exit 1
339+
fi
340+
dotnet run --project src/KeyLoad.AppHost --no-build --no-restore --configuration Release -- --KeyLoadTests:Suite=comparison '--KeyLoadTests:Filter=/*/*/IsolatedNativeComparisonTests/*' --KeyLoadTests:TimeoutMinutes=140
341+
- name: Record benchmark availability
342+
if: always()
343+
env:
344+
KEYLOAD_WORKLOAD_OUTCOME: ${{ steps.workload.outcome }}
345+
run: node scripts/Features/BenchmarkComparisons/finalize-worker.mjs
297346
- name: Save benchmark results
298347
if: always()
299348
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a
@@ -311,6 +360,7 @@ jobs:
311360
comparison-aggregate:
312361
name: Combine benchmark results
313362
needs: [comparison-build, comparison-plan, comparison-images, comparison-preflight, comparison-crud, comparison-specialized]
363+
if: ${{ always() && !cancelled() && needs.comparison-plan.result == 'success' && needs.comparison-images.result == 'success' }}
314364
runs-on: ubuntu-latest
315365
timeout-minutes: 150
316366
permissions: {contents: read, actions: read}
@@ -352,6 +402,7 @@ jobs:
352402
qualify:
353403
name: Check website
354404
needs: comparison-aggregate
405+
if: ${{ always() && !cancelled() && needs.comparison-aggregate.result == 'success' }}
355406
runs-on: ubuntu-latest
356407
timeout-minutes: 180
357408
permissions: {contents: read, actions: read}
@@ -374,7 +425,7 @@ jobs:
374425
deploy:
375426
name: Publish website
376427
needs: qualify
377-
if: needs.qualify.result == 'success' && needs.qualify.outputs.mode == 'publish'
428+
if: ${{ always() && !cancelled() && needs.qualify.result == 'success' && needs.qualify.outputs.mode == 'publish' }}
378429
runs-on: ubuntu-latest
379430
timeout-minutes: 90
380431
permissions: {contents: read, actions: read, pages: write, id-token: write}

‎AGENTS.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -401,6 +401,7 @@ For changes outside existing owner authorization, obtain direction before changi
401401
## Preferences
402402

403403
### Likes
404+
- Owner correction 2026-10-04 requires Benchmarks to finish every independent database workload and publish authenticated successful measurements even when another benchmark, including unfinished KeyLoad, fails. Record each failed workload as unavailable with no numeric data and its actual failure/job provenance; never invent results or label a failed job successful. This explicitly supersedes the all-success comparison-cohort publication restriction while preserving complete planned-cell accounting, source/run/attempt identity, native topology, fairness, site qualification and least-privilege deployment. Benchmark repair does not authorize taking over concurrent KeyLoad engine implementation.
404405
- Complete the coherent task's implementation before executing its runtime tests or benchmarks. Author the mapped tests with the code, then run the required validation after the implementation and self-review; do not let repeated intermediate verification replace implementation progress. This owner correction2026-10-03 supersedes test-first execution for the current work while preserving every final correctness, CI, fault and performance gate.
405406
- Each agent MUST follow the execution plan for its own user-assigned task and complete only that task; change its scope only when the owner explicitly redirects it (owner correction 2026-10-03).
406407
- Other chats' plans, progress, requests and concurrent changes MUST NOT expand an agent's task or redirect its implementation. Preserve their work and leave their tasks to their owners (owner correction 2026-10-03).
Lines changed: 56 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,56 @@
1+
# ADR-080: Isolate benchmark failures during publication
2+
3+
Status: Accepted; source implemented, delivered-source verification pending.
4+
Date: 2026-10-04. Related: REQ/AC-BC-FAIL-001..005, ADR-056/074/076.
5+
6+
## Decision
7+
8+
The owner requires all independent benchmarks to finish and publication to proceed
9+
when one database workload fails, including unfinished KeyLoad. Preserve every
10+
planned cell, actual native topology, bounded resources, workload contracts and
11+
source/run/attempt/job/artifact authentication. A failed workload emits a version4
12+
envelope with `disposition: failed`, a fixed safe reason and `report: null`. Its
13+
original job remains failed and workload step remains failure; result upload must
14+
succeed. Measured envelopes require successful workload jobs. Evidence mismatch,
15+
missing artifacts, mixed sources, failed image preparation or unsuccessful site
16+
qualification still fails publication. Partial success cannot establish a winner
17+
against an unavailable database. No engine implementation or persisted data changes.
18+
19+
The alternative of stopping all publication discards useful independent results.
20+
Treating failures as zero would invent comparative performance. Retaining raw
21+
exception text risks exposing credentials; use a fixed public reason and link the
22+
original GitHub job for diagnostics. Ordinary runner failures are finalized before
23+
upload; cancellation/timeouts that prevent artifacts remain explicit blockers.
24+
25+
## Implementation contract
26+
27+
1. Root records owner policy, feature acceptance, exact baseline and shared schema.
28+
2. Producer worker owns aggregation/GitHub selection modules and producer regression
29+
tests; preserve original step conclusions and require envelope/proof agreement.
30+
3. Site worker owns isolated browser contract/metadata/report/numeric/view modules
31+
and dedicated TUnit projection/oracle/browser regressions. Failed cells contain
32+
no numeric metrics and retain their failed job links.
33+
4. Root owns workflow always-finalization/aggregate condition, shared receipt
34+
validation, dependency/coverage inventory, source name tests and documentation.
35+
Diagnostic worker inspects native startup errors without changing KeyLoad engine.
36+
5. Root reviews all diffs, builds solution, runs formatter/governance and focused
37+
Aspire-owned suites, then checkpoints scoped changes on current main and pushes.
38+
6. Genuine Linux Benchmarks run qualifies all cells, aggregate, site coverage/browser
39+
and Pages publication. Local tests are development proof only.
40+
41+
Migration is additive to version4 dispositions; deploy producer/validators/site
42+
atomically. Rollback reverts this coherent change and restores the conservative
43+
publication gate, retaining immutable original artifacts. AC-BC-FAIL-001..005 map
44+
to automated and actual-provider evidence in the feature specification. Root alone
45+
owns integration and shared contract updates; workers never commit or push.
46+
47+
```mermaid
48+
flowchart TD
49+
Runner[Aspire native runner] -->|success| Measured[Measurement envelope]
50+
Runner -->|failure| Unavailable[Null report envelope]
51+
Provider[Original GitHub job and artifact] --> Validator[Strict agreement validator]
52+
Measured --> Validator
53+
Unavailable --> Validator
54+
Validator --> Site[Successful values and unavailable cells]
55+
Site --> Gates[Full site qualification and Pages]
56+
```

‎docs/ADR/README.md‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -151,3 +151,9 @@ peer fences. Homogeneous rollout, real prior-binary proof and qualification pend
151151
ZoneTree.FullTextSearch candidate generations under a node-local physical owner,
152152
exact canonical ranking, persisted-policy cuts and fail-closed publication.
153153
Source integration and qualification pending; acceleration is not claimed.
154+
155+
[ADR-080](ADR-080-benchmark-failure-isolation.md) accepts owner-directed independent
156+
benchmark failure handling: bounded transient GitHub GET retries, authenticated
157+
null-report failed cells and successful competitor publication. Full planned-cell
158+
accounting and complete site qualification remain required; delivered-source proof
159+
is pending.

‎docs/Architecture.md‎

Lines changed: 21 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -947,3 +947,24 @@ flowchart LR
947947
Build --> Publish[Immutable tag GHCR and GitHub Release]
948948
Gate --> Publish
949949
```
950+
951+
## Independent benchmark failure publication
952+
953+
[ADR-080](ADR/ADR-080-benchmark-failure-isolation.md) and
954+
[BenchmarkComparisons](Features/BenchmarkComparisons.md) preserve the full native
955+
planned-cell inventory while allowing terminal workload failures to coexist with
956+
successful measurements. The original GitHub job/step remains failed, its bounded
957+
worker envelope has a null report, and the site has no numeric result for that
958+
cell. The aggregate and site jobs explicitly wait for all matrices and run after
959+
failures; image authority, authenticated artifacts, fairness, full website
960+
qualification and freshness remain mandatory. KeyLoad engine repair is separate.
961+
962+
```mermaid
963+
flowchart LR
964+
Job[Independent native Aspire cell] --> Result[Measured or failed null report]
965+
GitHub[Actual job steps and immutable artifact] --> Proof[Same source run attempt validation]
966+
Result --> Proof
967+
Proof --> Aggregate[Complete planned-cell aggregate]
968+
Aggregate --> Qualification[Site tests browser coverage freshness]
969+
Qualification --> Site[Values or no data and original job links]
970+
```

0 commit comments

Comments
 (0)