Skip to content

Commit ee7c7be

Browse files
committed
Select the latest completed benchmark results after queued cancellations
1 parent da643fb commit ee7c7be

10 files changed

Lines changed: 67 additions & 5 deletions

File tree

‎.github/workflows/AGENTS.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -57,3 +57,6 @@
5757

5858
## Owner-directed failed-cell publication, 2026-10-04
5959
- ADR-080 and the latest explicit owner instruction supersede the all-success cohort requirement: finish every independent workload and publish authenticated successful measurements while failed workloads retain null reports and actual failed job/workload conclusions. Keep complete planned-cell artifact accounting, successful image authority, full site qualification, exact source/run/attempt identity and least privileges. Unrelated KeyLoad build/runtime failures do not by themselves skip aggregation or website jobs; missing/corrupt evidence and failed site gates still block publication.
60+
61+
## Latest completed result eligibility, 2026-10-04
62+
- Latest benchmark metrics means the newest completed own-main push/manual Benchmarks producer with success/failure conclusion. Pending, skipped and canceled workflows have no completed comparison cohort and MUST NOT displace ready JSON. Exclude them before choosing the latest eligible producer; then reject its missing/corrupt/failed aggregate without older fallback. A canceled workflow_run event still cannot authorize publication. This refines the latest-result rule without accepting incomplete or fabricated measurements.

‎AGENTS.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -456,3 +456,6 @@ ADR-063 additionally introduces `src/KeyLoad.Diagnostics`, bringing the shared w
456456
ADR-077 additionally introduces `src/KeyLoad.Storage.IO`, bringing the shared workspace count to27. Its local policy precedes implementation. The internal StorageRecovery primitive uses public OS interfaces to reject non-regular migration inputs without blocking and preserves existing owner-lock interoperability. It owns no engine, database state or replication; all prior policies, immutable formats and qualification gates remain mandatory.
457457

458458
A bounded website qualification candidate contains the20-project historical runtime base plus KeyLoad.Analyzers, KeyLoad.Analyzers.Tests and KeyLoad.SiteTests (23projects). Its derived inventory MUST record the shared25-project scope and omitted concurrent projects explicitly, preserve all28 included root/local policy blobs exactly, and qualify only the actual included website/analyzer source. This is an evidence-scope record, never an exception to any solution-wide architecture, project-policy or required product qualification rule.
459+
460+
## Latest completed result eligibility, 2026-10-04
461+
- Latest benchmark metrics means the newest completed own-main push/manual Benchmarks producer with success/failure conclusion. Pending, skipped and canceled workflows have no completed comparison cohort and MUST NOT displace ready JSON. Exclude them before choosing the latest eligible producer; then reject its missing/corrupt/failed aggregate without older fallback. A canceled workflow_run event still cannot authorize publication. This refines the latest-result rule without accepting incomplete or fabricated measurements.

‎docs/ADR/ADR-080-benchmark-failure-isolation.md‎

Lines changed: 13 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -19,8 +19,9 @@ no dependency on ordinary CI/RF3. Preserve every archive/source/TUnit/browser/
1919
coverage gate and least-privilege needs-gated deployment.
2020

2121
Every website build selects the newest completed own-main Benchmarks run by run
22-
number across push/workflow_dispatch producers; pending runs cannot replace ready
23-
JSON. Require its success/failure conclusion and successful complete aggregate.
22+
number across push/workflow_dispatch producers with success/failure conclusion;
23+
pending, skipped and canceled workflows have no completed comparison cohort and
24+
are excluded before choosing the latest. Require its successful complete aggregate.
2425
After selecting the newest completed run, missing/corrupt/expired/failed aggregate
2526
evidence rejects publication without older fallback. Authenticate an actual
2627
workflow_run trigger separately: it may refer to an older completed run while a
@@ -273,3 +274,13 @@ flowchart TD
273274
Validator --> Site[Successful values and unavailable cells]
274275
Site --> Gates[Full site qualification and Pages]
275276
```
277+
278+
## Delivered selection correction
279+
280+
Original GitHub metadata after873cd1a shows runs55/54 canceled while waiting,
281+
run56 active and the genuine completed270-cell JSON at run53/73. Cancellation
282+
is not a completed comparison result. Filter candidates to completed success/
283+
failure producers before selecting highest run number; the actual canceled event
284+
still cannot authorize publication. Preserve no-fallback rejection after the
285+
newest eligible producer has a failed/missing/corrupt aggregate. The original
286+
metadata digest is f8816dacf8a60b68a41b8911185b8c97f2e27cbd816980f1331972b929d19007.

‎docs/Features/BenchmarkComparisons.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -718,7 +718,7 @@ KeyLoad engine repair and concurrent series-codec work are outside this task.
718718
| REQ-BC-FAIL-018 independent automatic JSON consumer | AC-BC-FAIL-018 CI website jobs run on own-main push/manual and actual completed Benchmarks events, independently of ordinary CI/RF3 and producer workload failures. Consume only a successful aggregate and its exact original suite/provider archives. Chrome/site/build/deploy remain exclusively CI website work and cannot affect the producer | Updated workflow/source/producer-selection TUnit tests, genuine workflow_run CI and published JSON/browser join |
719719
| REQ-BC-FAIL-019 authenticate the cross-workflow producer | AC-BC-FAIL-019 native Linux CI qualify/deploy context validates the real GitHub workflow_run payload: repository477801965, own-main head repository/branch, Benchmarks path/name, completed status, allowed producer event, immutable source, positive exact run/attempt and success/failure conclusion. Reject cancellation, fork, other workflow, changed tuple, missing/malformed/nonregular/oversized event file and historical publication overrides before capture. Authenticate that trigger separately from the latest selected producer and current CI executor/control/website; reselect latest and reauthenticate before deploy | Actual native context/selection rejection tests and CI AcPipe003; original archive receipts,270/277 identity and existing freshness coverage |
720720

721-
| REQ-BC-FAIL-020 latest metrics for every website build | AC-BC-FAIL-020 select the highest run-number completed own-main Benchmarks run across push/workflow_dispatch producers; pending runs do not replace finished JSON. Require success/failure producer conclusion, successful aggregate and all original270/277 proofs. Reject missing/corrupt/latest failed aggregation without older fallback. Own-main CI push/manual and workflow_run all use this selection; changed latest tuple before deploy prevents stale refresh | Real producer-selection positive/negative TUnit cases, native CI capture and freshness/provider join |
721+
| REQ-BC-FAIL-020 latest metrics for every website build | AC-BC-FAIL-020 select the highest run-number completed own-main Benchmarks run with success/failure conclusion across push/workflow_dispatch producers; pending, skipped and canceled runs have no finished comparison cohort and are not candidates. Require success/failure producer conclusion, successful aggregate and all original270/277 proofs. Reject missing/corrupt/latest failed aggregation without older fallback. Own-main CI push/manual and workflow_run all use this selection; changed latest tuple before deploy prevents stale refresh | Real producer-selection positive/negative TUnit cases, native CI capture and freshness/provider join |
722722

723723
TASK-FAIL-SEPARATE-001 (root) records policy/requirements/ADR before edits.
724724
TASK-FAIL-SEPARATE-002 (root) removes site generation and qualify/deploy from

‎docs/implementation/benchmark-json-ci-development-2026-10-04.json‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -248,5 +248,24 @@
248248
"fullWebsiteCoverageAndPublication": "pending",
249249
"noLocalBrowserRunForThisStage": true,
250250
"representativeScaleQualification": false
251+
},
252+
"deliveredLatestEligibilityFollowup": {
253+
"parentCodeCheckpoint": "873cd1a36ad14ab966065c924c71a292ea681083",
254+
"actualOriginalCIJob": 111408451894,
255+
"originalEvidenceArtifactId": 11300230252,
256+
"originalEvidenceDigest": "sha256:afd6f4ebab3744b016c978945ee5fb2abfba0ebc3e05c426deade9a74bc4d021",
257+
"originalMetadataConfirmedNewestCompletedWasCanceled": 37190736149,
258+
"productionParserOverOriginalAPISelectedReadyJSON": 37184989107,
259+
"productionParserFileSha256": "a0454cc821f44b45ebc931d95a3fd415816d5932f555ee4161a3c1547e486818",
260+
"newestEligibleAggregateFailureDoesNotFallBack": true,
261+
"fullDevelopmentBuild": {
262+
"passed": true,
263+
"warnings": 0,
264+
"errors": 0,
265+
"logSha256": "149dfb0ec99bc886622d2a750a954a538af8b1dfce202108d914048cbf001922"
266+
},
267+
"formatPassed": true,
268+
"nativeCIRecheck": "pending next scoped push",
269+
"browserExecutedLocally": false
251270
}
252271
}

‎docs/implementation/status.json‎

Lines changed: 19 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -3847,6 +3847,25 @@
38473847
"skipped": 0,
38483848
"trxSha256": "91e50a30de911b506dd83962b509bf720d9adfebf527b742aefff4f3dc3faa2d",
38493849
"qualification": "development_only"
3850+
},
3851+
"checkpointRevision": "873cd1a36ad14ab966065c924c71a292ea681083",
3852+
"genuineCITriggers": {
3853+
"pushRun": 37192832065,
3854+
"completedBenchmarkWorkflowRun": 37192849271,
3855+
"separateWebsiteJobsCreated": true,
3856+
"ordinarySuitesNotRepeatedForWorkflowRun": true
3857+
},
3858+
"deliveredLatestSelectionRepair": {
3859+
"cause": "newer canceled pending workflows selected ahead of completed 73 JSON",
3860+
"originalJobId": 111408451894,
3861+
"originalArtifactId": 11300230252,
3862+
"originalArtifactDigest": "sha256:afd6f4ebab3744b016c978945ee5fb2abfba0ebc3e05c426deade9a74bc4d021",
3863+
"apiMetadataSha256": "f8816dacf8a60b68a41b8911185b8c97f2e27cbd816980f1331972b929d19007",
3864+
"repair": "filter_completed_success_or_failure_candidates_before_selecting_latest; no_older_fallback_after_invalid_selected_aggregate",
3865+
"originalAPIProductionParserSelected": 37184989107,
3866+
"developmentBuildWarnings": 0,
3867+
"developmentBuildErrors": 0,
3868+
"nativeCIReverification": "pending_next_scoped_push"
38503869
}
38513870
}
38523871
}

‎scripts/AGENTS.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -61,3 +61,6 @@ No repository skills are installed or applicable to this tooling module. Do not
6161

6262
## Latest benchmark selection, 2026-10-04
6363
- The latest owner clarification requires every separate CI website build, whether own-main push/manual or actual completed Benchmarks workflow_run, to consume the newest completed own-main Benchmarks push/manual producer. Authenticate the trigger separately from the latest selected run/attempt/source/event/conclusion; require its successful complete aggregate and original archives, reject invalid latest evidence without historical fallback, and repeat latest selection for predeploy freshness. This explicitly supersedes publication selection pinned blindly to the triggering producer.
64+
65+
## Latest completed result eligibility, 2026-10-04
66+
- Latest benchmark metrics means the newest completed own-main push/manual Benchmarks producer with success/failure conclusion. Pending, skipped and canceled workflows have no completed comparison cohort and MUST NOT displace ready JSON. Exclude them before choosing the latest eligible producer; then reject its missing/corrupt/failed aggregate without older fallback. A canceled workflow_run event still cannot authorize publication. This refines the latest-result rule without accepting incomplete or fabricated measurements.

‎scripts/Features/BenchmarkComparisons/site-isolated-github-runs.mjs‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -39,7 +39,8 @@ export function flattenSiteRuns(pages) {
3939
}
4040

4141
export function selectLatestSiteProducer(runs, workflow) {
42-
const completed = runs.map(run => validateSiteRun(run, workflow)).filter(run => run.status === 'completed');
42+
const completed = runs.map(run => validateSiteRun(run, workflow))
43+
.filter(run => run.status === 'completed' && SITE_GH.producerConclusions.includes(run.conclusion));
4344
completed.sort((left, right) => right.run_number - left.run_number);
4445
const run = completed[0];
4546
requireSite(run !== undefined && SITE_GH.producerConclusions.includes(run.conclusion));

‎site/AGENTS.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -47,3 +47,6 @@
4747

4848
## Latest benchmark and separate build action, 2026-10-04
4949
- The latest owner clarification explicitly permits static website building after benchmark JSON or a separate CI trigger, superseding the interim static-build prohibition above. The selected independent CI action runs on own-main push/manual and completed Benchmarks events, always consumes the newest completed own-main push/manual benchmark and its successful authenticated aggregate, and retains original source/run/attempt, failed/null cells, all website gates and predeploy freshness. Never fall back to older evidence after rejecting the latest completed producer.
50+
51+
## Latest completed result eligibility, 2026-10-04
52+
- Latest benchmark metrics means the newest completed own-main push/manual Benchmarks producer with success/failure conclusion. Pending, skipped and canceled workflows have no completed comparison cohort and MUST NOT displace ready JSON. Exclude them before choosing the latest eligible producer; then reject its missing/corrupt/failed aggregate without older fallback. A canceled workflow_run event still cannot authorize publication. This refines the latest-result rule without accepting incomplete or fabricated measurements.

‎tests/KeyLoad.SiteTests/Features/BenchmarkComparisons/SiteBenchmarkProducerEventTests.cs‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -137,6 +137,7 @@ internal sealed class SiteBenchmarkLatestProducerTests
137137
[Test]
138138
[Arguments("newerPending")]
139139
[Arguments("olderOtherEvent")]
140+
[Arguments("cancelledLatest")]
140141
public async Task AC_BC_FAIL_020_LatestCompletedProducerSpansPushAndManualWhileIgnoringPending(string scenario)
141142
{
142143
var token = TestContext.Current!.Execution.CancellationToken;
@@ -156,7 +157,6 @@ await Assert.That(result.GetProperty(SiteIsolatedGitHubFields.Result)
156157
[Test]
157158
[Arguments("failedAggregate")]
158159
[Arguments("missingAggregate")]
159-
[Arguments("cancelledLatest")]
160160
[Arguments("noCompleted")]
161161
public async Task AC_BC_FAIL_020_InvalidLatestCompletedProducerCannotFallBackToOlderMeasurements(string scenario)
162162
{

0 commit comments

Comments
 (0)