Skip to content

Clarify retry semantics: what etryAttempts[] contains, how etries is counted, and how to report a run where only failed tests are re-executed #58

Description

@Evangelink

Hi! We produce CTRF reports from Microsoft.Testing.Platform (the Microsoft.Testing.Extensions.CtrfReport extension) and we hit three points where the written spec and the published example disagree, or where we could not find an answer at all. All three affect the same feature — retries — so I'm grouping them in one issue. Happy to split if you prefer.

Context: our platform has two very different retry mechanisms, and we want both to produce spec-correct CTRF.

  1. In-process retries — the test framework re-executes a test inside a single run and publishes each attempt. One CTRF document, retryAttempts[] on the test.
  2. Orchestrated retries (--retry-failed-tests N) — the runner re-launches the whole test host process, and each re-launch executes only the tests that failed in the previous attempt. Every attempt is a separate process that writes its own CTRF document.

1. Does retryAttempts[] contain every attempt, or only the re-executions?

The spec text and the example point in opposite directions.

Section 11 (Retry Attempt Object) intro:

Each retry attempt represents a single re-execution of a test after an initial failure or according to a retry policy.

Terminology:

Retry / Retry Attempt: Any attempt with an attempt number greater than 1.

Both of those read as "retryAttempts[] holds attempts 2..N".

But §11.1 says:

Attempt number 1 represents the first execution; higher numbers represent retries.

…which only makes sense if attempt 1 can appear in the array. And examples/with-retries.json does exactly that — a test that ran three times lists all three attempts, including the first (attempt: 1, failed) and the final passing one (attempt: 3, passed), while the test object's own status is passed:

{
  "name": "flaky test",
  "status": "passed",
  "duration": 300,
  "flaky": true,
  "retries": 2,
  "retryAttempts": [
    { "attempt": 1, "status": "failed", "duration": 120, "message": "Connection timeout" },
    { "attempt": 2, "status": "failed", "duration": 130, "message": "Connection timeout" },
    { "attempt": 3, "status": "passed", "duration": 140 }
  ]
}

Question: is the example normative here — i.e. retryAttempts[] is the complete attempt history (1..N, including the attempt whose outcome is mirrored in the test object's status)? Or is it meant to hold only attempts 2..N?

This matters a lot for consumers: if a producer emits only 2..N, the diagnostics of the first failure (message/trace/stdout) have nowhere to live, because the test object itself carries the final outcome. Our current implementation emits attempts 1..N-1 (every attempt except the final one, whose data is on the test object), which is a third shape and satisfies §9.20 below but not the example — hence the question.

2. retries — count of retryAttempts[], or count minus one?

§9.20:

retries … SHOULD equal the count of entries in retryAttempts.

The example has "retries": 2 with three entries in retryAttempts. Under the "array holds every attempt" reading, retries is attempts - 1 (the number of re-executions), which contradicts §9.20 as written.

Question: should §9.20 be amended to retries SHOULD equal retryAttempts.length - 1 (when the array holds every attempt), or should the example be corrected? A consumer computing "how many times did this test run" needs one unambiguous rule.

3. Reporting a logical run when only the failed subset is re-executed, in separate processes

This is the case that has no worked example. Concretely, a 4-test suite with --retry-failed-tests 3:

attempt tests executed outcome
1 4 2 passed, 2 failed
2 2 1 passed, 1 failed
3 1 1 failed
4 1 1 failed

Each attempt writes its own CTRF document, so attempt 4's document legitimately reports summary.tests: 1. Reading §2.1 ("A logical run MAY consist of multiple physical executions … Results from such executions MAY be merged into a single CTRF document") together with §2.11 (a merge produces a new document with a different reportId, leaving the inputs immutable), our understanding is:

  • the per-attempt documents stay as they are (immutable artifacts, one per physical execution);
  • the runner additionally emits one merged document for the logical run, in which each logical test appears once — status = the final attempt's outcome, earlier attempts folded into retryAttempts[], flaky: true where a failure was followed by a pass;
  • summary.tests therefore counts logical tests (4 here), not executions (8), and summary.flaky counts the recovered ones — consistent with §8.1 ("SHOULD equal the length of the tests array") and §4.9.1 ("Producers SHOULD NOT emit duplicate identity values within a single CTRF document when those duplicates would make correlation ambiguous").

Questions:

  • Is that the intended representation? A short note or an example in the spec would help — today "shards" is the only merge scenario discussed, and shards have disjoint tests, whereas retry attempts overlap, which is precisely the case where naive concatenation inflates summary.tests.
  • Should the merged document and the per-attempt documents share a runId (§5.4), so a consumer can tell they describe the same logical run? The spec frames runId around shards/workers; retry attempts of the same suite feel like the same concept, but it's not spelled out.
  • Is there any expectation about the test object's duration in the merged document — the final attempt's duration, or the sum across attempts? (The example's 300 matches neither its own attempt durations nor their sum, so we assume it's illustrative only.)

Thanks a lot for the spec — and for the quick turnaround on #53, which we followed for labels. Whatever you confirm here, we'll implement in Microsoft.Testing.Extensions.CtrfReport and reference back from microsoft/testfx#10293.

Activity

  1. Ma11hewThomas commented on Jul 29, 2026

    @Ma11hewThomas
    Collaborator

    Thanks for raising these. You’re right about the inconsistencies. We’ll correct and align them.

    The intended model matches your implementation:

    retryAttempts[] is the attempt history preceding the final attempt. It contains attempts 1..N-1, including the initial execution and any earlier retries.

    The final attempt is excluded from retryAttempts[] because its outcome and associated information are represented by the test object itself.

    retries is the number of re-executions after the initial execution. It equals retryAttempts.length, and the final attempt number is retries + 1.

    This matches the approach currently taken by your implementation. We’ll update the specification, schema descriptions, and examples so they express this consistently.

  2. Ma11hewThomas commented on Jul 29, 2026

    @Ma11hewThomas
    Collaborator

    Regarding retries orchestrated across separate processes, your suggested approach seems sensible and aligns with the retry history model. We'll consider adding a short note or an example in the spec for this use case.

    The physical-execution documents and merged document represent the same logical run, so they should share a runId. Each remains a distinct immutable artifact and should have its own reportId.

    In this case, the final attempt's duration should be used for the test object duration in the merged report.

    Thanks again for sharing these use cases, which are really important to help shape CTRF.

  3. forevercrab321-svg commented on Aug 3, 2026

    @forevercrab321-svg

    On case 3: the merged document is the part I'd pin hardest, because a merge has a failure mode neither physical document has. It can be short one input and still be a perfectly valid CTRF report. retries then counts what the merger found rather than what ran, and a consumer computing flakiness gets a confident number off an incomplete set.

    We had that exact shape in an aggregate count. Our spend caps read a total out of a content-range header and never checked whether the request had succeeded, so a broken count came back as zero spent, which is the most permissive answer the system can give. Three caps failed open at once, one of them on an unauthenticated endpoint. The header also has a form, 0-24/*, that NaNs straight past the comparison without ever passing through zero.

    Does the merged report have anywhere to declare how many physical documents it expected? An incomplete merge is currently detectable only from the orchestrator's logs, and the report is the artifact that outlives them.

  4. Evangelink commented on Sep 28, 2026

    @Evangelink
    Author

    Following up on the remaining runId scope questions for Microsoft.Testing.Platform after the retry-history clarification landed in #62.

    The confirmed retry case is clear: raw retry-process documents and the merged document share one runId, while every immutable document has its own reportId. Two adjacent cases are still ambiguous for a generic producer:

    1. Modules launched by one CLI invocation. A multi-project dotnet test invocation starts each module as a separate root test application/process tree. MTP therefore has a per-module execution identifier, but no automatic outer invocation identifier. Should those sibling modules normally be considered one CTRF logical run and share a runId, or is that intentionally producer/workflow-defined? Today MTP only correlates them when the caller explicitly supplies a logical-run ID.
    2. Raw plus merged artifacts. Because source documents and their merged document share the same runId, a consumer that ingests every retained artifact can double-count the run. There is no standardized document role (source, partial, merged), source reportId list, or expected-input count. This also leaves the incomplete-merge concern raised in the existing comment above unresolved.

    Would you prefer the spec to define those boundaries and lineage/completeness fields here, or should producers keep them in namespaced extra metadata? The current MTP merger preserves runId only when every input agrees, which is safe but cannot describe why a merged report is complete or which immutable inputs it supersedes for analysis.

  5. Ma11hewThomas commented on Oct 6, 2026

    @Ma11hewThomas
    Collaborator

    For the multi-project dotnet test example, we consider the invocation one logical run. The individual modules and their process trees are physical executions contributing to that run, so their CTRF documents should share a runId, with each distinct document having its own reportId.

    This follows §2.1’s distinction between logical runs and physical executions, and §5.4’s guidance that documents from the same coordinated execution SHOULD share a runId. §5.3 defines the separate identity of each report artifact. We should make the multi-project invocation example explicit in the spec as it’s a great example.

    Because runId is optional, omitting it when that context is unavailable is acceptable; supplying it enables consumers to correlate the module reports correctly.

  6. Ma11hewThomas commented on Oct 6, 2026

    @Ma11hewThomas
    Collaborator

    @Evangelink: “There is no standardized document role (source, partial, merged), source reportId list, or expected-input count.”

    @forevercrab321-svg: “Does the merged report have anywhere to declare how many physical documents it expected?”

    I believe both of these points relate to provenance, though lineage and completeness need distinct treatment.

    Recording the source reportId values would let consumers identify which documents contributed to a merged report. We have a provenance proposal that introduces provenance.inputs for this purpose, but it is not yet part of the current specification.

    Although the intended list alone would not prove completeness: it describes the inputs the merger used, rather than the inputs it expected. This is the gap highlighted by @forevercrab321-svg, a structurally valid report can still represent an incomplete merge. This will require further thought.

    For now, namespaced extra metadata is the available mechanism for recording these details.

    I think these are worth assessing together as part of the provenance work, with explicit guidance for consumer behavior.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions