Repository navigation
Clarify identity semantics for parameterized tests and merged retries #66
Description
Activity
IMHO Evangelink raises important questions to be aware of as creator of a producer.
My opinion from the viewpoint of a HIL (hardware in the loop) test framework is as followed:
Questions 1 and 2
CTRF should keeping to strongly recommend that each set of parameters (or in Evangelink words: argument row) should be a distinct concrete (or in CTRF terms: logical) test case that should have its own
testId.
(For why I have choosen "concrete test case" over "logical test case" see #70.)That said, IMHO the CTRF spec should:
- keeping being indifferent on how a producer is achieving this (by row index, by creating a hash over the parameter set, by some kind of additional row identifier, …)
- keeping to allow to not follow this recommendation
- (in addition) explicitely state that
testIdmust be omitted if a producer may not provide a deterministic and stabletestIdat all
Why this matters
Without a strong opinion regarding determinism and stability of
testId,testIdwould be useless for consumers.Note
For producer implementations, for example, it may make sense to generate a hash over the parameter set’s values and, if that is not possible, falling back to use the parameter set's index as fallback.
I am aware that generating a hash of the parameter set's values can also lead to variabletestIdsif the parameter set (or the table, due to a new column) is expanded.The not so easily answerable question in the parameter set extension case is, if the concretely executed test (flow logic plus set of parameters) before and after the extension is testing „the same case“. Of course, this needs be decided by the developers/testers case by case.
Question 3
IMHO, CTRF should leave open the question of whether and how file system paths, class names, and method names are included in the
testId, as currently is the case.Note
Producer implementations can, for example, allow test methods to be assigned a (unique) identifier that is included in the calculation of the
testId, if available. Otherwise, they can fall back on using the path, class names, and method names.Parametrization and
executionIdTo make it more obvious maybe the spec should state explicitly that
executionIdandtestIdare orthogonal concepts. That said I strongly object „treating the method as the logical test and the row as an execution“.Questions 4 and 5
No opinion (yet) regarding retry attempts.
Thanks @Evangelink, and thanks @gregweb for the additional feedback. Here is the intended direction for each point.
- Identity of parameterized cases
Each parameterized test case should have its own distinct testId. For example, login(username="test"), login(username="dev") and login(username="prod") are separate cases, each represented by its own test object in results.tests.
They can pass or fail independently, so distinct identities allow consumers to track each case’s history and flakiness across runs. Retries of the same parameterized case should retain the same testId. We will clarify this recommendation in the schema.
- Constructing a stable testId
For parameterized tests, testId should remain associated with the same case across runs within the producer’s documented scope. The producer chooses how to achieve this; CTRF does not prescribe a particular generation algorithm.
Case-defining parameter values are one option, but a stable framework identifier or explicit case key may work equally well. Producers are not required to serialize parameter values to construct an ID, particularly when those values are complex or sensitive.
A row index is suitable only if it reliably identifies the same case within that scope. If reordering the data source causes an existing ID to refer to a different case, it undermines the ID’s usefulness for tracking history and flakiness. We will consider clarifying the schema wording to reflect this guidance.
- Renames and file moves
Producers should preserve testId across test renames and file moves where their identity mechanism supports it.
A producer that derives IDs from names or paths may legitimately generate a different ID after a refactor. CTRF encourages preserving identity for the same test case, but does not require producers to detect renames or moves, or to maintain continuity beyond their documented identity scope. We will clarify this guidance in the schema.
- Execution identity across retry processes
Known attempts of one coordinated execution should share an executionId, even when they run in separate processes or their reports remain separate.
For example, if an orchestrator runs a suite and then reruns only failed cases as retries, each case’s initial attempt, subsequent attempts, and merged test result, if produced, belong to the same execution lifecycle. They should retain the same executionId, while each distinct report document has its own reportId.
Merging does not determine whether executions are retries; that relationship comes from the coordinated workflow. An independent execution of the same case should have a different executionId. We will clarify this guidance in the schema.
This also makes testId and executionId distinct concepts: testId identifies the case across runs, while executionId identifies one execution lifecycle, including its retries.
- Identity of the final attempt
You’re right that the final attempt currently lacks an attemptId. We will add an optional attemptId property to the Test object to identify that final attempt, complementing the existing property on earlier entries in retryAttempts.
These clarifications and the new optional property are planned changes. The parameterized and cross-process retry examples you’ve shared will be useful for illustrating the intended model.
Thanks, this clarifies the distinction between a case, an execution lifecycle, and an attempt. We will align our coordinated cross-process retries so the raw reports and merged result retain one executionId per case, while each document retains its own reportId.
Two remaining points would help us make that implementation interoperable:
-
What is an acceptable documented stability scope for index-based parameterized cases, and when should testId be omitted?
For example, a data source returns
[login("alice"), login("bob")]in one run and[login("bob"), login("alice")]in the next. A framework UID based on method plus row index remains deterministic, but the same UID now identifies a different case. Is a documented "same assembly/method and unchanged data-source ordering" scope acceptable, or should a producer omit testId whenever it cannot establish that ordering is stable? Similarly, if a framework only supplies fresh per-run UIDs, should the generic report producer omit testId and retain that UID only under extra?We do not want to hash arbitrary argument values: some are sensitive, nonserializable, or have nondeterministic representations. An explicit stable case key would solve this when the framework provides one, but a generic producer cannot invent that guarantee.
-
Which tagged specification/schema version will introduce the final Test.attemptId, and what is the recommended transition for older reports?
Both the current main schema and v0.1.0 currently reject a top-level Test.attemptId. Consider two raw reports for the same coordinated lifecycle:
- initial failed attempt:
testId=T,executionId=E; - successful retry:
testId=T,executionId=E; - merged test:
testId=T,executionId=E, with the failed attempt in retryAttempts.
Without a distinct attempt identity on each raw test, a consumer cannot correlate those attempts with the merged history. Is storing an attempt key under extra until the new schema is released a reasonable bridge? For legacy input that has executionId but no attemptId, we would avoid treating executionId as an attemptId: once E is shared by the lifecycle, doing so would assign the same attemptId to multiple attempts. Is omitting the optional attemptId for those legacy attempts the intended behavior?
- initial failed attempt:
The parameterized-case recommendation and rename/file-move behavior are otherwise clear; neither seems to require a prescribed hashing algorithm or automatic rename detection.
-
Thanks for the follow-up @Evangelink,
We would prefer producers to include testId wherever meaningful case identity can be provided. As noted above, an index-based ID is suitable only when it reliably identifies the same case within the producer’s documented scope.
In your Alice/Bob example, reordering could associate an existing ID with a different case, leading consumers to attach the wrong history. Documenting the ordering limitation helps explain the scheme, but does not itself establish stability.
We haven’t established an omission rule where that stability cannot be guaranteed. This also applies to frameworks that provide only fresh per-run UIDs. We’d like to assess these cases further before prescribing a mapping or omission rule. We will clarify the schema wording at a later stage so we can ensure the guidance is clear.
For the final-attempt property, we are targeting CTRF specification 0.2.0 very soon, in the next day or so. This will add optional attemptId to the Test object.
Reports using it should declare specVersion: "0.2.0", since the 0.1.0 schema rejects that property.
Until then, storing an attempt key under producer-namespaced extra is a reasonable bridge. For legacy inputs without an attempt identity, leaving the optional attemptId absent is appropriate.
Producer context
Microsoft.Testing.Platform (MTP) produces CTRF for multiple test frameworks, so it cannot assume every framework assigns the same stability semantics to its native test UID.
Today the MTP CTRF reporter keeps the framework-provided UID under
tests[].extra.uidrather than claiming it is a CTRFtestId:For MSTest specifically, that UID is a versioned hash of the assembly file name, fully-qualified test name, parameter types, and (for data-driven tests) the row index:
This is stable for many rebuilds, but a rename, assembly rename, or reordered dynamic data source can change it.
What the current spec resolves
The identity model is clear that:
testIdidentifies a logical test case and should be stable across runs within the producer's documented scope;executionIdis unique for one execution of that logical test within a run;attemptIdidentifies an individual entry inretryAttempts;PR #57 and PR #62 provide useful general guidance, but we could not find guidance for parameterized cases or for reports transformed by retry merging.
Questions
testId, or should all rows share the method'stestIdand be distinguished byparametersplusexecutionId?testIdbe based on canonical parameter values (stable across data-source reordering) rather than a row index? How should producers handle values that cannot be serialized deterministically or safely?testId, or is preserving identity across renames an expected producer responsibility?executionId, or should each raw document have its own execution identity and the merged document get another one?attemptIdintentionally limited to previous attempts inretryAttempts? The final attempt is represented by the test object, which has noattemptId; that makes it difficult to correlate the final attempt between a raw per-process report and a merged report.Why this matters for MTP
Mapping an arbitrary framework UID to
testIdmay overstate its stability, while omitting standardized identities prevents CTRF consumers from correlating parameterized rows and retry artifacts. We would like to document and implement the intended model rather than establish a .NET-specific convention.Possible interpretations include treating every parameter row as a logical test case, or treating the method as the logical test and the row as an execution. We are not prescribing either; a normative recommendation and one worked parameterized/retry example would unblock the mapping.