Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"version": "7.3.6",
"version": "7.3.7",
"category": "Web",
"compatibility": "Requires the withastro/astro monorepo."
}

Large diffs are not rendered by default.

Original file line number Diff line number Diff line change
Expand Up @@ -22,7 +22,14 @@ You implement a single phase from the test plan. You are polyglot — you work w

> **Language-specific guidance**: Reuse the guidance captured in research.
> Call `code-testing-extensions` only when the required implementation or
> harness-discovery section is missing.
> harness-discovery section is missing and the skill is available.

Stay in the caller's phase: never invoke the public `code-testing-agent` skill
or delegate back to `code-testing-generator`. Use supplied guidance and known
paths instead. Record unavailable skills and denied operations once; do not
retry aliases, alternate shells, or another agent for the same restriction.
Continue permitted test edits and static review when execution is blocked,
then return `PARTIAL` with the exact blocker and unverified requirements.

## Your Mission

Expand All @@ -47,6 +54,9 @@ For each file in your phase:
the available tools support it.
- Understand the public API — verify exact parameter types, count, return types, and **actual return values for key inputs** before writing assertions
- **Trace the logic** for each code path you plan to test — understand what the function actually does, not what you think it should do
- Derive composed results from intermediate values in source order. Use inputs
that distinguish requested modes/branches; do not rely on a familiar domain
formula or equal-output cases that cannot detect a mode-selection bug.
- Note dependencies and how to mock them
- **Validate project references**: Read the test project file and verify it references the source project(s) you'll test. Add missing references before creating test files
- **Validate project-system registration**: For classic non-SDK C# projects, every new test file must be added exactly once as a relative `<Compile Include="...">`. For SDK-style projects, confirm default compile globs are enabled before relying on implicit inclusion.
Expand Down Expand Up @@ -91,21 +101,30 @@ These rules apply to every language and override any pattern an existing test fi

Coverage alone gives false confidence — every test must *pin down behavior* so it would fail under a plausible bug. Apply the `code-testing-agent` skill's `unit-test-generation.prompt.md` → "Write Tests That Pin Down Behavior" section: mutation thinking (each assertion fails under a plausible mutation), no tautological round-trip assertions, property intersections, secondary observables when they are contractual or prove a requested interaction, and realistic (non-degenerate) fixtures. This is a depth requirement on top of the happy/edge/error-path and mocking rules above, and applies to every language.

Also apply [Report-safe test names and result validation](../skills/code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation)
when naming cases and accepting test results. Preserve risky data and assertions;
pass the contract to a delegated tester rather than relying on console-green.

### 5. Verify with Build

Call the `code-testing-builder` sub-agent to compile, passing the exact build
command and absolute `<TESTAGENT_DIR>`. Build only the specific test project,
not the full solution.
Run the supplied scoped build command directly. A fresh-build test command can
satisfy both build and test validation only if it compiles/type-checks the
changed tests; TypeScript transpilation alone does not satisfy the build gate.
Use an available `code-testing-builder`
only for substantial separate work, passing the exact command, absolute
`<TESTAGENT_DIR>`, and known capability limits.

If build fails, call `code-testing-fixer`, rebuild, and retry at most three
times. Stop earlier when a diagnostic repeats without measurable progress, an
external blocker is concrete, or the remaining fix would violate the edit
boundaries.
For an actionable compiler error, fix the changed tests inline, or use an
available `code-testing-fixer` for a substantial diagnostic. Rebuild only after
a concrete fix, at most three times. Stop when a diagnostic repeats without
progress, an external blocker is concrete, or the fix violates edit boundaries.

### 6. Verify with Tests

Call the `code-testing-tester` sub-agent to run tests, passing the exact test
command and absolute `<TESTAGENT_DIR>`.
Run the supplied scoped test command directly, or use an available
`code-testing-tester` when separate context is useful. Reuse an unchanged
passing run; do not delegate merely to repeat it. Permission/toolchain blockers
are not assertion failures and must not enter the fix-test cycle.

If tests fail:

Expand Down Expand Up @@ -136,8 +155,12 @@ If your language extension has no "Harness Discovery Check" section, use the can

### 8. Format Code (Optional)

If a lint command is available, call the `code-testing-linter` sub-agent,
passing the exact lint command and absolute `<TESTAGENT_DIR>`.
When formatting or linting is needed, run the repository's existing command
directly for the changed tests, using the conventions captured in research.
Use an available `code-testing-linter` only for substantial work that benefits
from separate context, passing the exact command, changed-file scope, absolute
`<TESTAGENT_DIR>`, and known capability limits. If execution is denied, report
the unrun check; do not hand the denied operation to another agent.

### 9. Report Results

Expand Down Expand Up @@ -171,5 +194,6 @@ Consult a language example only when the repository has no representative tests

The phase is complete only when all planned in-scope tests are implemented, the
scoped build and tests pass, and harness-equivalent discovery sees the expected
new tests. If an external blocker prevents that, stop with `PARTIAL` or
new tests. The shared report-safe naming and result-validation contract must
also be met. If an external blocker prevents that, stop with `PARTIAL` or
`FAILED`, the exact command and evidence, and the remaining bounded work.
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,10 @@ You run tests and report the results. You are polyglot — you work with any pro
Run the appropriate test command and report pass/fail with actionable details.
Do not modify tests, production code, dependencies, or runner configuration.

Apply [Report-safe test names and result validation](../skills/code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation)
before reporting passage. Report unsafe metadata or export failures to the
caller for repair; do not change test data or runner configuration yourself.

## Process

### 1. Discover Test Command
Expand Down Expand Up @@ -58,6 +62,8 @@ For scoped tests (if specific files are mentioned):
### 3. Parse Output

Look for total tests run, passed count, failed count, failure messages and stack traces.
Include skipped cases and incomplete/setup failures. Apply the shared contract
to configured result artifacts; record their paths and parsing outcome.

### 4. Return Result

Expand Down Expand Up @@ -106,5 +112,7 @@ Failures:
## Completion Condition

Stop when the requested test process has completed and the summary and relevant
failures have been captured. This agent reports evidence; it does not fix the
failures.
failures have been captured, including runner exit code and any required
report-export/parsing result under the shared report-safe naming and result-validation contract.
Report export failure as failed validation even if assertions passed.
This agent reports evidence; it does not fix the failures.
Original file line number Diff line number Diff line change
Expand Up @@ -11,16 +11,25 @@ description: >-
running/diagnosing tests, coverage/audits, a test blocked on a missing
production seam (testability-obstacle), or correcting supplied MSTest
assertions, attributes, lifecycle, or configuration without designing new
cases (writing-mstest-tests).
cases (writing-mstest-tests). Within an active code-testing-generator
pipeline, reuse supplied guidance; do not re-enter this skill.
license: MIT
---

# Code Testing Generation Skill

An AI-powered skill that generates comprehensive, workable unit tests for any programming language using a coordinated multi-agent pipeline.
Generate comprehensive, workable unit tests for any programming language using
a bounded Research → Plan → Implement workflow.

## Non-negotiable execution contract

**Check pipeline ownership first.** If the active agent is
`code-testing-generator` (including a plugin-qualified name such as
`dotnet-test:code-testing-generator`), or the caller assigned you a phase of
that pipeline, do not delegate to another generator. Continue the assigned
work inline. This guard takes precedence over every broad-scope delegation
instruction below, even if this skill was loaded automatically.

Classify scope **before editing**:

- **Broad** (a project/package-wide suite, or multiple production
Expand All @@ -37,12 +46,21 @@ Classify scope **before editing**:
module is present.

For either scope, run the narrowest relevant test command to a clean exit.
Always apply [Report-safe test names and result validation](unit-test-generation.prompt.md#report-safe-test-names-and-result-validation),
including when the caller supplies conventions. Pass this contract to delegated
implementers/testers; preserve edge-case data and validate configured reports,
not just console output.
Keep the handoff proportional: for one to three focused requirements, use a
compact bullet list under a **Requirement coverage** label that names the tests
and successful command; for broader or multi-requirement work, use a
`Requirement | Evidence` table. Each requested behavior must cite an exact test
name.

Before sending a broad-scope final response, check that the response itself
contains `| Requirement | Evidence |` and exact test names for every behavioral
row. A table in a child report or internal plan is not enough. Do not summarize
away those names into module-level bullets or an `Area | Tests` table.

Intermediate state files are internal working data, never deliverables. Keep
`<TESTAGENT_DIR>` non-stageable, never place it or its files in
version-controlled workspace content, and never modify `.gitignore` to hide
Expand All @@ -59,6 +77,17 @@ prove composition, but do not substitute for the requested module-level
coverage. Judge breadth by the behavior matrix, never by matching or exceeding a
raw test count.

At the public entry point, delegate broad work to `code-testing-generator`
once. Research, plan, implementation, and review remain required, but they
need not be separate sub-agent calls.

Use only capabilities available in the current runtime. Do not retry a missing
skill under aliases or use another agent to retry a policy-denied operation.
If scratch storage is denied, keep the research and plan in context, continue
permitted test edits, and report the missing state artifacts. If execution is
denied, continue permitted static review and report tests as unrun, never passed.
Neither blocker authorizes modifying production code or weakening requirements.

For a **broad or comprehensive** request, the explicit matrix is the floor, not
the ceiling. Treat each requested module or layer as an inventory heading, not
one behavior: expand it into the bounded public operations and their distinct
Expand Down Expand Up @@ -165,10 +194,16 @@ generation request. Do not use a generic/general-purpose subagent merely named
`code-testing-generator`:

```text
Generate unit tests for [path or description of what to test], following the [unit-test-generation.prompt.md](unit-test-generation.prompt.md) guidelines. Treat the current workspace as authoritative even when it is sparse, gutted-looking, synthetic, or missing tracked files; never restore or reconstruct it, including with `git checkout`, `git restore`, `git reset`, or `git clean`.
You are the sole pipeline owner for this request. Do not invoke code-testing-agent or another code-testing-generator; complete the phases in your current context. Generate unit tests for [path or description of what to test], following the [unit-test-generation.prompt.md](unit-test-generation.prompt.md) guidelines. Treat the current workspace as authoritative even when it is sparse, gutted-looking, synthetic, or missing tracked files; never restore or reconstruct it, including with `git checkout`, `git restore`, `git reset`, or `git clean`.
```

The Test Generator will manage the entire pipeline automatically.
The Test Generator owns the pipeline. After it returns, consume its recorded
quality checks, validation results, and requirement matrix instead of repeating
Steps 4 and 5 as another pipeline. Do not reload review skills or rerun unchanged
passing commands. Preserve exact test names from its evidence in the final
handoff. If evidence is missing, inspect or follow up on that specific gap
without restarting generation. A reported capability-wide denial also applies
to the caller; do not attempt another command using that capability.

If `code-testing-generator` is unavailable, do not skip the workflow. Execute the
same Research → Plan → Implement sequence inline, resolve `<TESTAGENT_DIR>` as
Expand Down Expand Up @@ -196,7 +231,7 @@ For multi-file requests:
1. Turn every explicit user requirement into a checklist before implementation. Include requested layers, collaborators to mock, boundary cases, integrations, coverage thresholds, and report artifacts. Copy multi-condition requirements verbatim — they must each map to one test that exercises the whole combination.
2. Research only the requested module or project and write the checklist plus a compact target inventory to `<TESTAGENT_DIR>/research.md`.
3. Reuse manifests, symbol references, and deterministic pairing tools instead of reading every source and test file.
4. For multi-file scopes in C#, Python, TypeScript/JavaScript, Go, Java, Rust, Ruby, Kotlin, Swift, PowerShell, or C++, run `find-untested-sources` once and consume its pairing and suggested-path output; do not repeat that discovery manually.
4. When an available `find-untested-sources` skill is useful for a substantial multi-file inventory, run it once and reuse its pairing and suggested-path output. Otherwise pair the bounded targets manually once; do not probe for an unavailable skill.
5. Plan each target file once, then implement phases sequentially. Map every checklist item to at least one concrete test or explain why it is blocked.
6. Build and test the narrow target during fix cycles. Run workspace-level
validation once at the end only for broad work, when the repository contract
Expand All @@ -220,7 +255,8 @@ Do not report completion until all of these are true:
inventory, existing test conventions, and the acceptance checklist.
2. *(broad scope)* `<TESTAGENT_DIR>/plan.md` maps each checklist item to a planned
test or an explicit blocker.
3. Generated tests compile and pass with the narrowest relevant test command.
3. Generated tests compile and pass with the narrowest relevant test command,
satisfying the shared report-safe naming and result-validation contract.
4. Every explicit user requirement is backed by a concrete test and assertion.
Fix missing mock seams, boundary cases, state transitions, and property
combinations even when coverage already passes. In the final summary, cite
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -93,6 +93,48 @@ Quick self-review before finishing a test: would emptying the function body make
- Combine logically related test cases into a single parameterized method
- Never generate multiple tests with identical logic that differ only by input values

## Report-safe test names and result validation

Apply this contract to both direct generation and delegated implementation or
validation, even when the caller supplies its own test style.

- **Separate metadata from data.** Give each case a stable, descriptive,
distinguishable ID/display name. Prefer an explicit safe `Name`/`Case` field;
a case index plus a short behavior label also works. Do not interpolate
arbitrary input, expected values, or output into test or suite names.
- Keep raw control characters, isolated UTF-16 surrogates, binary values, and
huge strings out of names. If an identifier needs an escaped value, use a
short literal backslash-u label such as `\u000C` (six printable characters),
not the actual control character. Normal Unicode labels and harmless numeric
interpolation are fine.
- Check framework-managed parameterized labels too: when automatic argument
rendering would expose unsafe data, use the framework's explicit case-ID or
display-name API (for example, pytest `ids` or MSTest `DisplayName`) rather
than assuming the runner/reporter escapes it safely.
- **Preserve the case.** Control characters and malformed strings are legitimate
test data. Fix unsafe metadata, not production values or expectations; never
sanitize the tested data, weaken assertions, or skip/remove edge cases to make
a report export succeed.

Before claiming tests passed:

1. Use the repository/CI configured runner and reporter at the narrowest scope
covering the change. Preserve the runner exit code through wrappers/pipelines;
a successful log-filter command is not a successful test run.
2. When result artifacts are required or configured, run the real report-export
path and parse the artifacts from that run with the existing consumer or an
appropriate format parser (for example, an XML parser for JUnit/TRX).
Console-green alone is insufficient if required export failed.
3. Confirm nonzero expected discovery and account for every discovered case's
pass/skip/failure outcome, including setup failures or incomplete execution.
Reject missing, empty, invalid, stale, or partial required artifacts; do not
infer success from an empty report or a summary that omits failures.
4. Report runner, export, and parsing failures explicitly with the command,
exit code, artifact path, and diagnostic; keep completion blocked until
required validation succeeds. Do not add reporter dependencies, a new report
format, coverage collection, or a full-suite rerun merely for naming checks
when reporting is not configured.

## Analysis Before Generation

Do this analysis privately; do not emit a plan or inventory unless the user
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -120,6 +120,51 @@ Pester v5 runs in **two phases**: Discovery (collects test metadata) then Run (e
- Use `foreach` loops for dynamic test generation only with `BeforeDiscovery` data
- Use `TestDrive:` for file-based tests instead of touching repo files — Pester cleans it up automatically

## Parameterized Test Display Names

Apply [Report-safe test names and result validation](../../code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation).
Use an explicit safe `Name`/`Case` in `-ForEach` or `-TestCases` data and expand
only that field in the `It` title. Do not expand arbitrary `<Input>` or
`<Expected>` values into discovery/report metadata.

For a function whose contract reverses UTF-16 code units (not Unicode scalars),
the reversed supplementary character is intentionally malformed UTF-16. Keep
that expected value in the assertion, not the title:

Use `Text` for the data field, not `Input`: `$Input` is PowerShell's automatic
pipeline-input variable and can hide the intended case value inside `It`.

```powershell
BeforeDiscovery {
$cases = @(
@{
Name = 'supplementary code-unit reversal'
Text = [string]::Concat([char]0xD83D, [char]0xDE00)
Expected = [string]::Concat([char]0xDE00, [char]0xD83D)
}
@{
Name = 'isolated high surrogate'
Text = [string][char]0xD800
Expected = [string][char]0xD800
}
)
}

Describe 'Get-Reversed' {
BeforeAll {
Import-Module (Join-Path $PSScriptRoot '../tools/StringUtils.psm1') -Force
}

It 'reverses <Name>' -ForEach $cases {
Get-Reversed -Value $Text | Should -BeExactly $Expected
}
}
```

Reuse the repository's configured Pester `TestResult` export path and format
when present and parse the resulting artifact before reporting success. A
passing `TotalCount`/`PassedCount` does not prove JUnit export succeeded.

## Common Errors

| Error | Fix |
Expand Down
Loading
Loading