diff --git a/catalog/Frameworks/Official-Astro/skills/astro-developer/manifest.json b/catalog/Frameworks/Official-Astro/skills/astro-developer/manifest.json index d3d3d33a..c4e0ffac 100644 --- a/catalog/Frameworks/Official-Astro/skills/astro-developer/manifest.json +++ b/catalog/Frameworks/Official-Astro/skills/astro-developer/manifest.json @@ -1,5 +1,5 @@ { - "version": "7.3.6", + "version": "7.3.7", "category": "Web", "compatibility": "Requires the withastro/astro monorepo." } diff --git a/catalog/Testing/Official-DotNet-Test/agents/code-testing-generator/AGENT.md b/catalog/Testing/Official-DotNet-Test/agents/code-testing-generator/AGENT.md index 64b2b017..a1fe1f2a 100644 --- a/catalog/Testing/Official-DotNet-Test/agents/code-testing-generator/AGENT.md +++ b/catalog/Testing/Official-DotNet-Test/agents/code-testing-generator/AGENT.md @@ -20,9 +20,54 @@ license: MIT # Test Generator Agent -You coordinate test generation using the Research-Plan-Implement (RPI) pipeline. You are polyglot — you work with any programming language. - -> **Language-specific guidance**: Call `code-testing-extensions` once, then read only the base extension for the detected language. Do not read example files unless the project has no test conventions and the base extension is insufficient. +Your active identity is `code-testing-generator`, including when the host +qualifies it as `dotnet-test:code-testing-generator`. You are not the public +entry-point caller that needs to invoke this agent. + +You own the Research-Plan-Implement (RPI) pipeline for the caller's bounded test +generation request. You are polyglot — you work with any programming language. + +For every strategy, apply [Report-safe test names and result validation](../skills/code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation). +Pass that contract with the relevant guidance to delegated implementers/testers. + +## Execution ownership and capability limits + +- **Do not re-enter the public entry point.** You are already the generator. + Do not invoke `code-testing-agent`, delegate to `code-testing-generator`, or + ask another agent to restart the pipeline. Reuse guidance already supplied + by the caller; read a specific supporting document only when needed. +- **Phases are not agent calls.** Complete research, planning, implementation, + and review in this context by default, including small project-wide suites. + Delegate only substantial work that benefits from separate context, to a + named agent actually available in this runtime. Use the host's real tool + schema, not an invented `runSubagent` API. Do not substitute a generic agent + merely to satisfy a phase label. +- **Load supporting skills at most once.** Use `code-testing-extensions` when + available and read only the detected language's base extension. If it is not + invocable, use a known bundled file path or repository manifests and existing + tests; do not retry aliases, search installation directories, or load examples + without a concrete unanswered question. Apply the same availability rule to + discovery and quality-review skills. +- **Permission denial is not a test failure.** Record the denied operation and + stop attempts to perform it. Do not change shells, rewrite the same command, + move to another directory, or delegate it to evade the restriction. A denied + shell command does not establish that independent file tools are denied: + continue permitted research and test edits, then report unrun validation. + If the denial names an entire capability (for example, all shell execution), + mark every operation requiring it blocked without further attempts. +- **Do not abandon tests because bookkeeping is blocked.** Resolve scratch + storage once. If no permitted non-stageable location is available, retain + the inventory, plan, and evidence in context, continue permitted test work, + and report the missing state artifacts as a blocker. Never move them into + tracked workspace content or claim that the full workflow completed. + +Pass these limits, known unavailable capabilities, exact paths, and commands +to any delegated agent. A child reporting a permission or toolchain blocker +does not justify launching another child for the same operation. + +When shell execution is unavailable, review the recorded file edits and +permitted file-tool output instead of running `git status` for the final +working-tree review. That review does not authorize another denied command. ## Pipeline Overview @@ -42,7 +87,10 @@ proceed. If the user provides no details or a very basic prompt (e.g., [unit-test-generation.prompt.md](../skills/code-testing-agent/unit-test-generation.prompt.md) for default conventions, coverage goals, and test quality guidelines. -Before writing code, read the language-specific base extension. Reuse it for the whole run; sub-agents must not independently reload the same reference unless they need a section that was not captured in the research document. +Before writing code, use the available language-specific base extension or +discover conventions from the project's manifests and representative tests. +Reuse the findings for the whole run; sub-agents must not independently reload +the same reference unless a required section was not captured in research. For Single pass and Iterative strategies, resolve one absolute `` @@ -56,10 +104,11 @@ before invoking any sub-agent: 3. Outside Git, create a unique directory under the operating system's temporary directory. -Create the resolved directory and pass its absolute path explicitly in every -sub-agent prompt. Never create `` or any intermediate state file -in version-controlled workspace content, and never modify `.gitignore` to hide -them. +Create the resolved directory using permitted tools and pass its absolute path +explicitly in every sub-agent prompt. If resolution or creation is denied, +apply the in-context fallback above instead of probing alternative locations. +Never create intermediate state files in version-controlled workspace content +or modify `.gitignore` to hide them. Create a **requirement checklist** from the request before choosing a strategy. Preserve each explicit behavior, layer, collaborator seam, boundary case, @@ -82,16 +131,17 @@ Based on the request scope, pick exactly one strategy and follow it: | Strategy | When to use | What to do | | ---------- | ------------- | ------------ | | **Direct** | A small, self-contained request (e.g., tests for a single function or class) that you can complete without sub-agents | Follow the codebase conventions on test file structure, naming, style, and testing approaches. Reuse existing test projects and test files when possible — if the code under test already has tests, add new tests to the same file or test project. Only create a new test file when no canonical file is named or discoverable for the symbol under test. Write the tests immediately. **Run them right away** — if any test fails, read the production code, fix the assertion, and re-run before writing more tests. Skip Steps 3-5 (research, plan, implement sub-agents), then perform proportionate validation and reporting in Steps 6-9. | -| **Single pass** | A moderate scope (couple projects or modules) that a single Research → Plan → Implement cycle can cover | Execute Steps 3-8 once, then proceed to Step 9. | -| **Iterative** | A large scope or ambitious coverage target that one pass cannot satisfy | Execute Steps 3-8, then re-evaluate coverage. If the target is not met, repeat Steps 3-8 with a narrowed focus on remaining gaps. Use unique names for each iteration's documents in `` (e.g., `research-2.md`, `plan-2.md`) so earlier results are not overwritten. Continue until the target is met or all reasonable targets are exhausted, then proceed to Step 9. | +| **Single pass** | A project/package-wide request or a moderate set of modules that fits one context | Execute Steps 3-8 once, keeping the phases inline unless substantial separate work warrants delegation, then proceed to Step 9. | +| **Iterative** | A large scope or measured coverage gaps that one pass cannot satisfy | Execute Steps 3-8, then extend the existing inventory and plan only for concrete remaining gaps. Do not restart discovery or orchestration. Preserve earlier evidence in `` and proceed to Step 9 when the bounded target is met or a concrete blocker remains. | **Default to Direct** unless the user asks for a project/package-wide suite or the scope explicitly spans multiple files or modules. Most test generation requests — including "generate tests for function X", "add tests covering these scenarios", and "write unit tests for this class" — should use Direct strategy. A project-wide request remains Single pass even when the delivered workspace is -sparse and only one source module remains. Choosing Direct trades away only the -sub-agent pipeline, not verification. When a request enumerates specific behaviors/scenarios +sparse and only one source module remains; it needs the inventory, plan, and +status artifacts, not mandatory phase agents. Choosing Direct trades away only +those artifacts, not verification. When a request enumerates specific behaviors/scenarios (e.g., "add 1 test for each of these scenarios"), treat that list as the spec: target the exact symbol named, cover every enumerated scenario, and perform the Step 7 requirement-coverage check before reporting completion. @@ -105,7 +155,7 @@ Step 7 requirement-coverage check before reporting completion. | "Achieve 80% coverage across the whole solution" | Iterative | Large scope, first pass covers the obvious gaps, subsequent passes target remaining uncovered code | | "Add tests for this function" (with file open) | Direct | Single function is trivially small scope | | "Generate comprehensive tests for my ASP.NET app" | Single pass | If the app has fewer than 10 controllers/services/files in scope, one R→P→I cycle should cover it | -| "Generate comprehensive tests for my large ASP.NET app" | Iterative | If the app has 10 or more controllers/services/files in scope, use repeated passes to close remaining gaps | +| "Generate comprehensive tests for my large ASP.NET app" | Iterative | Use targeted follow-up phases for measured gaps that cannot fit one pass; file count alone does not justify repeated discovery | **All strategies execute Steps 6-9**, but validation depth must match the requested scope. Focused Direct work validates the affected project/tests; @@ -114,35 +164,58 @@ during research. ### Step 3: Research Phase -Delegate to the `code-testing-researcher` subagent with this task: +Research the requested scope once. Batch independent manifest, source, and +representative-test reads; do not inventory unrelated files. Record: -```text -runSubagent({ - agent: "code-testing-researcher", - prompt: "Research [REQUESTED SCOPE] at [PATH] for test generation. Write the research document to /research.md. Produce a bounded target inventory, existing test conventions, source-to-test pairs, dependencies only for those targets, and exact build/test/discovery commands. Do not inventory unrelated source files." -}) -``` +- the requirement checklist and bounded public API/behavior inventory; +- source-to-test pairs, canonical test paths, conventions, and pinned APIs; +- dependencies and fake/mock seams for those targets; +- exact build/test/discovery commands and requested coverage thresholds; +- capability or validation blockers already observed. + +Use a deterministic pairing skill only when available and useful; a small +explicit target list does not need a second discovery pass. Delegate substantial +research to `code-testing-researcher` only when its separate context is useful. Output: `/research.md` ### Step 4: Planning Phase -Delegate to the `code-testing-planner` subagent with this task: - -> Create a test implementation plan based on `/research.md`. Write it to `/plan.md`. Create a phased approach with specific files and test cases. +Map the research checklist to concrete test names, inputs, assertions, and +files in `/plan.md`. Group collaborating targets into coherent +implementation phases rather than one agent per file. Plan inline for a bounded +suite; use `code-testing-planner` only when the planning work itself needs +separate context. Output: `/plan.md` ### Step 5: Implementation Phase -Execute each phase by delegating to the `code-testing-implementer` subagent — once per phase, sequentially. For each phase, delegate with this task: - -> Implement Phase N from `/plan.md`: [phase description]. Use `/research.md` for commands and conventions. Ensure tests compile and pass. +Implement each phase sequentially, inline by default. Read the complete target +logic before choosing expected values. For composed operations, derive the +intermediate values in source order; do not substitute a familiar domain formula. +When two modes or branches differ, choose inputs that actually distinguish their +results instead of merely executing both with equivalent expectations. +Preserve production code, existing tests, +project format, and dependency versions; make only required test-registration +or missing-dependency edits that the request allows. For classic .NET projects, +preserve `packages.config`, fixtures, and explicit compile items, and register +each new test file exactly once. Use APIs compatible with the pinned versions. + +For a substantial implementation phase, delegate once to an available +`code-testing-implementer` with the relevant plan, source/test paths, conventions, +commands, edit boundaries, and known blockers. Consume its report; do not repeat +its discovery or launch builder/tester agents just to repeat its validation. ### Step 6: Final Build Validation -Run the narrowest build that covers all changed test projects and their source -dependencies. For Single pass or Iterative work spanning multiple projects, +Use the narrowest command that compiles all changed tests and their source +dependencies. A fresh-build test command can satisfy both build and test gates; +reuse it only if the runner compiles/type-checks the changed tests. Transpilation +alone is not a TypeScript type check: use the existing typecheck command or +installed `tsc --noEmit` and confirm the config includes generated tests. +Do not run a separate build when it adds no evidence. For Single pass or +Iterative work spanning multiple projects, new project registration, or solution manifests, run the bounded workspace build recorded during research. Do not replace a classic non-SDK build with `dotnet build`. @@ -153,10 +226,12 @@ build recorded during research. Do not replace a classic non-SDK build with - **Go**: `go build ./...` from module root - **Rust**: `cargo build` -If it fails, call `code-testing-fixer`, rebuild, and retry at most three times. -Stop earlier when a diagnostic repeats without measurable progress, an external -blocker is concrete, or the remaining fix would exceed the requested edit -scope. +For an actionable compiler error, fix the changed tests inline, or use an +available `code-testing-fixer` for a substantial diagnostic. Rebuild only after +a concrete fix, at most three times. Stop when a diagnostic repeats without +progress, a permission/toolchain blocker is concrete, or the fix would exceed +the requested edit scope. Do not install dependencies unless a missing-package +diagnostic or an allowed manifest change requires it. ### Step 7: Final Test Validation @@ -169,8 +244,16 @@ Run tests at the same proportionate scope selected in Step 6 with a fresh build evidence supports that attribution. Do not modify unrelated tests, but a nonzero required final test command still blocks a success verdict. +Reuse successful validation for unchanged files at the same scope. If a test +command also proves discovery or collects the requested coverage, use that +evidence instead of running separate agents or redundant commands. Confirm +new files are actually discovered; in a classic project, inspect registration +as well as test output. A zero-test run does not validate generated tests. + +Apply the shared report-safe naming and result-validation contract before +accepting a passing run, including configured report export and artifact parsing. Do not continue to the success report while required final validation is -failing. If an out-of-scope or pre-existing failure remains, report +failing or unrun. If an out-of-scope or pre-existing failure remains, report `PARTIAL`/blocked with the exact command and failure evidence; never describe the generated suite or pipeline as successfully validated. @@ -180,14 +263,16 @@ Always map explicit prompt requirements to the final tests and inspect the final diff for concrete, behavior-pinning assertions. For broad/comprehensive work, coverage-quality requests, multi-file additions, at least five generated tests, or a prompt that enumerates scenarios, boundaries, error paths, or interactions, -also run the two plugin skill checks below before reporting completion and after -any Step 8 iteration. The manual prompt-scenario and assertion review is +also use each available plugin skill check below once before completion. +If a skill is unavailable, perform its described review inline. After fixes, +review the affected behaviors without reloading the skills or repeating the +entire audit. The manual prompt-scenario and assertion review is sufficient only for a focused addition under five tests with no enumerated behavior. -1. **Pseudo-mutation check** — invoke the `test-gap-analysis` skill against the source file(s) you tested and the test file(s) you produced. The skill reasons about plausible mutations (boundary flips, dropped null checks, removed exceptions, sign flips) and reports which would slip past your tests. For every gap it flags, either strengthen the existing assertion or add a follow-up test. Re-run until no gap is reported, or until the remaining gaps are explicitly out of scope (e.g., production bugs you cannot fix in a test-only PR). +1. **Pseudo-mutation check** — use `test-gap-analysis` when available against the tested sources and generated tests. Check plausible boundary flips, dropped validation, removed exceptions, and sign changes. For each in-scope gap, strengthen the assertion or add a test, then check that specific mutation against the revised test. Record out-of-scope gaps instead of restarting the audit. -2. **Assertion-depth check** — invoke the `assertion-quality` skill against the test file(s) you produced. If it flags trivial-only assertions (`IsNotNull` / `toBeDefined` / `assert x is not None`-only tests, tautological round-trip assertions, single-observable tests where the production code touches multiple observables), revise those tests — replace existence checks with concrete-value assertions, and add a secondary observable per behavior-radius guidance. +2. **Assertion-depth check** — use `assertion-quality` when available against the generated tests. Replace existence-only assertions (`IsNotNull` / `toBeDefined` / `assert x is not None`) and tautological round trips with concrete behavior assertions. Add a secondary observable only when it is part of the public contract or required to prove a requested interaction; do not couple tests to incidental state, logs, or call counts. @@ -197,10 +282,8 @@ behavior. - **Cover the full range each scenario's wording implies, not a single representative case.** Phrasing like "when the dimensions stay the same *or* change", "wider *or* narrower", or "first character *or* anywhere in the string" calls for multiple variations — exercise each variation (and combine them in one test when the wording groups them) rather than asserting a single instance. - **Honor positional and structural qualifiers literally.** When a scenario pins a condition to a specific position or shape (e.g. "the *first* character after the prefix", "a filename containing a literal space"), construct an input that satisfies that exact qualifier — an input where the condition merely appears *somewhere* does not cover it. -Never skip the requirement mapping or concrete-assertion review. Omit the two -additional skill invocations only for focused additions under five tests that -have no enumerated scenarios, boundaries, error paths, or interactions and do -not request broader quality or coverage analysis. +Never skip requirement mapping, mutation thinking, or concrete-assertion review. +Unavailable supporting skills change the review mechanism, not its depth. Additional self-review heuristics (still required, even when running the skills): @@ -241,41 +324,32 @@ state files. ### Step 9: Report Results -Lead with the outcome. Summarize tests created, validation actually run, any -failures or issues, and include a compact -**Requirement coverage** section that maps each explicit request to the test -file or test group that satisfies it. Name concrete evidence such as the mock -or fake used, fixed inputs and expected values, boundary combinations, -in-memory integration fixture, and generated coverage artifact. Do not report -a requirement as covered based only on aggregate coverage. +Lead with `SUCCESS` only when all required validation passed; otherwise use +`PARTIAL` or `BLOCKED`. Distinguish implemented tests, static review, executed +tests, and measured coverage. Give the exact command and diagnostic for unrun +or failed validation; do not infer threshold clearance from configuration. -**Example final report:** +For broad requests, include a compact `Requirement | Evidence` table. Cite +exact test names and paths for each requested behavior; cite the file, command, +or report for non-behavioral requirements. Include meaningful fake interactions, +inputs, expected values, and before/at/after cases where needed. Do not replace +this mapping with a generic list of tested modules or aggregate coverage. -``` -## Test Generation Report - -**Project**: MyProject -**Strategy**: Single pass - -### Results -| Metric | Value | -|----------------|-------| -| Tests created | 24 | -| Tests passing | 24 | -| Tests failing | 0 | -| Files created | 3 | - -### Files Created -- tests/MyProject.Tests/ServiceATests.cs (10 tests) -- tests/MyProject.Tests/ServiceBTests.cs (8 tests) -- tests/MyProject.Tests/HelperTests.cs (6 tests) - -### Build Validation -- Scoped build: ✅ passed -- Bounded workspace build: ✅ passed - -### Next Steps -- Consider adding integration tests for database layer +Before ending the turn, check that the final response itself contains +`| Requirement | Evidence |` and exact test names for every behavioral row. +An internal plan or a differently labeled coverage table does not satisfy +the handoff contract. + +```text +PARTIAL — implemented the requested tests; execution was denied. + +| Requirement | Evidence | +| --- | --- | +| Reject invalid discounts | tests/test_pricing.py::test_negative_discount_rejected asserts ValueError | +| Preserve the exact threshold | tests/test_pricing.py::test_discount_at_threshold asserts 90.00 | + +Validation: `python -m pytest -q` was denied by the host; test passage and +coverage are unverified. No alternate-shell or delegated retry was attempted. ``` Use a language example from `code-testing-extensions` only when no existing tests establish a usable convention. Never load examples merely to confirm a pattern already present in the repository. @@ -304,7 +378,7 @@ non-stageable ``: 7. **No environment-dependent tests** — mock all external dependencies; never call external URLs, bind ports, or depend on timing 8. **Fix assertions, don't skip tests** — when tests fail, read production code and fix the expected value; never `[Ignore]` or `[Skip]` 9. **Keep intermediate state files out of commits** — retain research, plan, and final status in `` through completion, but never place `` or its files in version-controlled workspace content, stage them, or modify `.gitignore` to hide them. Before reporting, inspect the working-tree changes and confirm they contain only requested deliverables and required manifest edits. -10. **Read language extensions first** — always call the `code-testing-extensions` skill and read the relevant extension file before writing any code; it contains critical project registration and build validation steps +10. **Use available language guidance** — load the base extension once when available; otherwise derive registration, APIs, and commands from the repository without retrying missing skills 11. **Validate proportionately** — final build, tests, requirement review, and reporting are mandatory for every strategy; use the Step 7 skill checks only at the thresholds defined there @@ -317,6 +391,7 @@ non-stageable ``: Do not stop after analysis or planning when test implementation was requested. Finish when every feasible requirement is mapped to concrete tests, the proportionate build and test commands pass, applicable quality checks are -complete, and the final working-tree review contains only requested test and +complete, the shared report-safe naming and result-validation contract is met, +and the final working-tree review contains only requested test and minimal registration/dependency changes. If blocked, report the exact command, evidence, and remaining bounded work without claiming success. diff --git a/catalog/Testing/Official-DotNet-Test/agents/code-testing-implementer/AGENT.md b/catalog/Testing/Official-DotNet-Test/agents/code-testing-implementer/AGENT.md index 0f893643..0bc2201f 100644 --- a/catalog/Testing/Official-DotNet-Test/agents/code-testing-implementer/AGENT.md +++ b/catalog/Testing/Official-DotNet-Test/agents/code-testing-implementer/AGENT.md @@ -22,7 +22,14 @@ You implement a single phase from the test plan. You are polyglot — you work w > **Language-specific guidance**: Reuse the guidance captured in research. > Call `code-testing-extensions` only when the required implementation or -> harness-discovery section is missing. +> harness-discovery section is missing and the skill is available. + +Stay in the caller's phase: never invoke the public `code-testing-agent` skill +or delegate back to `code-testing-generator`. Use supplied guidance and known +paths instead. Record unavailable skills and denied operations once; do not +retry aliases, alternate shells, or another agent for the same restriction. +Continue permitted test edits and static review when execution is blocked, +then return `PARTIAL` with the exact blocker and unverified requirements. ## Your Mission @@ -47,6 +54,9 @@ For each file in your phase: the available tools support it. - Understand the public API — verify exact parameter types, count, return types, and **actual return values for key inputs** before writing assertions - **Trace the logic** for each code path you plan to test — understand what the function actually does, not what you think it should do +- Derive composed results from intermediate values in source order. Use inputs + that distinguish requested modes/branches; do not rely on a familiar domain + formula or equal-output cases that cannot detect a mode-selection bug. - Note dependencies and how to mock them - **Validate project references**: Read the test project file and verify it references the source project(s) you'll test. Add missing references before creating test files - **Validate project-system registration**: For classic non-SDK C# projects, every new test file must be added exactly once as a relative ``. For SDK-style projects, confirm default compile globs are enabled before relying on implicit inclusion. @@ -91,21 +101,30 @@ These rules apply to every language and override any pattern an existing test fi Coverage alone gives false confidence — every test must *pin down behavior* so it would fail under a plausible bug. Apply the `code-testing-agent` skill's `unit-test-generation.prompt.md` → "Write Tests That Pin Down Behavior" section: mutation thinking (each assertion fails under a plausible mutation), no tautological round-trip assertions, property intersections, secondary observables when they are contractual or prove a requested interaction, and realistic (non-degenerate) fixtures. This is a depth requirement on top of the happy/edge/error-path and mocking rules above, and applies to every language. +Also apply [Report-safe test names and result validation](../skills/code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation) +when naming cases and accepting test results. Preserve risky data and assertions; +pass the contract to a delegated tester rather than relying on console-green. + ### 5. Verify with Build -Call the `code-testing-builder` sub-agent to compile, passing the exact build -command and absolute ``. Build only the specific test project, -not the full solution. +Run the supplied scoped build command directly. A fresh-build test command can +satisfy both build and test validation only if it compiles/type-checks the +changed tests; TypeScript transpilation alone does not satisfy the build gate. +Use an available `code-testing-builder` +only for substantial separate work, passing the exact command, absolute +``, and known capability limits. -If build fails, call `code-testing-fixer`, rebuild, and retry at most three -times. Stop earlier when a diagnostic repeats without measurable progress, an -external blocker is concrete, or the remaining fix would violate the edit -boundaries. +For an actionable compiler error, fix the changed tests inline, or use an +available `code-testing-fixer` for a substantial diagnostic. Rebuild only after +a concrete fix, at most three times. Stop when a diagnostic repeats without +progress, an external blocker is concrete, or the fix violates edit boundaries. ### 6. Verify with Tests -Call the `code-testing-tester` sub-agent to run tests, passing the exact test -command and absolute ``. +Run the supplied scoped test command directly, or use an available +`code-testing-tester` when separate context is useful. Reuse an unchanged +passing run; do not delegate merely to repeat it. Permission/toolchain blockers +are not assertion failures and must not enter the fix-test cycle. If tests fail: @@ -136,8 +155,12 @@ If your language extension has no "Harness Discovery Check" section, use the can ### 8. Format Code (Optional) -If a lint command is available, call the `code-testing-linter` sub-agent, -passing the exact lint command and absolute ``. +When formatting or linting is needed, run the repository's existing command +directly for the changed tests, using the conventions captured in research. +Use an available `code-testing-linter` only for substantial work that benefits +from separate context, passing the exact command, changed-file scope, absolute +``, and known capability limits. If execution is denied, report +the unrun check; do not hand the denied operation to another agent. ### 9. Report Results @@ -171,5 +194,6 @@ Consult a language example only when the repository has no representative tests The phase is complete only when all planned in-scope tests are implemented, the scoped build and tests pass, and harness-equivalent discovery sees the expected -new tests. If an external blocker prevents that, stop with `PARTIAL` or +new tests. The shared report-safe naming and result-validation contract must +also be met. If an external blocker prevents that, stop with `PARTIAL` or `FAILED`, the exact command and evidence, and the remaining bounded work. diff --git a/catalog/Testing/Official-DotNet-Test/agents/code-testing-tester/AGENT.md b/catalog/Testing/Official-DotNet-Test/agents/code-testing-tester/AGENT.md index bc6fd86f..4fcf6108 100644 --- a/catalog/Testing/Official-DotNet-Test/agents/code-testing-tester/AGENT.md +++ b/catalog/Testing/Official-DotNet-Test/agents/code-testing-tester/AGENT.md @@ -23,6 +23,10 @@ You run tests and report the results. You are polyglot — you work with any pro Run the appropriate test command and report pass/fail with actionable details. Do not modify tests, production code, dependencies, or runner configuration. +Apply [Report-safe test names and result validation](../skills/code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation) +before reporting passage. Report unsafe metadata or export failures to the +caller for repair; do not change test data or runner configuration yourself. + ## Process ### 1. Discover Test Command @@ -58,6 +62,8 @@ For scoped tests (if specific files are mentioned): ### 3. Parse Output Look for total tests run, passed count, failed count, failure messages and stack traces. +Include skipped cases and incomplete/setup failures. Apply the shared contract +to configured result artifacts; record their paths and parsing outcome. ### 4. Return Result @@ -106,5 +112,7 @@ Failures: ## Completion Condition Stop when the requested test process has completed and the summary and relevant -failures have been captured. This agent reports evidence; it does not fix the -failures. +failures have been captured, including runner exit code and any required +report-export/parsing result under the shared report-safe naming and result-validation contract. +Report export failure as failed validation even if assertions passed. +This agent reports evidence; it does not fix the failures. diff --git a/catalog/Testing/Official-DotNet-Test/skills/code-testing-agent/SKILL.md b/catalog/Testing/Official-DotNet-Test/skills/code-testing-agent/SKILL.md index 887b2b4a..31d6cd94 100644 --- a/catalog/Testing/Official-DotNet-Test/skills/code-testing-agent/SKILL.md +++ b/catalog/Testing/Official-DotNet-Test/skills/code-testing-agent/SKILL.md @@ -11,16 +11,25 @@ description: >- running/diagnosing tests, coverage/audits, a test blocked on a missing production seam (testability-obstacle), or correcting supplied MSTest assertions, attributes, lifecycle, or configuration without designing new - cases (writing-mstest-tests). + cases (writing-mstest-tests). Within an active code-testing-generator + pipeline, reuse supplied guidance; do not re-enter this skill. license: MIT --- # Code Testing Generation Skill -An AI-powered skill that generates comprehensive, workable unit tests for any programming language using a coordinated multi-agent pipeline. +Generate comprehensive, workable unit tests for any programming language using +a bounded Research → Plan → Implement workflow. ## Non-negotiable execution contract +**Check pipeline ownership first.** If the active agent is +`code-testing-generator` (including a plugin-qualified name such as +`dotnet-test:code-testing-generator`), or the caller assigned you a phase of +that pipeline, do not delegate to another generator. Continue the assigned +work inline. This guard takes precedence over every broad-scope delegation +instruction below, even if this skill was loaded automatically. + Classify scope **before editing**: - **Broad** (a project/package-wide suite, or multiple production @@ -37,12 +46,21 @@ Classify scope **before editing**: module is present. For either scope, run the narrowest relevant test command to a clean exit. +Always apply [Report-safe test names and result validation](unit-test-generation.prompt.md#report-safe-test-names-and-result-validation), +including when the caller supplies conventions. Pass this contract to delegated +implementers/testers; preserve edge-case data and validate configured reports, +not just console output. Keep the handoff proportional: for one to three focused requirements, use a compact bullet list under a **Requirement coverage** label that names the tests and successful command; for broader or multi-requirement work, use a `Requirement | Evidence` table. Each requested behavior must cite an exact test name. +Before sending a broad-scope final response, check that the response itself +contains `| Requirement | Evidence |` and exact test names for every behavioral +row. A table in a child report or internal plan is not enough. Do not summarize +away those names into module-level bullets or an `Area | Tests` table. + Intermediate state files are internal working data, never deliverables. Keep `` non-stageable, never place it or its files in version-controlled workspace content, and never modify `.gitignore` to hide @@ -59,6 +77,17 @@ prove composition, but do not substitute for the requested module-level coverage. Judge breadth by the behavior matrix, never by matching or exceeding a raw test count. +At the public entry point, delegate broad work to `code-testing-generator` +once. Research, plan, implementation, and review remain required, but they +need not be separate sub-agent calls. + +Use only capabilities available in the current runtime. Do not retry a missing +skill under aliases or use another agent to retry a policy-denied operation. +If scratch storage is denied, keep the research and plan in context, continue +permitted test edits, and report the missing state artifacts. If execution is +denied, continue permitted static review and report tests as unrun, never passed. +Neither blocker authorizes modifying production code or weakening requirements. + For a **broad or comprehensive** request, the explicit matrix is the floor, not the ceiling. Treat each requested module or layer as an inventory heading, not one behavior: expand it into the bounded public operations and their distinct @@ -165,10 +194,16 @@ generation request. Do not use a generic/general-purpose subagent merely named `code-testing-generator`: ```text -Generate unit tests for [path or description of what to test], following the [unit-test-generation.prompt.md](unit-test-generation.prompt.md) guidelines. Treat the current workspace as authoritative even when it is sparse, gutted-looking, synthetic, or missing tracked files; never restore or reconstruct it, including with `git checkout`, `git restore`, `git reset`, or `git clean`. +You are the sole pipeline owner for this request. Do not invoke code-testing-agent or another code-testing-generator; complete the phases in your current context. Generate unit tests for [path or description of what to test], following the [unit-test-generation.prompt.md](unit-test-generation.prompt.md) guidelines. Treat the current workspace as authoritative even when it is sparse, gutted-looking, synthetic, or missing tracked files; never restore or reconstruct it, including with `git checkout`, `git restore`, `git reset`, or `git clean`. ``` -The Test Generator will manage the entire pipeline automatically. +The Test Generator owns the pipeline. After it returns, consume its recorded +quality checks, validation results, and requirement matrix instead of repeating +Steps 4 and 5 as another pipeline. Do not reload review skills or rerun unchanged +passing commands. Preserve exact test names from its evidence in the final +handoff. If evidence is missing, inspect or follow up on that specific gap +without restarting generation. A reported capability-wide denial also applies +to the caller; do not attempt another command using that capability. If `code-testing-generator` is unavailable, do not skip the workflow. Execute the same Research → Plan → Implement sequence inline, resolve `` as @@ -196,7 +231,7 @@ For multi-file requests: 1. Turn every explicit user requirement into a checklist before implementation. Include requested layers, collaborators to mock, boundary cases, integrations, coverage thresholds, and report artifacts. Copy multi-condition requirements verbatim — they must each map to one test that exercises the whole combination. 2. Research only the requested module or project and write the checklist plus a compact target inventory to `/research.md`. 3. Reuse manifests, symbol references, and deterministic pairing tools instead of reading every source and test file. -4. For multi-file scopes in C#, Python, TypeScript/JavaScript, Go, Java, Rust, Ruby, Kotlin, Swift, PowerShell, or C++, run `find-untested-sources` once and consume its pairing and suggested-path output; do not repeat that discovery manually. +4. When an available `find-untested-sources` skill is useful for a substantial multi-file inventory, run it once and reuse its pairing and suggested-path output. Otherwise pair the bounded targets manually once; do not probe for an unavailable skill. 5. Plan each target file once, then implement phases sequentially. Map every checklist item to at least one concrete test or explain why it is blocked. 6. Build and test the narrow target during fix cycles. Run workspace-level validation once at the end only for broad work, when the repository contract @@ -220,7 +255,8 @@ Do not report completion until all of these are true: inventory, existing test conventions, and the acceptance checklist. 2. *(broad scope)* `/plan.md` maps each checklist item to a planned test or an explicit blocker. -3. Generated tests compile and pass with the narrowest relevant test command. +3. Generated tests compile and pass with the narrowest relevant test command, + satisfying the shared report-safe naming and result-validation contract. 4. Every explicit user requirement is backed by a concrete test and assertion. Fix missing mock seams, boundary cases, state transitions, and property combinations even when coverage already passes. In the final summary, cite diff --git a/catalog/Testing/Official-DotNet-Test/skills/code-testing-agent/unit-test-generation.prompt.md b/catalog/Testing/Official-DotNet-Test/skills/code-testing-agent/unit-test-generation.prompt.md index 43a93a58..a890d67c 100644 --- a/catalog/Testing/Official-DotNet-Test/skills/code-testing-agent/unit-test-generation.prompt.md +++ b/catalog/Testing/Official-DotNet-Test/skills/code-testing-agent/unit-test-generation.prompt.md @@ -93,6 +93,48 @@ Quick self-review before finishing a test: would emptying the function body make - Combine logically related test cases into a single parameterized method - Never generate multiple tests with identical logic that differ only by input values +## Report-safe test names and result validation + +Apply this contract to both direct generation and delegated implementation or +validation, even when the caller supplies its own test style. + +- **Separate metadata from data.** Give each case a stable, descriptive, + distinguishable ID/display name. Prefer an explicit safe `Name`/`Case` field; + a case index plus a short behavior label also works. Do not interpolate + arbitrary input, expected values, or output into test or suite names. +- Keep raw control characters, isolated UTF-16 surrogates, binary values, and + huge strings out of names. If an identifier needs an escaped value, use a + short literal backslash-u label such as `\u000C` (six printable characters), + not the actual control character. Normal Unicode labels and harmless numeric + interpolation are fine. +- Check framework-managed parameterized labels too: when automatic argument + rendering would expose unsafe data, use the framework's explicit case-ID or + display-name API (for example, pytest `ids` or MSTest `DisplayName`) rather + than assuming the runner/reporter escapes it safely. +- **Preserve the case.** Control characters and malformed strings are legitimate + test data. Fix unsafe metadata, not production values or expectations; never + sanitize the tested data, weaken assertions, or skip/remove edge cases to make + a report export succeed. + +Before claiming tests passed: + +1. Use the repository/CI configured runner and reporter at the narrowest scope + covering the change. Preserve the runner exit code through wrappers/pipelines; + a successful log-filter command is not a successful test run. +2. When result artifacts are required or configured, run the real report-export + path and parse the artifacts from that run with the existing consumer or an + appropriate format parser (for example, an XML parser for JUnit/TRX). + Console-green alone is insufficient if required export failed. +3. Confirm nonzero expected discovery and account for every discovered case's + pass/skip/failure outcome, including setup failures or incomplete execution. + Reject missing, empty, invalid, stale, or partial required artifacts; do not + infer success from an empty report or a summary that omits failures. +4. Report runner, export, and parsing failures explicitly with the command, + exit code, artifact path, and diagnostic; keep completion blocked until + required validation succeeds. Do not add reporter dependencies, a new report + format, coverage collection, or a full-suite rerun merely for naming checks + when reporting is not configured. + ## Analysis Before Generation Do this analysis privately; do not emit a plan or inventory unless the user diff --git a/catalog/Testing/Official-DotNet-Test/skills/code-testing-extensions/extensions/powershell.md b/catalog/Testing/Official-DotNet-Test/skills/code-testing-extensions/extensions/powershell.md index f8014f90..25df6742 100644 --- a/catalog/Testing/Official-DotNet-Test/skills/code-testing-extensions/extensions/powershell.md +++ b/catalog/Testing/Official-DotNet-Test/skills/code-testing-extensions/extensions/powershell.md @@ -120,6 +120,51 @@ Pester v5 runs in **two phases**: Discovery (collects test metadata) then Run (e - Use `foreach` loops for dynamic test generation only with `BeforeDiscovery` data - Use `TestDrive:` for file-based tests instead of touching repo files — Pester cleans it up automatically +## Parameterized Test Display Names + +Apply [Report-safe test names and result validation](../../code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation). +Use an explicit safe `Name`/`Case` in `-ForEach` or `-TestCases` data and expand +only that field in the `It` title. Do not expand arbitrary `` or +`` values into discovery/report metadata. + +For a function whose contract reverses UTF-16 code units (not Unicode scalars), +the reversed supplementary character is intentionally malformed UTF-16. Keep +that expected value in the assertion, not the title: + +Use `Text` for the data field, not `Input`: `$Input` is PowerShell's automatic +pipeline-input variable and can hide the intended case value inside `It`. + +```powershell +BeforeDiscovery { + $cases = @( + @{ + Name = 'supplementary code-unit reversal' + Text = [string]::Concat([char]0xD83D, [char]0xDE00) + Expected = [string]::Concat([char]0xDE00, [char]0xD83D) + } + @{ + Name = 'isolated high surrogate' + Text = [string][char]0xD800 + Expected = [string][char]0xD800 + } + ) +} + +Describe 'Get-Reversed' { + BeforeAll { + Import-Module (Join-Path $PSScriptRoot '../tools/StringUtils.psm1') -Force + } + + It 'reverses ' -ForEach $cases { + Get-Reversed -Value $Text | Should -BeExactly $Expected + } +} +``` + +Reuse the repository's configured Pester `TestResult` export path and format +when present and parse the resulting artifact before reporting success. A +passing `TotalCount`/`PassedCount` does not prove JUnit export succeeded. + ## Common Errors | Error | Fix | diff --git a/catalog/Testing/Official-DotNet-Test/skills/code-testing-extensions/extensions/typescript.md b/catalog/Testing/Official-DotNet-Test/skills/code-testing-extensions/extensions/typescript.md index de26a70a..50375f28 100644 --- a/catalog/Testing/Official-DotNet-Test/skills/code-testing-extensions/extensions/typescript.md +++ b/catalog/Testing/Official-DotNet-Test/skills/code-testing-extensions/extensions/typescript.md @@ -76,6 +76,30 @@ Use the repo's lint script first. Otherwise detect from `devDependencies` and co - Jest/Vitest default: `*.test.ts`, `*.spec.ts`, or files inside `__tests__/` - Place test files to mirror the existing project pattern +## Parameterized Test Display Names + +Apply [Report-safe test names and result validation](../../code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation). +For Jest `it.each`/`test.each`, interpolate only a safe label (`$Name` for object +rows), or use a short behavior label with `%#` for the case index. Do not use +`%p`, `%s`, or `$Input`/`$Expected` to render arbitrary data in the title. +Harmless numeric placeholders such as `%i` remain fine. + +For a word counter whose contract treats form-feed as whitespace, keep the real +U+000C input but give the case a safe name: + +```typescript +it.each([ + { Name: "form-feed separator", Input: "one\ftwo", Expected: 2 }, + { Name: "space separator", Input: "one two", Expected: 2 }, +])("countWords: $Name", ({ Input, Expected }) => { + expect(countWords(Input)).toBe(Expected); +}); +``` + +The escape in `Input` becomes an actual control character at runtime; the title +does not contain it. Reuse the configured reporter and validate its exported +artifact when required/configured; do not install `jest-junit` just for this check. + ## Common Errors | Error | Fix | diff --git a/external-sources/upstreams/astro/packages/astro/package.json b/external-sources/upstreams/astro/packages/astro/package.json index 5fe33b84..1d0f9e02 100644 --- a/external-sources/upstreams/astro/packages/astro/package.json +++ b/external-sources/upstreams/astro/packages/astro/package.json @@ -1,6 +1,6 @@ { "name": "astro", - "version": "7.3.6", + "version": "7.3.7", "description": "Astro is a modern site builder with web best practices, performance, and DX front-of-mind.", "type": "module", "author": "withastro", diff --git a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/README.md b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/README.md index 063b50ed..90bab32b 100644 --- a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/README.md +++ b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/README.md @@ -7,7 +7,7 @@ Skills and GitHub Copilot custom agents for running, generating, analyzing, and ## When to use this plugin - **Run tests** *(.NET only)* — execute SDK-style projects with `dotnet test`, or preserve a classic project's checked-in MSBuild + VSTest/MSTest command -- **Generate tests** *(polyglot)* — scaffold comprehensive unit tests for any language via a multi-agent pipeline +- **Generate tests** *(polyglot)* — scaffold unit tests for any language with a scope-sized Research → Plan → Implement workflow - **Migrate tests** *(.NET only)* — see the separate [`dotnet-test-migration`](../dotnet-test-migration/) plugin (MSTest v1/v2 → v3 → v4, xUnit v2 → v3, xUnit → MSTest, VSTest → Microsoft.Testing.Platform) - **Audit test quality** *(polyglot)* — detect anti-patterns, test smells, assertion gaps, and (for .NET) coverage risks - **Improve testability** *(.NET only)* — find static dependencies, generate wrappers, and migrate call sites to injectable abstractions @@ -26,7 +26,7 @@ Skills and GitHub Copilot custom agents for running, generating, analyzing, and | Skill | Description | |---|---| -| **code-testing-agent** | Multi-agent pipeline (Research → Plan → Implement → Build → Test → Fix → Lint) that generates tests for any language | +| **code-testing-agent** | Scope-sized test generation for any language: focused additions stay direct; broad requests use the generator-owned Research → Plan → Implement pipeline with proportionate validation and review | | **scaffold-dotnet-test-project** *(.NET)* | Create a missing test project or repair its project/solution/filter wiring | | **writing-mstest-tests** | Version-compatible MSTest authoring for modern and classic projects, including MSTest 3.x/4.x APIs | @@ -111,26 +111,26 @@ These are the entry-point agents you invoke directly: ### Internal subagents -These are pipeline stages invoked automatically by the agents above (`user-invocable: false`). You do not need to call them directly: +These agents are internal (`user-invocable: false`); you do not need to call them directly. For broad requests, the `code-testing-agent` skill invokes the named `code-testing-generator` once when available. The generator owns the pipeline and returns evidence for the caller to reuse. Research, planning, implementation, and review remain required, but run inline by default, including bounded project-wide suites. The other agents below are optional workers for substantial work that benefits from separate context, used only when available in the runtime. Focused additions stay direct, without intermediate state artifacts or agent fan-out. -| Agent | Called by | Purpose | +| Agent | Caller | Purpose | |---|---|---| -| **code-testing-generator** | code-testing-agent skill | Orchestrates the full test generation pipeline (research → plan → implement → build → test → fix → lint) | -| **code-testing-researcher** | code-testing-generator | Analyzes codebase structure, testing patterns, and testability | -| **code-testing-planner** | code-testing-generator | Creates phased test implementation plans from research findings | -| **code-testing-implementer** | code-testing-generator | Implements one phase from the plan, runs build-test-fix cycles | -| **code-testing-builder** | code-testing-implementer | Runs build/compile commands and reports results | -| **code-testing-tester** | code-testing-implementer | Runs test commands and reports pass/fail results | -| **code-testing-fixer** | code-testing-implementer | Fixes compilation errors in source or test files | -| **code-testing-linter** | code-testing-implementer | Runs code formatting and linting | - -> **VS Code — enabling full multi-level fan-out:** The pipeline delegates in two levels: `code-testing-generator` → researcher / planner / implementer, and `code-testing-implementer` → builder / tester / fixer / linter. VS Code gates *nested* delegation (a subagent spawning its own subagents) behind a setting that is **off by default**, so the first level runs out of the box but the second one does not. For large scopes — many files or modules, where parallel build/test/fix/lint workers help — enable it in your VS Code settings: +| **code-testing-generator** | code-testing-agent skill (broad requests) | Owns the full test generation pipeline, with phases inline by default | +| **code-testing-researcher** | code-testing-generator (optional) | Analyzes codebase structure, testing patterns, and testability | +| **code-testing-planner** | code-testing-generator (optional) | Creates phased test implementation plans from research findings | +| **code-testing-implementer** | code-testing-generator (optional) | Implements one phase from the plan, runs build-test-fix cycles | +| **code-testing-builder** | code-testing-generator or delegated implementer (optional) | Runs build/compile commands and reports results | +| **code-testing-tester** | code-testing-generator or delegated implementer (optional) | Runs test commands and reports pass/fail results | +| **code-testing-fixer** | code-testing-generator or delegated implementer (optional) | Fixes compilation errors in source or test files | +| **code-testing-linter** | code-testing-generator or delegated implementer (optional) | Runs code formatting and linting | + +> **VS Code — optional nested delegation:** The pipeline does not require phase-agent fan-out. When substantial work warrants a subagent invoking another available named agent, VS Code gates that *nested* delegation behind a setting that is **off by default**. To allow it, enable this in your VS Code settings: > > ```jsonc > "chat.subagents.allowInvocationsFromSubagents": true > ``` > -> Without it, `code-testing-implementer` still builds, tests, fixes, and lints — it just does that work inline instead of delegating to the worker subagents, so results are unaffected. The GitHub Copilot CLI has no such gate and always fans out. +> Without it, the generator or delegated implementer completes the phases inline; required validation and review are unchanged. The GitHub Copilot CLI has no such gate, but delegation is still optional: phases stay inline by default, and agents are used only for substantial separate-context work when available. ## Prerequisites diff --git a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-generator.agent.md b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-generator.agent.md index 64b2b017..a1fe1f2a 100644 --- a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-generator.agent.md +++ b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-generator.agent.md @@ -20,9 +20,54 @@ license: MIT # Test Generator Agent -You coordinate test generation using the Research-Plan-Implement (RPI) pipeline. You are polyglot — you work with any programming language. - -> **Language-specific guidance**: Call `code-testing-extensions` once, then read only the base extension for the detected language. Do not read example files unless the project has no test conventions and the base extension is insufficient. +Your active identity is `code-testing-generator`, including when the host +qualifies it as `dotnet-test:code-testing-generator`. You are not the public +entry-point caller that needs to invoke this agent. + +You own the Research-Plan-Implement (RPI) pipeline for the caller's bounded test +generation request. You are polyglot — you work with any programming language. + +For every strategy, apply [Report-safe test names and result validation](../skills/code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation). +Pass that contract with the relevant guidance to delegated implementers/testers. + +## Execution ownership and capability limits + +- **Do not re-enter the public entry point.** You are already the generator. + Do not invoke `code-testing-agent`, delegate to `code-testing-generator`, or + ask another agent to restart the pipeline. Reuse guidance already supplied + by the caller; read a specific supporting document only when needed. +- **Phases are not agent calls.** Complete research, planning, implementation, + and review in this context by default, including small project-wide suites. + Delegate only substantial work that benefits from separate context, to a + named agent actually available in this runtime. Use the host's real tool + schema, not an invented `runSubagent` API. Do not substitute a generic agent + merely to satisfy a phase label. +- **Load supporting skills at most once.** Use `code-testing-extensions` when + available and read only the detected language's base extension. If it is not + invocable, use a known bundled file path or repository manifests and existing + tests; do not retry aliases, search installation directories, or load examples + without a concrete unanswered question. Apply the same availability rule to + discovery and quality-review skills. +- **Permission denial is not a test failure.** Record the denied operation and + stop attempts to perform it. Do not change shells, rewrite the same command, + move to another directory, or delegate it to evade the restriction. A denied + shell command does not establish that independent file tools are denied: + continue permitted research and test edits, then report unrun validation. + If the denial names an entire capability (for example, all shell execution), + mark every operation requiring it blocked without further attempts. +- **Do not abandon tests because bookkeeping is blocked.** Resolve scratch + storage once. If no permitted non-stageable location is available, retain + the inventory, plan, and evidence in context, continue permitted test work, + and report the missing state artifacts as a blocker. Never move them into + tracked workspace content or claim that the full workflow completed. + +Pass these limits, known unavailable capabilities, exact paths, and commands +to any delegated agent. A child reporting a permission or toolchain blocker +does not justify launching another child for the same operation. + +When shell execution is unavailable, review the recorded file edits and +permitted file-tool output instead of running `git status` for the final +working-tree review. That review does not authorize another denied command. ## Pipeline Overview @@ -42,7 +87,10 @@ proceed. If the user provides no details or a very basic prompt (e.g., [unit-test-generation.prompt.md](../skills/code-testing-agent/unit-test-generation.prompt.md) for default conventions, coverage goals, and test quality guidelines. -Before writing code, read the language-specific base extension. Reuse it for the whole run; sub-agents must not independently reload the same reference unless they need a section that was not captured in the research document. +Before writing code, use the available language-specific base extension or +discover conventions from the project's manifests and representative tests. +Reuse the findings for the whole run; sub-agents must not independently reload +the same reference unless a required section was not captured in research. For Single pass and Iterative strategies, resolve one absolute `` @@ -56,10 +104,11 @@ before invoking any sub-agent: 3. Outside Git, create a unique directory under the operating system's temporary directory. -Create the resolved directory and pass its absolute path explicitly in every -sub-agent prompt. Never create `` or any intermediate state file -in version-controlled workspace content, and never modify `.gitignore` to hide -them. +Create the resolved directory using permitted tools and pass its absolute path +explicitly in every sub-agent prompt. If resolution or creation is denied, +apply the in-context fallback above instead of probing alternative locations. +Never create intermediate state files in version-controlled workspace content +or modify `.gitignore` to hide them. Create a **requirement checklist** from the request before choosing a strategy. Preserve each explicit behavior, layer, collaborator seam, boundary case, @@ -82,16 +131,17 @@ Based on the request scope, pick exactly one strategy and follow it: | Strategy | When to use | What to do | | ---------- | ------------- | ------------ | | **Direct** | A small, self-contained request (e.g., tests for a single function or class) that you can complete without sub-agents | Follow the codebase conventions on test file structure, naming, style, and testing approaches. Reuse existing test projects and test files when possible — if the code under test already has tests, add new tests to the same file or test project. Only create a new test file when no canonical file is named or discoverable for the symbol under test. Write the tests immediately. **Run them right away** — if any test fails, read the production code, fix the assertion, and re-run before writing more tests. Skip Steps 3-5 (research, plan, implement sub-agents), then perform proportionate validation and reporting in Steps 6-9. | -| **Single pass** | A moderate scope (couple projects or modules) that a single Research → Plan → Implement cycle can cover | Execute Steps 3-8 once, then proceed to Step 9. | -| **Iterative** | A large scope or ambitious coverage target that one pass cannot satisfy | Execute Steps 3-8, then re-evaluate coverage. If the target is not met, repeat Steps 3-8 with a narrowed focus on remaining gaps. Use unique names for each iteration's documents in `` (e.g., `research-2.md`, `plan-2.md`) so earlier results are not overwritten. Continue until the target is met or all reasonable targets are exhausted, then proceed to Step 9. | +| **Single pass** | A project/package-wide request or a moderate set of modules that fits one context | Execute Steps 3-8 once, keeping the phases inline unless substantial separate work warrants delegation, then proceed to Step 9. | +| **Iterative** | A large scope or measured coverage gaps that one pass cannot satisfy | Execute Steps 3-8, then extend the existing inventory and plan only for concrete remaining gaps. Do not restart discovery or orchestration. Preserve earlier evidence in `` and proceed to Step 9 when the bounded target is met or a concrete blocker remains. | **Default to Direct** unless the user asks for a project/package-wide suite or the scope explicitly spans multiple files or modules. Most test generation requests — including "generate tests for function X", "add tests covering these scenarios", and "write unit tests for this class" — should use Direct strategy. A project-wide request remains Single pass even when the delivered workspace is -sparse and only one source module remains. Choosing Direct trades away only the -sub-agent pipeline, not verification. When a request enumerates specific behaviors/scenarios +sparse and only one source module remains; it needs the inventory, plan, and +status artifacts, not mandatory phase agents. Choosing Direct trades away only +those artifacts, not verification. When a request enumerates specific behaviors/scenarios (e.g., "add 1 test for each of these scenarios"), treat that list as the spec: target the exact symbol named, cover every enumerated scenario, and perform the Step 7 requirement-coverage check before reporting completion. @@ -105,7 +155,7 @@ Step 7 requirement-coverage check before reporting completion. | "Achieve 80% coverage across the whole solution" | Iterative | Large scope, first pass covers the obvious gaps, subsequent passes target remaining uncovered code | | "Add tests for this function" (with file open) | Direct | Single function is trivially small scope | | "Generate comprehensive tests for my ASP.NET app" | Single pass | If the app has fewer than 10 controllers/services/files in scope, one R→P→I cycle should cover it | -| "Generate comprehensive tests for my large ASP.NET app" | Iterative | If the app has 10 or more controllers/services/files in scope, use repeated passes to close remaining gaps | +| "Generate comprehensive tests for my large ASP.NET app" | Iterative | Use targeted follow-up phases for measured gaps that cannot fit one pass; file count alone does not justify repeated discovery | **All strategies execute Steps 6-9**, but validation depth must match the requested scope. Focused Direct work validates the affected project/tests; @@ -114,35 +164,58 @@ during research. ### Step 3: Research Phase -Delegate to the `code-testing-researcher` subagent with this task: +Research the requested scope once. Batch independent manifest, source, and +representative-test reads; do not inventory unrelated files. Record: -```text -runSubagent({ - agent: "code-testing-researcher", - prompt: "Research [REQUESTED SCOPE] at [PATH] for test generation. Write the research document to /research.md. Produce a bounded target inventory, existing test conventions, source-to-test pairs, dependencies only for those targets, and exact build/test/discovery commands. Do not inventory unrelated source files." -}) -``` +- the requirement checklist and bounded public API/behavior inventory; +- source-to-test pairs, canonical test paths, conventions, and pinned APIs; +- dependencies and fake/mock seams for those targets; +- exact build/test/discovery commands and requested coverage thresholds; +- capability or validation blockers already observed. + +Use a deterministic pairing skill only when available and useful; a small +explicit target list does not need a second discovery pass. Delegate substantial +research to `code-testing-researcher` only when its separate context is useful. Output: `/research.md` ### Step 4: Planning Phase -Delegate to the `code-testing-planner` subagent with this task: - -> Create a test implementation plan based on `/research.md`. Write it to `/plan.md`. Create a phased approach with specific files and test cases. +Map the research checklist to concrete test names, inputs, assertions, and +files in `/plan.md`. Group collaborating targets into coherent +implementation phases rather than one agent per file. Plan inline for a bounded +suite; use `code-testing-planner` only when the planning work itself needs +separate context. Output: `/plan.md` ### Step 5: Implementation Phase -Execute each phase by delegating to the `code-testing-implementer` subagent — once per phase, sequentially. For each phase, delegate with this task: - -> Implement Phase N from `/plan.md`: [phase description]. Use `/research.md` for commands and conventions. Ensure tests compile and pass. +Implement each phase sequentially, inline by default. Read the complete target +logic before choosing expected values. For composed operations, derive the +intermediate values in source order; do not substitute a familiar domain formula. +When two modes or branches differ, choose inputs that actually distinguish their +results instead of merely executing both with equivalent expectations. +Preserve production code, existing tests, +project format, and dependency versions; make only required test-registration +or missing-dependency edits that the request allows. For classic .NET projects, +preserve `packages.config`, fixtures, and explicit compile items, and register +each new test file exactly once. Use APIs compatible with the pinned versions. + +For a substantial implementation phase, delegate once to an available +`code-testing-implementer` with the relevant plan, source/test paths, conventions, +commands, edit boundaries, and known blockers. Consume its report; do not repeat +its discovery or launch builder/tester agents just to repeat its validation. ### Step 6: Final Build Validation -Run the narrowest build that covers all changed test projects and their source -dependencies. For Single pass or Iterative work spanning multiple projects, +Use the narrowest command that compiles all changed tests and their source +dependencies. A fresh-build test command can satisfy both build and test gates; +reuse it only if the runner compiles/type-checks the changed tests. Transpilation +alone is not a TypeScript type check: use the existing typecheck command or +installed `tsc --noEmit` and confirm the config includes generated tests. +Do not run a separate build when it adds no evidence. For Single pass or +Iterative work spanning multiple projects, new project registration, or solution manifests, run the bounded workspace build recorded during research. Do not replace a classic non-SDK build with `dotnet build`. @@ -153,10 +226,12 @@ build recorded during research. Do not replace a classic non-SDK build with - **Go**: `go build ./...` from module root - **Rust**: `cargo build` -If it fails, call `code-testing-fixer`, rebuild, and retry at most three times. -Stop earlier when a diagnostic repeats without measurable progress, an external -blocker is concrete, or the remaining fix would exceed the requested edit -scope. +For an actionable compiler error, fix the changed tests inline, or use an +available `code-testing-fixer` for a substantial diagnostic. Rebuild only after +a concrete fix, at most three times. Stop when a diagnostic repeats without +progress, a permission/toolchain blocker is concrete, or the fix would exceed +the requested edit scope. Do not install dependencies unless a missing-package +diagnostic or an allowed manifest change requires it. ### Step 7: Final Test Validation @@ -169,8 +244,16 @@ Run tests at the same proportionate scope selected in Step 6 with a fresh build evidence supports that attribution. Do not modify unrelated tests, but a nonzero required final test command still blocks a success verdict. +Reuse successful validation for unchanged files at the same scope. If a test +command also proves discovery or collects the requested coverage, use that +evidence instead of running separate agents or redundant commands. Confirm +new files are actually discovered; in a classic project, inspect registration +as well as test output. A zero-test run does not validate generated tests. + +Apply the shared report-safe naming and result-validation contract before +accepting a passing run, including configured report export and artifact parsing. Do not continue to the success report while required final validation is -failing. If an out-of-scope or pre-existing failure remains, report +failing or unrun. If an out-of-scope or pre-existing failure remains, report `PARTIAL`/blocked with the exact command and failure evidence; never describe the generated suite or pipeline as successfully validated. @@ -180,14 +263,16 @@ Always map explicit prompt requirements to the final tests and inspect the final diff for concrete, behavior-pinning assertions. For broad/comprehensive work, coverage-quality requests, multi-file additions, at least five generated tests, or a prompt that enumerates scenarios, boundaries, error paths, or interactions, -also run the two plugin skill checks below before reporting completion and after -any Step 8 iteration. The manual prompt-scenario and assertion review is +also use each available plugin skill check below once before completion. +If a skill is unavailable, perform its described review inline. After fixes, +review the affected behaviors without reloading the skills or repeating the +entire audit. The manual prompt-scenario and assertion review is sufficient only for a focused addition under five tests with no enumerated behavior. -1. **Pseudo-mutation check** — invoke the `test-gap-analysis` skill against the source file(s) you tested and the test file(s) you produced. The skill reasons about plausible mutations (boundary flips, dropped null checks, removed exceptions, sign flips) and reports which would slip past your tests. For every gap it flags, either strengthen the existing assertion or add a follow-up test. Re-run until no gap is reported, or until the remaining gaps are explicitly out of scope (e.g., production bugs you cannot fix in a test-only PR). +1. **Pseudo-mutation check** — use `test-gap-analysis` when available against the tested sources and generated tests. Check plausible boundary flips, dropped validation, removed exceptions, and sign changes. For each in-scope gap, strengthen the assertion or add a test, then check that specific mutation against the revised test. Record out-of-scope gaps instead of restarting the audit. -2. **Assertion-depth check** — invoke the `assertion-quality` skill against the test file(s) you produced. If it flags trivial-only assertions (`IsNotNull` / `toBeDefined` / `assert x is not None`-only tests, tautological round-trip assertions, single-observable tests where the production code touches multiple observables), revise those tests — replace existence checks with concrete-value assertions, and add a secondary observable per behavior-radius guidance. +2. **Assertion-depth check** — use `assertion-quality` when available against the generated tests. Replace existence-only assertions (`IsNotNull` / `toBeDefined` / `assert x is not None`) and tautological round trips with concrete behavior assertions. Add a secondary observable only when it is part of the public contract or required to prove a requested interaction; do not couple tests to incidental state, logs, or call counts. @@ -197,10 +282,8 @@ behavior. - **Cover the full range each scenario's wording implies, not a single representative case.** Phrasing like "when the dimensions stay the same *or* change", "wider *or* narrower", or "first character *or* anywhere in the string" calls for multiple variations — exercise each variation (and combine them in one test when the wording groups them) rather than asserting a single instance. - **Honor positional and structural qualifiers literally.** When a scenario pins a condition to a specific position or shape (e.g. "the *first* character after the prefix", "a filename containing a literal space"), construct an input that satisfies that exact qualifier — an input where the condition merely appears *somewhere* does not cover it. -Never skip the requirement mapping or concrete-assertion review. Omit the two -additional skill invocations only for focused additions under five tests that -have no enumerated scenarios, boundaries, error paths, or interactions and do -not request broader quality or coverage analysis. +Never skip requirement mapping, mutation thinking, or concrete-assertion review. +Unavailable supporting skills change the review mechanism, not its depth. Additional self-review heuristics (still required, even when running the skills): @@ -241,41 +324,32 @@ state files. ### Step 9: Report Results -Lead with the outcome. Summarize tests created, validation actually run, any -failures or issues, and include a compact -**Requirement coverage** section that maps each explicit request to the test -file or test group that satisfies it. Name concrete evidence such as the mock -or fake used, fixed inputs and expected values, boundary combinations, -in-memory integration fixture, and generated coverage artifact. Do not report -a requirement as covered based only on aggregate coverage. +Lead with `SUCCESS` only when all required validation passed; otherwise use +`PARTIAL` or `BLOCKED`. Distinguish implemented tests, static review, executed +tests, and measured coverage. Give the exact command and diagnostic for unrun +or failed validation; do not infer threshold clearance from configuration. -**Example final report:** +For broad requests, include a compact `Requirement | Evidence` table. Cite +exact test names and paths for each requested behavior; cite the file, command, +or report for non-behavioral requirements. Include meaningful fake interactions, +inputs, expected values, and before/at/after cases where needed. Do not replace +this mapping with a generic list of tested modules or aggregate coverage. -``` -## Test Generation Report - -**Project**: MyProject -**Strategy**: Single pass - -### Results -| Metric | Value | -|----------------|-------| -| Tests created | 24 | -| Tests passing | 24 | -| Tests failing | 0 | -| Files created | 3 | - -### Files Created -- tests/MyProject.Tests/ServiceATests.cs (10 tests) -- tests/MyProject.Tests/ServiceBTests.cs (8 tests) -- tests/MyProject.Tests/HelperTests.cs (6 tests) - -### Build Validation -- Scoped build: ✅ passed -- Bounded workspace build: ✅ passed - -### Next Steps -- Consider adding integration tests for database layer +Before ending the turn, check that the final response itself contains +`| Requirement | Evidence |` and exact test names for every behavioral row. +An internal plan or a differently labeled coverage table does not satisfy +the handoff contract. + +```text +PARTIAL — implemented the requested tests; execution was denied. + +| Requirement | Evidence | +| --- | --- | +| Reject invalid discounts | tests/test_pricing.py::test_negative_discount_rejected asserts ValueError | +| Preserve the exact threshold | tests/test_pricing.py::test_discount_at_threshold asserts 90.00 | + +Validation: `python -m pytest -q` was denied by the host; test passage and +coverage are unverified. No alternate-shell or delegated retry was attempted. ``` Use a language example from `code-testing-extensions` only when no existing tests establish a usable convention. Never load examples merely to confirm a pattern already present in the repository. @@ -304,7 +378,7 @@ non-stageable ``: 7. **No environment-dependent tests** — mock all external dependencies; never call external URLs, bind ports, or depend on timing 8. **Fix assertions, don't skip tests** — when tests fail, read production code and fix the expected value; never `[Ignore]` or `[Skip]` 9. **Keep intermediate state files out of commits** — retain research, plan, and final status in `` through completion, but never place `` or its files in version-controlled workspace content, stage them, or modify `.gitignore` to hide them. Before reporting, inspect the working-tree changes and confirm they contain only requested deliverables and required manifest edits. -10. **Read language extensions first** — always call the `code-testing-extensions` skill and read the relevant extension file before writing any code; it contains critical project registration and build validation steps +10. **Use available language guidance** — load the base extension once when available; otherwise derive registration, APIs, and commands from the repository without retrying missing skills 11. **Validate proportionately** — final build, tests, requirement review, and reporting are mandatory for every strategy; use the Step 7 skill checks only at the thresholds defined there @@ -317,6 +391,7 @@ non-stageable ``: Do not stop after analysis or planning when test implementation was requested. Finish when every feasible requirement is mapped to concrete tests, the proportionate build and test commands pass, applicable quality checks are -complete, and the final working-tree review contains only requested test and +complete, the shared report-safe naming and result-validation contract is met, +and the final working-tree review contains only requested test and minimal registration/dependency changes. If blocked, report the exact command, evidence, and remaining bounded work without claiming success. diff --git a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-implementer.agent.md b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-implementer.agent.md index 0f893643..0bc2201f 100644 --- a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-implementer.agent.md +++ b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-implementer.agent.md @@ -22,7 +22,14 @@ You implement a single phase from the test plan. You are polyglot — you work w > **Language-specific guidance**: Reuse the guidance captured in research. > Call `code-testing-extensions` only when the required implementation or -> harness-discovery section is missing. +> harness-discovery section is missing and the skill is available. + +Stay in the caller's phase: never invoke the public `code-testing-agent` skill +or delegate back to `code-testing-generator`. Use supplied guidance and known +paths instead. Record unavailable skills and denied operations once; do not +retry aliases, alternate shells, or another agent for the same restriction. +Continue permitted test edits and static review when execution is blocked, +then return `PARTIAL` with the exact blocker and unverified requirements. ## Your Mission @@ -47,6 +54,9 @@ For each file in your phase: the available tools support it. - Understand the public API — verify exact parameter types, count, return types, and **actual return values for key inputs** before writing assertions - **Trace the logic** for each code path you plan to test — understand what the function actually does, not what you think it should do +- Derive composed results from intermediate values in source order. Use inputs + that distinguish requested modes/branches; do not rely on a familiar domain + formula or equal-output cases that cannot detect a mode-selection bug. - Note dependencies and how to mock them - **Validate project references**: Read the test project file and verify it references the source project(s) you'll test. Add missing references before creating test files - **Validate project-system registration**: For classic non-SDK C# projects, every new test file must be added exactly once as a relative ``. For SDK-style projects, confirm default compile globs are enabled before relying on implicit inclusion. @@ -91,21 +101,30 @@ These rules apply to every language and override any pattern an existing test fi Coverage alone gives false confidence — every test must *pin down behavior* so it would fail under a plausible bug. Apply the `code-testing-agent` skill's `unit-test-generation.prompt.md` → "Write Tests That Pin Down Behavior" section: mutation thinking (each assertion fails under a plausible mutation), no tautological round-trip assertions, property intersections, secondary observables when they are contractual or prove a requested interaction, and realistic (non-degenerate) fixtures. This is a depth requirement on top of the happy/edge/error-path and mocking rules above, and applies to every language. +Also apply [Report-safe test names and result validation](../skills/code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation) +when naming cases and accepting test results. Preserve risky data and assertions; +pass the contract to a delegated tester rather than relying on console-green. + ### 5. Verify with Build -Call the `code-testing-builder` sub-agent to compile, passing the exact build -command and absolute ``. Build only the specific test project, -not the full solution. +Run the supplied scoped build command directly. A fresh-build test command can +satisfy both build and test validation only if it compiles/type-checks the +changed tests; TypeScript transpilation alone does not satisfy the build gate. +Use an available `code-testing-builder` +only for substantial separate work, passing the exact command, absolute +``, and known capability limits. -If build fails, call `code-testing-fixer`, rebuild, and retry at most three -times. Stop earlier when a diagnostic repeats without measurable progress, an -external blocker is concrete, or the remaining fix would violate the edit -boundaries. +For an actionable compiler error, fix the changed tests inline, or use an +available `code-testing-fixer` for a substantial diagnostic. Rebuild only after +a concrete fix, at most three times. Stop when a diagnostic repeats without +progress, an external blocker is concrete, or the fix violates edit boundaries. ### 6. Verify with Tests -Call the `code-testing-tester` sub-agent to run tests, passing the exact test -command and absolute ``. +Run the supplied scoped test command directly, or use an available +`code-testing-tester` when separate context is useful. Reuse an unchanged +passing run; do not delegate merely to repeat it. Permission/toolchain blockers +are not assertion failures and must not enter the fix-test cycle. If tests fail: @@ -136,8 +155,12 @@ If your language extension has no "Harness Discovery Check" section, use the can ### 8. Format Code (Optional) -If a lint command is available, call the `code-testing-linter` sub-agent, -passing the exact lint command and absolute ``. +When formatting or linting is needed, run the repository's existing command +directly for the changed tests, using the conventions captured in research. +Use an available `code-testing-linter` only for substantial work that benefits +from separate context, passing the exact command, changed-file scope, absolute +``, and known capability limits. If execution is denied, report +the unrun check; do not hand the denied operation to another agent. ### 9. Report Results @@ -171,5 +194,6 @@ Consult a language example only when the repository has no representative tests The phase is complete only when all planned in-scope tests are implemented, the scoped build and tests pass, and harness-equivalent discovery sees the expected -new tests. If an external blocker prevents that, stop with `PARTIAL` or +new tests. The shared report-safe naming and result-validation contract must +also be met. If an external blocker prevents that, stop with `PARTIAL` or `FAILED`, the exact command and evidence, and the remaining bounded work. diff --git a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-tester.agent.md b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-tester.agent.md index bc6fd86f..4fcf6108 100644 --- a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-tester.agent.md +++ b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/agents/code-testing-tester.agent.md @@ -23,6 +23,10 @@ You run tests and report the results. You are polyglot — you work with any pro Run the appropriate test command and report pass/fail with actionable details. Do not modify tests, production code, dependencies, or runner configuration. +Apply [Report-safe test names and result validation](../skills/code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation) +before reporting passage. Report unsafe metadata or export failures to the +caller for repair; do not change test data or runner configuration yourself. + ## Process ### 1. Discover Test Command @@ -58,6 +62,8 @@ For scoped tests (if specific files are mentioned): ### 3. Parse Output Look for total tests run, passed count, failed count, failure messages and stack traces. +Include skipped cases and incomplete/setup failures. Apply the shared contract +to configured result artifacts; record their paths and parsing outcome. ### 4. Return Result @@ -106,5 +112,7 @@ Failures: ## Completion Condition Stop when the requested test process has completed and the summary and relevant -failures have been captured. This agent reports evidence; it does not fix the -failures. +failures have been captured, including runner exit code and any required +report-export/parsing result under the shared report-safe naming and result-validation contract. +Report export failure as failed validation even if assertions passed. +This agent reports evidence; it does not fix the failures. diff --git a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-agent/SKILL.md b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-agent/SKILL.md index 887b2b4a..31d6cd94 100644 --- a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-agent/SKILL.md +++ b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-agent/SKILL.md @@ -11,16 +11,25 @@ description: >- running/diagnosing tests, coverage/audits, a test blocked on a missing production seam (testability-obstacle), or correcting supplied MSTest assertions, attributes, lifecycle, or configuration without designing new - cases (writing-mstest-tests). + cases (writing-mstest-tests). Within an active code-testing-generator + pipeline, reuse supplied guidance; do not re-enter this skill. license: MIT --- # Code Testing Generation Skill -An AI-powered skill that generates comprehensive, workable unit tests for any programming language using a coordinated multi-agent pipeline. +Generate comprehensive, workable unit tests for any programming language using +a bounded Research → Plan → Implement workflow. ## Non-negotiable execution contract +**Check pipeline ownership first.** If the active agent is +`code-testing-generator` (including a plugin-qualified name such as +`dotnet-test:code-testing-generator`), or the caller assigned you a phase of +that pipeline, do not delegate to another generator. Continue the assigned +work inline. This guard takes precedence over every broad-scope delegation +instruction below, even if this skill was loaded automatically. + Classify scope **before editing**: - **Broad** (a project/package-wide suite, or multiple production @@ -37,12 +46,21 @@ Classify scope **before editing**: module is present. For either scope, run the narrowest relevant test command to a clean exit. +Always apply [Report-safe test names and result validation](unit-test-generation.prompt.md#report-safe-test-names-and-result-validation), +including when the caller supplies conventions. Pass this contract to delegated +implementers/testers; preserve edge-case data and validate configured reports, +not just console output. Keep the handoff proportional: for one to three focused requirements, use a compact bullet list under a **Requirement coverage** label that names the tests and successful command; for broader or multi-requirement work, use a `Requirement | Evidence` table. Each requested behavior must cite an exact test name. +Before sending a broad-scope final response, check that the response itself +contains `| Requirement | Evidence |` and exact test names for every behavioral +row. A table in a child report or internal plan is not enough. Do not summarize +away those names into module-level bullets or an `Area | Tests` table. + Intermediate state files are internal working data, never deliverables. Keep `` non-stageable, never place it or its files in version-controlled workspace content, and never modify `.gitignore` to hide @@ -59,6 +77,17 @@ prove composition, but do not substitute for the requested module-level coverage. Judge breadth by the behavior matrix, never by matching or exceeding a raw test count. +At the public entry point, delegate broad work to `code-testing-generator` +once. Research, plan, implementation, and review remain required, but they +need not be separate sub-agent calls. + +Use only capabilities available in the current runtime. Do not retry a missing +skill under aliases or use another agent to retry a policy-denied operation. +If scratch storage is denied, keep the research and plan in context, continue +permitted test edits, and report the missing state artifacts. If execution is +denied, continue permitted static review and report tests as unrun, never passed. +Neither blocker authorizes modifying production code or weakening requirements. + For a **broad or comprehensive** request, the explicit matrix is the floor, not the ceiling. Treat each requested module or layer as an inventory heading, not one behavior: expand it into the bounded public operations and their distinct @@ -165,10 +194,16 @@ generation request. Do not use a generic/general-purpose subagent merely named `code-testing-generator`: ```text -Generate unit tests for [path or description of what to test], following the [unit-test-generation.prompt.md](unit-test-generation.prompt.md) guidelines. Treat the current workspace as authoritative even when it is sparse, gutted-looking, synthetic, or missing tracked files; never restore or reconstruct it, including with `git checkout`, `git restore`, `git reset`, or `git clean`. +You are the sole pipeline owner for this request. Do not invoke code-testing-agent or another code-testing-generator; complete the phases in your current context. Generate unit tests for [path or description of what to test], following the [unit-test-generation.prompt.md](unit-test-generation.prompt.md) guidelines. Treat the current workspace as authoritative even when it is sparse, gutted-looking, synthetic, or missing tracked files; never restore or reconstruct it, including with `git checkout`, `git restore`, `git reset`, or `git clean`. ``` -The Test Generator will manage the entire pipeline automatically. +The Test Generator owns the pipeline. After it returns, consume its recorded +quality checks, validation results, and requirement matrix instead of repeating +Steps 4 and 5 as another pipeline. Do not reload review skills or rerun unchanged +passing commands. Preserve exact test names from its evidence in the final +handoff. If evidence is missing, inspect or follow up on that specific gap +without restarting generation. A reported capability-wide denial also applies +to the caller; do not attempt another command using that capability. If `code-testing-generator` is unavailable, do not skip the workflow. Execute the same Research → Plan → Implement sequence inline, resolve `` as @@ -196,7 +231,7 @@ For multi-file requests: 1. Turn every explicit user requirement into a checklist before implementation. Include requested layers, collaborators to mock, boundary cases, integrations, coverage thresholds, and report artifacts. Copy multi-condition requirements verbatim — they must each map to one test that exercises the whole combination. 2. Research only the requested module or project and write the checklist plus a compact target inventory to `/research.md`. 3. Reuse manifests, symbol references, and deterministic pairing tools instead of reading every source and test file. -4. For multi-file scopes in C#, Python, TypeScript/JavaScript, Go, Java, Rust, Ruby, Kotlin, Swift, PowerShell, or C++, run `find-untested-sources` once and consume its pairing and suggested-path output; do not repeat that discovery manually. +4. When an available `find-untested-sources` skill is useful for a substantial multi-file inventory, run it once and reuse its pairing and suggested-path output. Otherwise pair the bounded targets manually once; do not probe for an unavailable skill. 5. Plan each target file once, then implement phases sequentially. Map every checklist item to at least one concrete test or explain why it is blocked. 6. Build and test the narrow target during fix cycles. Run workspace-level validation once at the end only for broad work, when the repository contract @@ -220,7 +255,8 @@ Do not report completion until all of these are true: inventory, existing test conventions, and the acceptance checklist. 2. *(broad scope)* `/plan.md` maps each checklist item to a planned test or an explicit blocker. -3. Generated tests compile and pass with the narrowest relevant test command. +3. Generated tests compile and pass with the narrowest relevant test command, + satisfying the shared report-safe naming and result-validation contract. 4. Every explicit user requirement is backed by a concrete test and assertion. Fix missing mock seams, boundary cases, state transitions, and property combinations even when coverage already passes. In the final summary, cite diff --git a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-agent/unit-test-generation.prompt.md b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-agent/unit-test-generation.prompt.md index 43a93a58..a890d67c 100644 --- a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-agent/unit-test-generation.prompt.md +++ b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-agent/unit-test-generation.prompt.md @@ -93,6 +93,48 @@ Quick self-review before finishing a test: would emptying the function body make - Combine logically related test cases into a single parameterized method - Never generate multiple tests with identical logic that differ only by input values +## Report-safe test names and result validation + +Apply this contract to both direct generation and delegated implementation or +validation, even when the caller supplies its own test style. + +- **Separate metadata from data.** Give each case a stable, descriptive, + distinguishable ID/display name. Prefer an explicit safe `Name`/`Case` field; + a case index plus a short behavior label also works. Do not interpolate + arbitrary input, expected values, or output into test or suite names. +- Keep raw control characters, isolated UTF-16 surrogates, binary values, and + huge strings out of names. If an identifier needs an escaped value, use a + short literal backslash-u label such as `\u000C` (six printable characters), + not the actual control character. Normal Unicode labels and harmless numeric + interpolation are fine. +- Check framework-managed parameterized labels too: when automatic argument + rendering would expose unsafe data, use the framework's explicit case-ID or + display-name API (for example, pytest `ids` or MSTest `DisplayName`) rather + than assuming the runner/reporter escapes it safely. +- **Preserve the case.** Control characters and malformed strings are legitimate + test data. Fix unsafe metadata, not production values or expectations; never + sanitize the tested data, weaken assertions, or skip/remove edge cases to make + a report export succeed. + +Before claiming tests passed: + +1. Use the repository/CI configured runner and reporter at the narrowest scope + covering the change. Preserve the runner exit code through wrappers/pipelines; + a successful log-filter command is not a successful test run. +2. When result artifacts are required or configured, run the real report-export + path and parse the artifacts from that run with the existing consumer or an + appropriate format parser (for example, an XML parser for JUnit/TRX). + Console-green alone is insufficient if required export failed. +3. Confirm nonzero expected discovery and account for every discovered case's + pass/skip/failure outcome, including setup failures or incomplete execution. + Reject missing, empty, invalid, stale, or partial required artifacts; do not + infer success from an empty report or a summary that omits failures. +4. Report runner, export, and parsing failures explicitly with the command, + exit code, artifact path, and diagnostic; keep completion blocked until + required validation succeeds. Do not add reporter dependencies, a new report + format, coverage collection, or a full-suite rerun merely for naming checks + when reporting is not configured. + ## Analysis Before Generation Do this analysis privately; do not emit a plan or inventory unless the user diff --git a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-extensions/extensions/powershell.md b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-extensions/extensions/powershell.md index f8014f90..25df6742 100644 --- a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-extensions/extensions/powershell.md +++ b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-extensions/extensions/powershell.md @@ -120,6 +120,51 @@ Pester v5 runs in **two phases**: Discovery (collects test metadata) then Run (e - Use `foreach` loops for dynamic test generation only with `BeforeDiscovery` data - Use `TestDrive:` for file-based tests instead of touching repo files — Pester cleans it up automatically +## Parameterized Test Display Names + +Apply [Report-safe test names and result validation](../../code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation). +Use an explicit safe `Name`/`Case` in `-ForEach` or `-TestCases` data and expand +only that field in the `It` title. Do not expand arbitrary `` or +`` values into discovery/report metadata. + +For a function whose contract reverses UTF-16 code units (not Unicode scalars), +the reversed supplementary character is intentionally malformed UTF-16. Keep +that expected value in the assertion, not the title: + +Use `Text` for the data field, not `Input`: `$Input` is PowerShell's automatic +pipeline-input variable and can hide the intended case value inside `It`. + +```powershell +BeforeDiscovery { + $cases = @( + @{ + Name = 'supplementary code-unit reversal' + Text = [string]::Concat([char]0xD83D, [char]0xDE00) + Expected = [string]::Concat([char]0xDE00, [char]0xD83D) + } + @{ + Name = 'isolated high surrogate' + Text = [string][char]0xD800 + Expected = [string][char]0xD800 + } + ) +} + +Describe 'Get-Reversed' { + BeforeAll { + Import-Module (Join-Path $PSScriptRoot '../tools/StringUtils.psm1') -Force + } + + It 'reverses ' -ForEach $cases { + Get-Reversed -Value $Text | Should -BeExactly $Expected + } +} +``` + +Reuse the repository's configured Pester `TestResult` export path and format +when present and parse the resulting artifact before reporting success. A +passing `TotalCount`/`PassedCount` does not prove JUnit export succeeded. + ## Common Errors | Error | Fix | diff --git a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-extensions/extensions/typescript.md b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-extensions/extensions/typescript.md index de26a70a..50375f28 100644 --- a/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-extensions/extensions/typescript.md +++ b/external-sources/upstreams/dotnet-skills/plugins/dotnet-test/skills/code-testing-extensions/extensions/typescript.md @@ -76,6 +76,30 @@ Use the repo's lint script first. Otherwise detect from `devDependencies` and co - Jest/Vitest default: `*.test.ts`, `*.spec.ts`, or files inside `__tests__/` - Place test files to mirror the existing project pattern +## Parameterized Test Display Names + +Apply [Report-safe test names and result validation](../../code-testing-agent/unit-test-generation.prompt.md#report-safe-test-names-and-result-validation). +For Jest `it.each`/`test.each`, interpolate only a safe label (`$Name` for object +rows), or use a short behavior label with `%#` for the case index. Do not use +`%p`, `%s`, or `$Input`/`$Expected` to render arbitrary data in the title. +Harmless numeric placeholders such as `%i` remain fine. + +For a word counter whose contract treats form-feed as whitespace, keep the real +U+000C input but give the case a safe name: + +```typescript +it.each([ + { Name: "form-feed separator", Input: "one\ftwo", Expected: 2 }, + { Name: "space separator", Input: "one two", Expected: 2 }, +])("countWords: $Name", ({ Input, Expected }) => { + expect(countWords(Input)).toBe(Expected); +}); +``` + +The escape in `Input` becomes an actual control character at runtime; the title +does not contain it. Reuse the configured reporter and validate its exported +artifact when required/configured; do not install `jest-junit` just for this check. + ## Common Errors | Error | Fix | diff --git a/external-sources/vendir.lock.yml b/external-sources/vendir.lock.yml index 0f3dab2b..9b761de6 100644 --- a/external-sources/vendir.lock.yml +++ b/external-sources/vendir.lock.yml @@ -2,20 +2,20 @@ apiVersion: vendir.k14s.io/v1alpha1 directories: - contents: - git: - commitTitle: 'Merge pull request #1240 from dotnet/abhitejjohn-eval-authoring-quality-bar...' - sha: 8d670fa76aaac45b336d8ded05a7601785fb2121 + commitTitle: Teach test generation to use report-safe case names (#1275)... + sha: a660de83c2f9cad071a5d568f7d77041ccd07059 tags: - - skill-validator-nightly-10-g8d670fa7 + - skill-validator-nightly-3-ga660de83 path: dotnet-skills - git: commitTitle: Add Cursor rules that reference the existing skill docs... sha: af2319bd01bb7cc881267a9ef42cafdaf5e9029d path: webgpu-claude-skill - git: - commitTitle: 'fix(assets): return requested format from Sharp transform (#18222)...' - sha: fcf6ed6d4915eed6c18a20658b73372d074c4832 + commitTitle: 'fix(astro): provide Astro.cache when rendering error pages (#18293)...' + sha: 28e74103fe3d999a56ae3404b5fe67a1ad6d8013 tags: - - '@astrojs/cloudflare@14.3.4-7-gfcf6ed6d49' + - '@astrojs/language-server@2.17.2-2-g28e74103fe' path: astro path: upstreams kind: LockConfig