From 1e7ed7575f8422fd4279e94128fc8dba5d703a87 Mon Sep 17 00:00:00 2001 From: "github-actions[bot]" <41898282+github-actions[bot]@users.noreply.github.com> Date: Wed, 7 Oct 2026 06:04:15 +0000 Subject: [PATCH] chore: refresh upstream skill catalog --- .../skills/astro-developer/manifest.json | 2 +- .../skills/create-skill-test/SKILL.md | 97 ++++++++++++++++--- .../skills/improve-skill-quality/SKILL.md | 34 ++++++- .../references/eval-triage.md | 10 +- .../astro/packages/astro/package.json | 2 +- .../.agents/skills/create-skill-test/SKILL.md | 97 ++++++++++++++++--- .../skills/improve-skill-quality/SKILL.md | 34 ++++++- .../references/eval-triage.md | 10 +- external-sources/vendir.lock.yml | 13 ++- 9 files changed, 252 insertions(+), 47 deletions(-) diff --git a/catalog/Frameworks/Official-Astro/skills/astro-developer/manifest.json b/catalog/Frameworks/Official-Astro/skills/astro-developer/manifest.json index de8cfdd..d3d3d33 100644 --- a/catalog/Frameworks/Official-Astro/skills/astro-developer/manifest.json +++ b/catalog/Frameworks/Official-Astro/skills/astro-developer/manifest.json @@ -1,5 +1,5 @@ { - "version": "7.3.5", + "version": "7.3.6", "category": "Web", "compatibility": "Requires the withastro/astro monorepo." } diff --git a/catalog/Tools/Official-DotNet-Create-Skill-Test/skills/create-skill-test/SKILL.md b/catalog/Tools/Official-DotNet-Create-Skill-Test/skills/create-skill-test/SKILL.md index e1f4f30..b5eaa5b 100644 --- a/catalog/Tools/Official-DotNet-Create-Skill-Test/skills/create-skill-test/SKILL.md +++ b/catalog/Tools/Official-DotNet-Create-Skill-Test/skills/create-skill-test/SKILL.md @@ -30,11 +30,33 @@ and does not overfit to the skill's own wording. | Skill or agent name | Yes | Must exist under `plugins//skills/` or `plugins//agents/` | | Plugin name | Yes | e.g. `dotnet-msbuild` | | Skill content | Yes | Read it — you cannot write non-overfitted rubric items without it | -| Failure modes to discriminate | Recommended | Each becomes one stimulus | +| Scenario hypothesis | Yes | State the expected improvement for a preference case, or the invariant protected by a guard | +| Failure modes to discriminate | Recommended | Each becomes one distinct stimulus | ## Workflow -### Step 1: Locate the target and the test directory +### Step 1: Prove the eval belongs to the target + +Before writing YAML, state what the stimulus proves. A preference stimulus is necessary when the +target should improve the answer, action, restraint, or validation result compared with the same +model without the target. A non-voting activation contract or no-op guard is necessary when it +protects a meaningful invariant, even if correct behavior is baseline-equivalent. + +Do not add a stimulus when it measures: + +- generic knowledge the base model already has; +- path recall for a skill that only points to reference files; +- output volume rather than correctness; +- a renamed or lightly reworded copy of an existing case; or +- a `disable-model-invocation: true` reference skill in isolation. + +Map each proposed case to **capability**, **risk**, and **customer journey** tags. Preference cases +must add distinct voting value. Activation contracts and no-op guards may share a capability when +they protect a separate routing or preservation invariant. If two cases have the same inputs, +expected outcome, failure mode, and grader path, keep the stronger one. The five-stimulus floor +never justifies padding. + +Then locate the target and test directory: ```text tests///eval.yaml # skills @@ -73,6 +95,10 @@ defaults: stimuli: - name: prompt: + tags: + capability: + risk: + journey: environment: files: - src: fixtures//Project.csproj @@ -87,11 +113,9 @@ stimuli: - ``` -> **`defaults:` replaces `config:` — it does not join it.** `config` is a deprecated alias for the -> same block and vally **throws** on a spec declaring both. Some existing evals still open with -> `config:`; when you change settings, replace it with one `defaults:` block. The failure is -> invisible otherwise: the job exits 0 with no verdicts and the PR comment -> blames "transient infrastructure". +> **Use `defaults:` only.** `config:` is a deprecated alias that the repository gate rejects. Vally +> warns when the alias appears alone and throws when a spec declares both keys. Replace `config:` +> with one `defaults:` block and preserve its settings. ### Step 3: Size the eval for power before writing content @@ -126,10 +150,15 @@ own value rather than defaulting it. cued prompts inflate the overfit score and bias the baseline. - Each stimulus should discriminate a **different** property of the skill. Five stimuli covering one property give arithmetic, not evidence. +- Give every capability stimulus non-empty `capability`, `risk`, and `journey` tags. Use stable + lowercase kebab-case values. A useful portfolio crosses distinct rows or columns in that matrix; + it does not repeat one journey with cosmetic wording changes. - Give every stimulus a stable, unique `name`. Vally pairs comparison trajectories by `(stimulus name, trial index)`; duplicate names make slot identity ambiguous. - Include a boundary / no-op stimulus for any skill that migrates or rewrites code, proving it leaves already-correct input alone. +- Add a dormancy stimulus for each real routing boundary. No-op proves restraint on an on-target, + already-correct input; dormancy proves the target stays inactive on an off-target request. ### Step 5: Configure the environment @@ -174,6 +203,9 @@ Fixture rules — each one has already cost a real result: (`|| exit 0`), or vally drops the trial. - A cleanup command that strips sources must skip directories containing `SKILL.md` — the staged skill lives there, and deleting it aborts only the skilled arm. +- **Preserve the complete file set.** When the prompt limits edits or requires source/test + preservation, snapshot or compare every in-scope file, not one representative file. A grader that + checks only the main output can miss deletion, truncation, or edits to sibling files. ### Step 6: Write graders @@ -199,6 +231,21 @@ Rules: `Recommendation:` line can silently stop doing so while the eval still passes. - Use `file-not-contains` / `file-not-exists` to prove the agent avoided an incorrect action. +Define the deterministic contract before writing the prompt grader: + +1. **Golden acceptance:** materialize the fixture and apply the `golden_patch`, if any. The golden + workspace and final `golden_trajectory` response must pass every deterministic grader that + applies to them. +2. **Mutation rejection:** make one realistic defect that the eval exists to catch, such as a + missing file, zero discovered tests, an out-of-scope edit, or a changed semantic value. The + relevant deterministic grader must fail. +3. **Complete-state check:** cover all files and artifacts named by the request. Do not accept a + partial artifact because one positive substring exists. + +Use `golden_patch` for replayable workspace state and `golden_trajectory` for the expected final +response. A narrated edit, build, or test is not proof: completion claims need a patch or a +`run-command` grader that replays the evidence. + ### Step 7: Write rubric items Rubric items are judged pairwise (baseline vs. skilled). The overfitting judge classifies each item: @@ -305,10 +352,28 @@ dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- evaluate \ CI adapts this result through `eng/vally-adapter/adapt-agent-results.mjs`, which applies the same distinct-stimulus sign-test policy used by skill results. -`check_eval_quality.py` blocks eleven structural defect classes that can corrupt a result: +Validation must cover four layers: + +1. **Deterministic structure:** run `check_eval_quality.py` and the relevant checker self-tests when + the checker changes. +2. **Production parsing and golden replay:** run skill evals through the repository's Vally entry + point. For agent evals, run `skill-validator evaluate` to prove the native SDK lane accepts the + executable scenario fields. That parser does not read `golden_trajectory` or `golden_patch`, so + validate the references separately: run `check_eval_quality.py`, then materialize the fixture, + apply the golden patch, and run every applicable deterministic file and command grader against + the golden workspace. Confirm the final golden response passes its output graders. +3. **Normal execution:** use the normal worker concurrency and the declared `defaults.timeout`. + Do not certify an eval only with one worker or a larger ad hoc time budget. If normal concurrency + exposes a race or timeout, classify it as reliability evidence. +4. **Cross-family sensitivity:** for broad routing or behavior changes, evaluate at least one GPT + family and one Claude family executor. Report each result separately. Different executor or + judge families do not increase the independent stimulus count. + +`check_eval_quality.py` blocks 22 structural defect classes that can corrupt a result: missing or untracked fixtures, self-contradicting coverage fixtures, empty grader configs, dormancy guards with `reject_skills`, sub-floor stimulus counts, duplicate YAML keys or stimulus names, and -`config:`/`defaults:` collisions. Do not add a new eval to +invalid defaults, tags, golden evidence, test commands, or ATIF trajectories. See +[`eng/eval-quality/README.md`](../../../eng/eval-quality/README.md) for the complete list. Do not add a new eval to `eng/eval-quality/underpowered-allowlist.txt` — the gate rejects allowlist entries that are new relative to the base branch. @@ -317,16 +382,22 @@ For the official run, submit a PR review containing `/evaluate` so it binds to t ## Validation Checklist - [ ] Directory is `tests///` or `tests//agent./` -- [ ] Spec uses `stimuli:` / `graders:`, and exactly one of `defaults:` or `config:` +- [ ] Spec uses `stimuli:` / `graders:` and the current `defaults:` settings block - [ ] At least 5 preference-eligible distinct stimuli exist; dormancy contracts do not count toward this floor -- [ ] Each stimulus discriminates a different property and has a stable, unique name +- [ ] Every stimulus is necessary and fits the target; each preference case adds distinct voting value +- [ ] Each capability stimulus has stable `capability`, `risk`, and `journey` tags and a unique name - [ ] Prompts never name the skill, the agent, or its vocabulary - [ ] Every referenced fixture exists and is tracked by `git ls-files` - [ ] Every fixture behaves as its stimulus assumes — healthy ones build, deliberately broken ones fail only for the stated reason +- [ ] Preservation and scope graders cover the complete in-scope file set - [ ] Every grader has its required `config` key - [ ] Any output shape the skill mandates has a grader +- [ ] Golden evidence passes deterministic graders, and a realistic mutation fails them - [ ] Rubric items are outcome-shaped and never reward using the skill -- [ ] Dormancy guards use `expect_activation: false` alone +- [ ] Rewrite skills have a no-op case; routing boundaries use `expect_activation: false` alone +- [ ] The production runner accepts the executable spec and completes under normal concurrency and time limits +- [ ] Golden references pass the standalone checker and deterministic replay +- [ ] Broad routing or behavior changes have separate GPT-family and Claude-family evidence - [ ] `skill-validator check` and `check_eval_quality.py` pass ## Common Pitfalls @@ -334,7 +405,7 @@ For the official run, submit a PR review containing `/evaluate` so it binds to t | Pitfall | Solution | |---------|----------| | Writing `scenarios:` / `assertions:` | That format no longer loads; use `stimuli:` / `graders:` | -| Adding `defaults: runs:` beside an existing `config:` | Merge into one `defaults:` block | +| Using the deprecated top-level `config:` alias | Rename it to `defaults:` and preserve its settings | | Landing an eval at exactly 5 stimuli | A single tie makes a pass unreachable; size for the effect and tie rate | | Raising `runs` to clear the floor | Repeats measure reliability for one task; add stimuli | | Prompt mentions the skill or agent by name | Rewrite as a natural developer request | diff --git a/catalog/Tools/Official-DotNet-Improve-Skill-Quality/skills/improve-skill-quality/SKILL.md b/catalog/Tools/Official-DotNet-Improve-Skill-Quality/skills/improve-skill-quality/SKILL.md index 76e9441..f93009c 100644 --- a/catalog/Tools/Official-DotNet-Improve-Skill-Quality/skills/improve-skill-quality/SKILL.md +++ b/catalog/Tools/Official-DotNet-Improve-Skill-Quality/skills/improve-skill-quality/SKILL.md @@ -76,8 +76,8 @@ a regression. Confirm that before reading a record as a power problem. See [references/eval-triage.md](references/eval-triage.md) for the full catalogue. The recurring ones: -- A spec declaring both `config:` and `defaults:` is rejected by vally, the job still exits 0, and - the PR comment blames "transient infrastructure". Merge them into one `defaults:` block. +- The repository gate rejects the deprecated top-level `config:` alias, and Vally rejects a spec + that declares both `config:` and `defaults:`. Replace the alias with one `defaults:` block. - An errored trial is not automatically a fixture problem — judge-side auth and `session.idle` failures look identical from the verdict and need harness fixes, not SDK pins. - `expect_tools: [bash]` on an advisory question forces a restore or build and turns an answer into @@ -87,10 +87,17 @@ See [references/eval-triage.md](references/eval-triage.md) for the full catalogu - Unmatched trajectories, an errored trial, or a summary that disagrees make the comparison **inconclusive**: the remaining matched trials are biased, so the record is not a measured null and must not be read as a power or content problem. +- A generic YAML parser is not the production loader. If Vally rejects a skill eval, or the native + SDK lane rejects an agent eval's executable scenario fields, classify it as harness / spec-load + before changing content. The native agent parser ignores golden references, so validate those + separately with `check_eval_quality.py` and deterministic golden-workspace replay. +- A pass with one worker or an enlarged local timeout is not normal execution evidence. Reproduce + with the repository's normal concurrency and declared suite budget; a failure there is a + reliability defect. ### Step 4: Verify the fixtures before touching the skill -Run `python eng/eval-quality/check_eval_quality.py` — it blocks eleven defect classes that can +Run `python eng/eval-quality/check_eval_quality.py` — it blocks 22 defect classes that can cost a real result here. Then confirm by hand: - every fixture behaves as its stimulus assumes — a fixture meant to be healthy builds, and one @@ -100,6 +107,10 @@ cost a real result here. Then confirm by hand: - a fixture never states the same fact in two places that disagree — a Cobertura report whose declared `line-rate`, summary totals and `` elements differ is the canonical case — or the two arms legitimately read different truths. +- every preservation or scope assertion covers the complete in-scope file set, not one + representative file. +- the golden trajectory and patch pass the deterministic graders, and a realistic broken mutation + fails the grader that is meant to protect the behavior. ### Step 5: Check whether the eval could ever have passed @@ -175,6 +186,14 @@ python eng/eval-quality/check_eval_quality.py ./eng/run-skill-evals.sh ``` +Use the production path at normal worker concurrency and with the declared `defaults.timeout`: +Vally for skill evals, and `skill-validator evaluate` for agent evals. For an agent eval, separately +run `check_eval_quality.py`, apply each golden patch to its materialized fixture, and run the +applicable deterministic file, output, and command graders against the golden result. Do not use a +serial-only pass or a larger ad hoc budget as completion evidence. For broad routing or behavior +changes, collect separate GPT-family and Claude-family results. Do not pool model families into +extra stimulus votes. + Then request the official run by submitting a PR review containing `/evaluate` (Files changed → Review changes), which binds the run to the reviewed commit. Before declaring a regression on the result, confirm the skill payload actually changed — reruns on byte-identical content have shifted @@ -185,8 +204,13 @@ result, confirm the skill payload actually changed — reruns on byte-identical - [ ] For a content fix, a losing trial and the judge's stated reason are quoted in the PR description. - [ ] The failure was classified before any content was edited. - [ ] `check_eval_quality.py` and `skill-validator check` both pass. +- [ ] The production skill or agent runner accepts the executable spec. +- [ ] Golden references pass the standalone checker and deterministic replay. +- [ ] The eval completes under normal concurrency and its declared time budget. +- [ ] Golden acceptance and mutation rejection have been demonstrated. - [ ] Distinct-stimulus count clears the power bar for the target effect and observed tie rate. - [ ] Isolated **and** plugin activation are both reported. +- [ ] Broad changes have separate GPT-family and Claude-family evidence. - [ ] The PR body records root cause, fix, and validation so the lesson is reusable. ## Common Pitfalls @@ -194,7 +218,7 @@ result, confirm the skill payload actually changed — reruns on byte-identical | Pitfall | Solution | |---------|----------| | Rewriting skill prose in response to an underpowered verdict | Underpowered means too few distinct stimuli; add discriminating stimuli instead | -| Adding `defaults: runs:` to a spec that already has `config:` | Merge into a single `defaults:` block; vally rejects specs with both | +| Using the deprecated top-level `config:` alias | Rename it to `defaults:` and preserve its settings; the repository gate rejects the alias | | Padding `runs` to clear the stimulus floor | Repeats measure reliability for one task; add stimuli | | Treating an errored trial as fixture nondeterminism | Read the stderr first; judge-side auth failures need harness fixes | | Fixing a "wrong" answer that the fixture actually made wrong | Check fixture self-consistency before blaming the response | @@ -205,5 +229,5 @@ result, confirm the skill payload actually changed — reruns on byte-identical - [references/writing-for-baseline-delta.md](references/writing-for-baseline-delta.md) — content patterns that beat the unskilled model - [references/eval-triage.md](references/eval-triage.md) — symptom, cause and fix catalogue with PR citations -- [eng/eval-quality/README.md](../../../eng/eval-quality/README.md) — the eleven structural gate checks and why each exists +- [eng/eval-quality/README.md](../../../eng/eval-quality/README.md) — the 22 structural gate checks and why each exists - [eng/vally-adapter/InvestigatingResults.md](../../../eng/vally-adapter/InvestigatingResults.md) — downloading artifacts and reading `results.json`. This is the current guide; the similarly-named `eng/skill-validator/src/docs/InvestigatingResults.md` documents the retired `skill-validator evaluate` schema and does not describe today's results. diff --git a/catalog/Tools/Official-DotNet-Improve-Skill-Quality/skills/improve-skill-quality/references/eval-triage.md b/catalog/Tools/Official-DotNet-Improve-Skill-Quality/skills/improve-skill-quality/references/eval-triage.md index 46ae934..5b666b3 100644 --- a/catalog/Tools/Official-DotNet-Improve-Skill-Quality/skills/improve-skill-quality/references/eval-triage.md +++ b/catalog/Tools/Official-DotNet-Improve-Skill-Quality/skills/improve-skill-quality/references/eval-triage.md @@ -7,11 +7,15 @@ Symptom → cause → fix, with the PR where each was diagnosed. Use with | Symptom | Cause | Fix | Evidence | |---------|-------|-----|----------| -| "Evaluation ran but produced no results", advice says transient infrastructure | Spec declares both `config:` and `defaults:`; vally throws, the job still exits 0 | Merge into one `defaults:` block carrying `timeout` and `runs` | PR #971 | +| "Evaluation ran but produced no results", advice says transient infrastructure | The spec declares both `config:` and `defaults:`; Vally rejects the mixed settings keys | Merge the settings into one `defaults:` block carrying `timeout` and `runs` | PR #971 | +| Authoring gate rejects an otherwise loadable spec | The spec uses the deprecated top-level `config:` alias | Rename `config:` to `defaults:` and preserve its settings | `eng/eval-quality/README.md` | | Same message, nothing in `plugins/` changed | Genuine LLM-session auth failure | Re-post `/evaluate`; inspect job logs before touching content | PR #932 | | Trial errored, avgN unusually low | Judge-side CAPI / `session.idle` timeout, not fixture nondeterminism | Read the trial stderr first; fix the harness, do not pin an SDK | PR #907 | | Timeout on an advisory question | `expect_tools: [bash]` forced a restore or build | Drop the tool requirement; the answer was always textual | PR #861 | | Every grader fails and output is empty | Code-generation stimulus timed out | Raise to ~360s | PR #862, PR #863 | +| Generic YAML parsing passes but a skill eval produces no result | Vally rejected the spec, golden reference, or ATIF trajectory | Reproduce through the repository Vally path and fix the first loader error | `eng/eval-quality/README.md` | +| Native agent execution passes but golden validation fails | The native agent parser ignores `golden_trajectory` and `golden_patch` | Run `check_eval_quality.py`, then replay the golden patch and deterministic graders separately | `eng/eval-quality/README.md` | +| Eval passes with one worker or a larger local timeout only | Concurrency race, shared-state leak, or an unrealistic suite budget | Run with normal workers and `defaults.timeout`; fix the reliability defect instead of certifying the special case | PR #1214 | | Only the skilled arm aborts with "contains no SKILL.md" | A setup cleanup command deleted the staged skill directory | Skip directories carrying `SKILL.md` when stripping sources | PR #878 | | Trials silently dropped | Setup command exited non-zero although its artifact was produced | Guard intentional failures, e.g. `dotnet build -bl \|\| exit 0` | PR #878 | | One arm has unmatched trajectories | Comparison judge error or arm timeout biases the remaining trials | Report inconclusive, do not read it as a regression | PR #887 | @@ -65,6 +69,9 @@ Consequences seen in real runs: | Overfit score high, user value unclear | Rubric items reward using the skill, or prompts echo skill vocabulary | Drop them: the harness already reports activation separately, so a rubric never needs to. Keep rubric items outcome-shaped and de-cue the prompt | PR #904 | | Both arms produce the same kind of artifact and the judge falls back on comparing volume | The rubric rewards raw output instead of the property under test | Add anti-hijack criteria: do not invoke the skill, and do not reward quantity (number of tests, findings, or lines produced) | PR #945 | | A grader appears to enforce something but does not | `config:` is missing its required key after an indentation slip | `check_eval_quality.py` blocks it; verify the key is present | `eng/eval-quality/README.md` | +| Golden evidence is GREEN only in prose | The trajectory claims edits or execution that its patch and command graders do not replay | Make the golden workspace and response pass every deterministic grader | PR #1214 | +| A plausible broken output still passes | The deterministic grader checks a marker, not the protected behavior | Add one realistic mutation and strengthen the grader until that mutation fails | PR #1213 | +| A scope-preservation scenario checks only one file | The agent can change or delete sibling files without detection | Snapshot or compare the complete in-scope file set | PR #1213 | | Two stimuli behave identically | Duplicate YAML key — a leftover `prompt:`/`graders:` block overwrites the following stimulus field by field | Delete the stray block after confirming it is not a distinct stimulus that lost its `- name:` | PR #971 | | Eval measures path recall | The skill is a map to reference files | Do not create the eval; test the consumer's outcome instead | PR #974 | @@ -86,6 +93,7 @@ Consequences seen in real runs: | Trigger `/evaluate` by submitting a PR review (Files changed → Review changes) so the run binds to the reviewed commit | PR #956, PR #949 | | Before declaring a regression, confirm the invoked payload changed — reruns on byte-identical content moved 7W/2T/2L to 4W/5T/2L | PR #974 | | Use cross-family evaluation for broad rewrites; single-family passes hide model-specific regressions | issue #899, PR #947 | +| Keep executor-family results separate; cross-family combinations are sensitivity evidence, not extra task votes | `eng/vally-adapter/README.md` | | Workflow changes cannot be validated by the PR's own evaluation (GitHub runs workflow definitions from `main`) — use manual dispatches | PR #872 | | Prefer deterministic scripts over agentic workflows for deterministic policy | PR #928 | | Agent `tools:` allowlists are host-specific and case-sensitive; an allowlist can grant zero tools | PR #856, PR #847 | diff --git a/external-sources/upstreams/astro/packages/astro/package.json b/external-sources/upstreams/astro/packages/astro/package.json index 5a44df3..5fe33b8 100644 --- a/external-sources/upstreams/astro/packages/astro/package.json +++ b/external-sources/upstreams/astro/packages/astro/package.json @@ -1,6 +1,6 @@ { "name": "astro", - "version": "7.3.5", + "version": "7.3.6", "description": "Astro is a modern site builder with web best practices, performance, and DX front-of-mind.", "type": "module", "author": "withastro", diff --git a/external-sources/upstreams/dotnet-skills/.agents/skills/create-skill-test/SKILL.md b/external-sources/upstreams/dotnet-skills/.agents/skills/create-skill-test/SKILL.md index e1f4f30..b5eaa5b 100644 --- a/external-sources/upstreams/dotnet-skills/.agents/skills/create-skill-test/SKILL.md +++ b/external-sources/upstreams/dotnet-skills/.agents/skills/create-skill-test/SKILL.md @@ -30,11 +30,33 @@ and does not overfit to the skill's own wording. | Skill or agent name | Yes | Must exist under `plugins//skills/` or `plugins//agents/` | | Plugin name | Yes | e.g. `dotnet-msbuild` | | Skill content | Yes | Read it — you cannot write non-overfitted rubric items without it | -| Failure modes to discriminate | Recommended | Each becomes one stimulus | +| Scenario hypothesis | Yes | State the expected improvement for a preference case, or the invariant protected by a guard | +| Failure modes to discriminate | Recommended | Each becomes one distinct stimulus | ## Workflow -### Step 1: Locate the target and the test directory +### Step 1: Prove the eval belongs to the target + +Before writing YAML, state what the stimulus proves. A preference stimulus is necessary when the +target should improve the answer, action, restraint, or validation result compared with the same +model without the target. A non-voting activation contract or no-op guard is necessary when it +protects a meaningful invariant, even if correct behavior is baseline-equivalent. + +Do not add a stimulus when it measures: + +- generic knowledge the base model already has; +- path recall for a skill that only points to reference files; +- output volume rather than correctness; +- a renamed or lightly reworded copy of an existing case; or +- a `disable-model-invocation: true` reference skill in isolation. + +Map each proposed case to **capability**, **risk**, and **customer journey** tags. Preference cases +must add distinct voting value. Activation contracts and no-op guards may share a capability when +they protect a separate routing or preservation invariant. If two cases have the same inputs, +expected outcome, failure mode, and grader path, keep the stronger one. The five-stimulus floor +never justifies padding. + +Then locate the target and test directory: ```text tests///eval.yaml # skills @@ -73,6 +95,10 @@ defaults: stimuli: - name: prompt: + tags: + capability: + risk: + journey: environment: files: - src: fixtures//Project.csproj @@ -87,11 +113,9 @@ stimuli: - ``` -> **`defaults:` replaces `config:` — it does not join it.** `config` is a deprecated alias for the -> same block and vally **throws** on a spec declaring both. Some existing evals still open with -> `config:`; when you change settings, replace it with one `defaults:` block. The failure is -> invisible otherwise: the job exits 0 with no verdicts and the PR comment -> blames "transient infrastructure". +> **Use `defaults:` only.** `config:` is a deprecated alias that the repository gate rejects. Vally +> warns when the alias appears alone and throws when a spec declares both keys. Replace `config:` +> with one `defaults:` block and preserve its settings. ### Step 3: Size the eval for power before writing content @@ -126,10 +150,15 @@ own value rather than defaulting it. cued prompts inflate the overfit score and bias the baseline. - Each stimulus should discriminate a **different** property of the skill. Five stimuli covering one property give arithmetic, not evidence. +- Give every capability stimulus non-empty `capability`, `risk`, and `journey` tags. Use stable + lowercase kebab-case values. A useful portfolio crosses distinct rows or columns in that matrix; + it does not repeat one journey with cosmetic wording changes. - Give every stimulus a stable, unique `name`. Vally pairs comparison trajectories by `(stimulus name, trial index)`; duplicate names make slot identity ambiguous. - Include a boundary / no-op stimulus for any skill that migrates or rewrites code, proving it leaves already-correct input alone. +- Add a dormancy stimulus for each real routing boundary. No-op proves restraint on an on-target, + already-correct input; dormancy proves the target stays inactive on an off-target request. ### Step 5: Configure the environment @@ -174,6 +203,9 @@ Fixture rules — each one has already cost a real result: (`|| exit 0`), or vally drops the trial. - A cleanup command that strips sources must skip directories containing `SKILL.md` — the staged skill lives there, and deleting it aborts only the skilled arm. +- **Preserve the complete file set.** When the prompt limits edits or requires source/test + preservation, snapshot or compare every in-scope file, not one representative file. A grader that + checks only the main output can miss deletion, truncation, or edits to sibling files. ### Step 6: Write graders @@ -199,6 +231,21 @@ Rules: `Recommendation:` line can silently stop doing so while the eval still passes. - Use `file-not-contains` / `file-not-exists` to prove the agent avoided an incorrect action. +Define the deterministic contract before writing the prompt grader: + +1. **Golden acceptance:** materialize the fixture and apply the `golden_patch`, if any. The golden + workspace and final `golden_trajectory` response must pass every deterministic grader that + applies to them. +2. **Mutation rejection:** make one realistic defect that the eval exists to catch, such as a + missing file, zero discovered tests, an out-of-scope edit, or a changed semantic value. The + relevant deterministic grader must fail. +3. **Complete-state check:** cover all files and artifacts named by the request. Do not accept a + partial artifact because one positive substring exists. + +Use `golden_patch` for replayable workspace state and `golden_trajectory` for the expected final +response. A narrated edit, build, or test is not proof: completion claims need a patch or a +`run-command` grader that replays the evidence. + ### Step 7: Write rubric items Rubric items are judged pairwise (baseline vs. skilled). The overfitting judge classifies each item: @@ -305,10 +352,28 @@ dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- evaluate \ CI adapts this result through `eng/vally-adapter/adapt-agent-results.mjs`, which applies the same distinct-stimulus sign-test policy used by skill results. -`check_eval_quality.py` blocks eleven structural defect classes that can corrupt a result: +Validation must cover four layers: + +1. **Deterministic structure:** run `check_eval_quality.py` and the relevant checker self-tests when + the checker changes. +2. **Production parsing and golden replay:** run skill evals through the repository's Vally entry + point. For agent evals, run `skill-validator evaluate` to prove the native SDK lane accepts the + executable scenario fields. That parser does not read `golden_trajectory` or `golden_patch`, so + validate the references separately: run `check_eval_quality.py`, then materialize the fixture, + apply the golden patch, and run every applicable deterministic file and command grader against + the golden workspace. Confirm the final golden response passes its output graders. +3. **Normal execution:** use the normal worker concurrency and the declared `defaults.timeout`. + Do not certify an eval only with one worker or a larger ad hoc time budget. If normal concurrency + exposes a race or timeout, classify it as reliability evidence. +4. **Cross-family sensitivity:** for broad routing or behavior changes, evaluate at least one GPT + family and one Claude family executor. Report each result separately. Different executor or + judge families do not increase the independent stimulus count. + +`check_eval_quality.py` blocks 22 structural defect classes that can corrupt a result: missing or untracked fixtures, self-contradicting coverage fixtures, empty grader configs, dormancy guards with `reject_skills`, sub-floor stimulus counts, duplicate YAML keys or stimulus names, and -`config:`/`defaults:` collisions. Do not add a new eval to +invalid defaults, tags, golden evidence, test commands, or ATIF trajectories. See +[`eng/eval-quality/README.md`](../../../eng/eval-quality/README.md) for the complete list. Do not add a new eval to `eng/eval-quality/underpowered-allowlist.txt` — the gate rejects allowlist entries that are new relative to the base branch. @@ -317,16 +382,22 @@ For the official run, submit a PR review containing `/evaluate` so it binds to t ## Validation Checklist - [ ] Directory is `tests///` or `tests//agent./` -- [ ] Spec uses `stimuli:` / `graders:`, and exactly one of `defaults:` or `config:` +- [ ] Spec uses `stimuli:` / `graders:` and the current `defaults:` settings block - [ ] At least 5 preference-eligible distinct stimuli exist; dormancy contracts do not count toward this floor -- [ ] Each stimulus discriminates a different property and has a stable, unique name +- [ ] Every stimulus is necessary and fits the target; each preference case adds distinct voting value +- [ ] Each capability stimulus has stable `capability`, `risk`, and `journey` tags and a unique name - [ ] Prompts never name the skill, the agent, or its vocabulary - [ ] Every referenced fixture exists and is tracked by `git ls-files` - [ ] Every fixture behaves as its stimulus assumes — healthy ones build, deliberately broken ones fail only for the stated reason +- [ ] Preservation and scope graders cover the complete in-scope file set - [ ] Every grader has its required `config` key - [ ] Any output shape the skill mandates has a grader +- [ ] Golden evidence passes deterministic graders, and a realistic mutation fails them - [ ] Rubric items are outcome-shaped and never reward using the skill -- [ ] Dormancy guards use `expect_activation: false` alone +- [ ] Rewrite skills have a no-op case; routing boundaries use `expect_activation: false` alone +- [ ] The production runner accepts the executable spec and completes under normal concurrency and time limits +- [ ] Golden references pass the standalone checker and deterministic replay +- [ ] Broad routing or behavior changes have separate GPT-family and Claude-family evidence - [ ] `skill-validator check` and `check_eval_quality.py` pass ## Common Pitfalls @@ -334,7 +405,7 @@ For the official run, submit a PR review containing `/evaluate` so it binds to t | Pitfall | Solution | |---------|----------| | Writing `scenarios:` / `assertions:` | That format no longer loads; use `stimuli:` / `graders:` | -| Adding `defaults: runs:` beside an existing `config:` | Merge into one `defaults:` block | +| Using the deprecated top-level `config:` alias | Rename it to `defaults:` and preserve its settings | | Landing an eval at exactly 5 stimuli | A single tie makes a pass unreachable; size for the effect and tie rate | | Raising `runs` to clear the floor | Repeats measure reliability for one task; add stimuli | | Prompt mentions the skill or agent by name | Rewrite as a natural developer request | diff --git a/external-sources/upstreams/dotnet-skills/.agents/skills/improve-skill-quality/SKILL.md b/external-sources/upstreams/dotnet-skills/.agents/skills/improve-skill-quality/SKILL.md index 76e9441..f93009c 100644 --- a/external-sources/upstreams/dotnet-skills/.agents/skills/improve-skill-quality/SKILL.md +++ b/external-sources/upstreams/dotnet-skills/.agents/skills/improve-skill-quality/SKILL.md @@ -76,8 +76,8 @@ a regression. Confirm that before reading a record as a power problem. See [references/eval-triage.md](references/eval-triage.md) for the full catalogue. The recurring ones: -- A spec declaring both `config:` and `defaults:` is rejected by vally, the job still exits 0, and - the PR comment blames "transient infrastructure". Merge them into one `defaults:` block. +- The repository gate rejects the deprecated top-level `config:` alias, and Vally rejects a spec + that declares both `config:` and `defaults:`. Replace the alias with one `defaults:` block. - An errored trial is not automatically a fixture problem — judge-side auth and `session.idle` failures look identical from the verdict and need harness fixes, not SDK pins. - `expect_tools: [bash]` on an advisory question forces a restore or build and turns an answer into @@ -87,10 +87,17 @@ See [references/eval-triage.md](references/eval-triage.md) for the full catalogu - Unmatched trajectories, an errored trial, or a summary that disagrees make the comparison **inconclusive**: the remaining matched trials are biased, so the record is not a measured null and must not be read as a power or content problem. +- A generic YAML parser is not the production loader. If Vally rejects a skill eval, or the native + SDK lane rejects an agent eval's executable scenario fields, classify it as harness / spec-load + before changing content. The native agent parser ignores golden references, so validate those + separately with `check_eval_quality.py` and deterministic golden-workspace replay. +- A pass with one worker or an enlarged local timeout is not normal execution evidence. Reproduce + with the repository's normal concurrency and declared suite budget; a failure there is a + reliability defect. ### Step 4: Verify the fixtures before touching the skill -Run `python eng/eval-quality/check_eval_quality.py` — it blocks eleven defect classes that can +Run `python eng/eval-quality/check_eval_quality.py` — it blocks 22 defect classes that can cost a real result here. Then confirm by hand: - every fixture behaves as its stimulus assumes — a fixture meant to be healthy builds, and one @@ -100,6 +107,10 @@ cost a real result here. Then confirm by hand: - a fixture never states the same fact in two places that disagree — a Cobertura report whose declared `line-rate`, summary totals and `` elements differ is the canonical case — or the two arms legitimately read different truths. +- every preservation or scope assertion covers the complete in-scope file set, not one + representative file. +- the golden trajectory and patch pass the deterministic graders, and a realistic broken mutation + fails the grader that is meant to protect the behavior. ### Step 5: Check whether the eval could ever have passed @@ -175,6 +186,14 @@ python eng/eval-quality/check_eval_quality.py ./eng/run-skill-evals.sh ``` +Use the production path at normal worker concurrency and with the declared `defaults.timeout`: +Vally for skill evals, and `skill-validator evaluate` for agent evals. For an agent eval, separately +run `check_eval_quality.py`, apply each golden patch to its materialized fixture, and run the +applicable deterministic file, output, and command graders against the golden result. Do not use a +serial-only pass or a larger ad hoc budget as completion evidence. For broad routing or behavior +changes, collect separate GPT-family and Claude-family results. Do not pool model families into +extra stimulus votes. + Then request the official run by submitting a PR review containing `/evaluate` (Files changed → Review changes), which binds the run to the reviewed commit. Before declaring a regression on the result, confirm the skill payload actually changed — reruns on byte-identical content have shifted @@ -185,8 +204,13 @@ result, confirm the skill payload actually changed — reruns on byte-identical - [ ] For a content fix, a losing trial and the judge's stated reason are quoted in the PR description. - [ ] The failure was classified before any content was edited. - [ ] `check_eval_quality.py` and `skill-validator check` both pass. +- [ ] The production skill or agent runner accepts the executable spec. +- [ ] Golden references pass the standalone checker and deterministic replay. +- [ ] The eval completes under normal concurrency and its declared time budget. +- [ ] Golden acceptance and mutation rejection have been demonstrated. - [ ] Distinct-stimulus count clears the power bar for the target effect and observed tie rate. - [ ] Isolated **and** plugin activation are both reported. +- [ ] Broad changes have separate GPT-family and Claude-family evidence. - [ ] The PR body records root cause, fix, and validation so the lesson is reusable. ## Common Pitfalls @@ -194,7 +218,7 @@ result, confirm the skill payload actually changed — reruns on byte-identical | Pitfall | Solution | |---------|----------| | Rewriting skill prose in response to an underpowered verdict | Underpowered means too few distinct stimuli; add discriminating stimuli instead | -| Adding `defaults: runs:` to a spec that already has `config:` | Merge into a single `defaults:` block; vally rejects specs with both | +| Using the deprecated top-level `config:` alias | Rename it to `defaults:` and preserve its settings; the repository gate rejects the alias | | Padding `runs` to clear the stimulus floor | Repeats measure reliability for one task; add stimuli | | Treating an errored trial as fixture nondeterminism | Read the stderr first; judge-side auth failures need harness fixes | | Fixing a "wrong" answer that the fixture actually made wrong | Check fixture self-consistency before blaming the response | @@ -205,5 +229,5 @@ result, confirm the skill payload actually changed — reruns on byte-identical - [references/writing-for-baseline-delta.md](references/writing-for-baseline-delta.md) — content patterns that beat the unskilled model - [references/eval-triage.md](references/eval-triage.md) — symptom, cause and fix catalogue with PR citations -- [eng/eval-quality/README.md](../../../eng/eval-quality/README.md) — the eleven structural gate checks and why each exists +- [eng/eval-quality/README.md](../../../eng/eval-quality/README.md) — the 22 structural gate checks and why each exists - [eng/vally-adapter/InvestigatingResults.md](../../../eng/vally-adapter/InvestigatingResults.md) — downloading artifacts and reading `results.json`. This is the current guide; the similarly-named `eng/skill-validator/src/docs/InvestigatingResults.md` documents the retired `skill-validator evaluate` schema and does not describe today's results. diff --git a/external-sources/upstreams/dotnet-skills/.agents/skills/improve-skill-quality/references/eval-triage.md b/external-sources/upstreams/dotnet-skills/.agents/skills/improve-skill-quality/references/eval-triage.md index 46ae934..5b666b3 100644 --- a/external-sources/upstreams/dotnet-skills/.agents/skills/improve-skill-quality/references/eval-triage.md +++ b/external-sources/upstreams/dotnet-skills/.agents/skills/improve-skill-quality/references/eval-triage.md @@ -7,11 +7,15 @@ Symptom → cause → fix, with the PR where each was diagnosed. Use with | Symptom | Cause | Fix | Evidence | |---------|-------|-----|----------| -| "Evaluation ran but produced no results", advice says transient infrastructure | Spec declares both `config:` and `defaults:`; vally throws, the job still exits 0 | Merge into one `defaults:` block carrying `timeout` and `runs` | PR #971 | +| "Evaluation ran but produced no results", advice says transient infrastructure | The spec declares both `config:` and `defaults:`; Vally rejects the mixed settings keys | Merge the settings into one `defaults:` block carrying `timeout` and `runs` | PR #971 | +| Authoring gate rejects an otherwise loadable spec | The spec uses the deprecated top-level `config:` alias | Rename `config:` to `defaults:` and preserve its settings | `eng/eval-quality/README.md` | | Same message, nothing in `plugins/` changed | Genuine LLM-session auth failure | Re-post `/evaluate`; inspect job logs before touching content | PR #932 | | Trial errored, avgN unusually low | Judge-side CAPI / `session.idle` timeout, not fixture nondeterminism | Read the trial stderr first; fix the harness, do not pin an SDK | PR #907 | | Timeout on an advisory question | `expect_tools: [bash]` forced a restore or build | Drop the tool requirement; the answer was always textual | PR #861 | | Every grader fails and output is empty | Code-generation stimulus timed out | Raise to ~360s | PR #862, PR #863 | +| Generic YAML parsing passes but a skill eval produces no result | Vally rejected the spec, golden reference, or ATIF trajectory | Reproduce through the repository Vally path and fix the first loader error | `eng/eval-quality/README.md` | +| Native agent execution passes but golden validation fails | The native agent parser ignores `golden_trajectory` and `golden_patch` | Run `check_eval_quality.py`, then replay the golden patch and deterministic graders separately | `eng/eval-quality/README.md` | +| Eval passes with one worker or a larger local timeout only | Concurrency race, shared-state leak, or an unrealistic suite budget | Run with normal workers and `defaults.timeout`; fix the reliability defect instead of certifying the special case | PR #1214 | | Only the skilled arm aborts with "contains no SKILL.md" | A setup cleanup command deleted the staged skill directory | Skip directories carrying `SKILL.md` when stripping sources | PR #878 | | Trials silently dropped | Setup command exited non-zero although its artifact was produced | Guard intentional failures, e.g. `dotnet build -bl \|\| exit 0` | PR #878 | | One arm has unmatched trajectories | Comparison judge error or arm timeout biases the remaining trials | Report inconclusive, do not read it as a regression | PR #887 | @@ -65,6 +69,9 @@ Consequences seen in real runs: | Overfit score high, user value unclear | Rubric items reward using the skill, or prompts echo skill vocabulary | Drop them: the harness already reports activation separately, so a rubric never needs to. Keep rubric items outcome-shaped and de-cue the prompt | PR #904 | | Both arms produce the same kind of artifact and the judge falls back on comparing volume | The rubric rewards raw output instead of the property under test | Add anti-hijack criteria: do not invoke the skill, and do not reward quantity (number of tests, findings, or lines produced) | PR #945 | | A grader appears to enforce something but does not | `config:` is missing its required key after an indentation slip | `check_eval_quality.py` blocks it; verify the key is present | `eng/eval-quality/README.md` | +| Golden evidence is GREEN only in prose | The trajectory claims edits or execution that its patch and command graders do not replay | Make the golden workspace and response pass every deterministic grader | PR #1214 | +| A plausible broken output still passes | The deterministic grader checks a marker, not the protected behavior | Add one realistic mutation and strengthen the grader until that mutation fails | PR #1213 | +| A scope-preservation scenario checks only one file | The agent can change or delete sibling files without detection | Snapshot or compare the complete in-scope file set | PR #1213 | | Two stimuli behave identically | Duplicate YAML key — a leftover `prompt:`/`graders:` block overwrites the following stimulus field by field | Delete the stray block after confirming it is not a distinct stimulus that lost its `- name:` | PR #971 | | Eval measures path recall | The skill is a map to reference files | Do not create the eval; test the consumer's outcome instead | PR #974 | @@ -86,6 +93,7 @@ Consequences seen in real runs: | Trigger `/evaluate` by submitting a PR review (Files changed → Review changes) so the run binds to the reviewed commit | PR #956, PR #949 | | Before declaring a regression, confirm the invoked payload changed — reruns on byte-identical content moved 7W/2T/2L to 4W/5T/2L | PR #974 | | Use cross-family evaluation for broad rewrites; single-family passes hide model-specific regressions | issue #899, PR #947 | +| Keep executor-family results separate; cross-family combinations are sensitivity evidence, not extra task votes | `eng/vally-adapter/README.md` | | Workflow changes cannot be validated by the PR's own evaluation (GitHub runs workflow definitions from `main`) — use manual dispatches | PR #872 | | Prefer deterministic scripts over agentic workflows for deterministic policy | PR #928 | | Agent `tools:` allowlists are host-specific and case-sensitive; an allowlist can grant zero tools | PR #856, PR #847 | diff --git a/external-sources/vendir.lock.yml b/external-sources/vendir.lock.yml index bb8ee41..0f3dab2 100644 --- a/external-sources/vendir.lock.yml +++ b/external-sources/vendir.lock.yml @@ -2,21 +2,20 @@ apiVersion: vendir.k14s.io/v1alpha1 directories: - contents: - git: - commitTitle: 'Merge pull request #1194 from dotnet/bot/weekly-version-sync...' - sha: c7cc6617da6c6cadc4ebe7b06926a02caa108e82 + commitTitle: 'Merge pull request #1240 from dotnet/abhitejjohn-eval-authoring-quality-bar...' + sha: 8d670fa76aaac45b336d8ded05a7601785fb2121 tags: - - skill-validator-nightly-23-gc7cc6617 + - skill-validator-nightly-10-g8d670fa7 path: dotnet-skills - git: commitTitle: Add Cursor rules that reference the existing skill docs... sha: af2319bd01bb7cc881267a9ef42cafdaf5e9029d path: webgpu-claude-skill - git: - commitTitle: Skip dev toolbar source annotations when compiling in Vite test - mode (#18162) (#18242)... - sha: c5275e3d5a688a7d246c9576b594870675b70e37 + commitTitle: 'fix(assets): return requested format from Sharp transform (#18222)...' + sha: fcf6ed6d4915eed6c18a20658b73372d074c4832 tags: - - astro@7.3.5-67-gc5275e3d5a + - '@astrojs/cloudflare@14.3.4-7-gfcf6ed6d49' path: astro path: upstreams kind: LockConfig