Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"version": "7.3.5",
"version": "7.3.6",
"category": "Web",
"compatibility": "Requires the withastro/astro monorepo."
}
Original file line number Diff line number Diff line change
Expand Up @@ -30,11 +30,33 @@ and does not overfit to the skill's own wording.
| Skill or agent name | Yes | Must exist under `plugins/<plugin>/skills/` or `plugins/<plugin>/agents/` |
| Plugin name | Yes | e.g. `dotnet-msbuild` |
| Skill content | Yes | Read it — you cannot write non-overfitted rubric items without it |
| Failure modes to discriminate | Recommended | Each becomes one stimulus |
| Scenario hypothesis | Yes | State the expected improvement for a preference case, or the invariant protected by a guard |
| Failure modes to discriminate | Recommended | Each becomes one distinct stimulus |

## Workflow

### Step 1: Locate the target and the test directory
### Step 1: Prove the eval belongs to the target

Before writing YAML, state what the stimulus proves. A preference stimulus is necessary when the
target should improve the answer, action, restraint, or validation result compared with the same
model without the target. A non-voting activation contract or no-op guard is necessary when it
protects a meaningful invariant, even if correct behavior is baseline-equivalent.

Do not add a stimulus when it measures:

- generic knowledge the base model already has;
- path recall for a skill that only points to reference files;
- output volume rather than correctness;
- a renamed or lightly reworded copy of an existing case; or
- a `disable-model-invocation: true` reference skill in isolation.

Map each proposed case to **capability**, **risk**, and **customer journey** tags. Preference cases
must add distinct voting value. Activation contracts and no-op guards may share a capability when
they protect a separate routing or preservation invariant. If two cases have the same inputs,
expected outcome, failure mode, and grader path, keep the stronger one. The five-stimulus floor
never justifies padding.

Then locate the target and test directory:

```text
tests/<plugin>/<skill-name>/eval.yaml # skills
Expand Down Expand Up @@ -73,6 +95,10 @@ defaults:
stimuli:
- name: <what the agent must accomplish>
prompt: <natural developer request>
tags:
capability: <distinct-capability>
risk: <failure-being-prevented>
journey: <customer-task>
environment:
files:
- src: fixtures/<case>/Project.csproj
Expand All @@ -87,11 +113,9 @@ stimuli:
- <outcome the agent should have reached>
```

> **`defaults:` replaces `config:` — it does not join it.** `config` is a deprecated alias for the
> same block and vally **throws** on a spec declaring both. Some existing evals still open with
> `config:`; when you change settings, replace it with one `defaults:` block. The failure is
> invisible otherwise: the job exits 0 with no verdicts and the PR comment
> blames "transient infrastructure".
> **Use `defaults:` only.** `config:` is a deprecated alias that the repository gate rejects. Vally
> warns when the alias appears alone and throws when a spec declares both keys. Replace `config:`
> with one `defaults:` block and preserve its settings.

### Step 3: Size the eval for power before writing content

Expand Down Expand Up @@ -126,10 +150,15 @@ own value rather than defaulting it.
cued prompts inflate the overfit score and bias the baseline.
- Each stimulus should discriminate a **different** property of the skill. Five stimuli covering one
property give arithmetic, not evidence.
- Give every capability stimulus non-empty `capability`, `risk`, and `journey` tags. Use stable
lowercase kebab-case values. A useful portfolio crosses distinct rows or columns in that matrix;
it does not repeat one journey with cosmetic wording changes.
- Give every stimulus a stable, unique `name`. Vally pairs comparison trajectories by
`(stimulus name, trial index)`; duplicate names make slot identity ambiguous.
- Include a boundary / no-op stimulus for any skill that migrates or rewrites code, proving it
leaves already-correct input alone.
- Add a dormancy stimulus for each real routing boundary. No-op proves restraint on an on-target,
already-correct input; dormancy proves the target stays inactive on an off-target request.

### Step 5: Configure the environment

Expand Down Expand Up @@ -174,6 +203,9 @@ Fixture rules — each one has already cost a real result:
(`|| exit 0`), or vally drops the trial.
- A cleanup command that strips sources must skip directories containing `SKILL.md` — the staged
skill lives there, and deleting it aborts only the skilled arm.
- **Preserve the complete file set.** When the prompt limits edits or requires source/test
preservation, snapshot or compare every in-scope file, not one representative file. A grader that
checks only the main output can miss deletion, truncation, or edits to sibling files.

### Step 6: Write graders

Expand All @@ -199,6 +231,21 @@ Rules:
`Recommendation:` line can silently stop doing so while the eval still passes.
- Use `file-not-contains` / `file-not-exists` to prove the agent avoided an incorrect action.

Define the deterministic contract before writing the prompt grader:

1. **Golden acceptance:** materialize the fixture and apply the `golden_patch`, if any. The golden
workspace and final `golden_trajectory` response must pass every deterministic grader that
applies to them.
2. **Mutation rejection:** make one realistic defect that the eval exists to catch, such as a
missing file, zero discovered tests, an out-of-scope edit, or a changed semantic value. The
relevant deterministic grader must fail.
3. **Complete-state check:** cover all files and artifacts named by the request. Do not accept a
partial artifact because one positive substring exists.

Use `golden_patch` for replayable workspace state and `golden_trajectory` for the expected final
response. A narrated edit, build, or test is not proof: completion claims need a patch or a
`run-command` grader that replays the evidence.

### Step 7: Write rubric items

Rubric items are judged pairwise (baseline vs. skilled). The overfitting judge classifies each item:
Expand Down Expand Up @@ -305,10 +352,28 @@ dotnet run --project eng/skill-validator/src/SkillValidator.csproj -- evaluate \
CI adapts this result through `eng/vally-adapter/adapt-agent-results.mjs`,
which applies the same distinct-stimulus sign-test policy used by skill results.

`check_eval_quality.py` blocks eleven structural defect classes that can corrupt a result:
Validation must cover four layers:

1. **Deterministic structure:** run `check_eval_quality.py` and the relevant checker self-tests when
the checker changes.
2. **Production parsing and golden replay:** run skill evals through the repository's Vally entry
point. For agent evals, run `skill-validator evaluate` to prove the native SDK lane accepts the
executable scenario fields. That parser does not read `golden_trajectory` or `golden_patch`, so
validate the references separately: run `check_eval_quality.py`, then materialize the fixture,
apply the golden patch, and run every applicable deterministic file and command grader against
the golden workspace. Confirm the final golden response passes its output graders.
3. **Normal execution:** use the normal worker concurrency and the declared `defaults.timeout`.
Do not certify an eval only with one worker or a larger ad hoc time budget. If normal concurrency
exposes a race or timeout, classify it as reliability evidence.
4. **Cross-family sensitivity:** for broad routing or behavior changes, evaluate at least one GPT
family and one Claude family executor. Report each result separately. Different executor or
judge families do not increase the independent stimulus count.

`check_eval_quality.py` blocks 22 structural defect classes that can corrupt a result:
missing or untracked fixtures, self-contradicting coverage fixtures, empty grader configs, dormancy
guards with `reject_skills`, sub-floor stimulus counts, duplicate YAML keys or stimulus names, and
`config:`/`defaults:` collisions. Do not add a new eval to
invalid defaults, tags, golden evidence, test commands, or ATIF trajectories. See
[`eng/eval-quality/README.md`](../../../eng/eval-quality/README.md) for the complete list. Do not add a new eval to
`eng/eval-quality/underpowered-allowlist.txt` — the gate rejects
allowlist entries that are new relative to the base branch.

Expand All @@ -317,24 +382,30 @@ For the official run, submit a PR review containing `/evaluate` so it binds to t
## Validation Checklist

- [ ] Directory is `tests/<plugin>/<skill-name>/` or `tests/<plugin>/agent.<agent-name>/`
- [ ] Spec uses `stimuli:` / `graders:`, and exactly one of `defaults:` or `config:`
- [ ] Spec uses `stimuli:` / `graders:` and the current `defaults:` settings block
- [ ] At least 5 preference-eligible distinct stimuli exist; dormancy contracts do not count toward this floor
- [ ] Each stimulus discriminates a different property and has a stable, unique name
- [ ] Every stimulus is necessary and fits the target; each preference case adds distinct voting value
- [ ] Each capability stimulus has stable `capability`, `risk`, and `journey` tags and a unique name
- [ ] Prompts never name the skill, the agent, or its vocabulary
- [ ] Every referenced fixture exists and is tracked by `git ls-files`
- [ ] Every fixture behaves as its stimulus assumes — healthy ones build, deliberately broken ones fail only for the stated reason
- [ ] Preservation and scope graders cover the complete in-scope file set
- [ ] Every grader has its required `config` key
- [ ] Any output shape the skill mandates has a grader
- [ ] Golden evidence passes deterministic graders, and a realistic mutation fails them
- [ ] Rubric items are outcome-shaped and never reward using the skill
- [ ] Dormancy guards use `expect_activation: false` alone
- [ ] Rewrite skills have a no-op case; routing boundaries use `expect_activation: false` alone
- [ ] The production runner accepts the executable spec and completes under normal concurrency and time limits
- [ ] Golden references pass the standalone checker and deterministic replay
- [ ] Broad routing or behavior changes have separate GPT-family and Claude-family evidence
- [ ] `skill-validator check` and `check_eval_quality.py` pass

## Common Pitfalls

| Pitfall | Solution |
|---------|----------|
| Writing `scenarios:` / `assertions:` | That format no longer loads; use `stimuli:` / `graders:` |
| Adding `defaults: runs:` beside an existing `config:` | Merge into one `defaults:` block |
| Using the deprecated top-level `config:` alias | Rename it to `defaults:` and preserve its settings |
| Landing an eval at exactly 5 stimuli | A single tie makes a pass unreachable; size for the effect and tie rate |
| Raising `runs` to clear the floor | Repeats measure reliability for one task; add stimuli |
| Prompt mentions the skill or agent by name | Rewrite as a natural developer request |
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -76,8 +76,8 @@ a regression. Confirm that before reading a record as a power problem.

See [references/eval-triage.md](references/eval-triage.md) for the full catalogue. The recurring ones:

- A spec declaring both `config:` and `defaults:` is rejected by vally, the job still exits 0, and
the PR comment blames "transient infrastructure". Merge them into one `defaults:` block.
- The repository gate rejects the deprecated top-level `config:` alias, and Vally rejects a spec
that declares both `config:` and `defaults:`. Replace the alias with one `defaults:` block.
- An errored trial is not automatically a fixture problem — judge-side auth and `session.idle`
failures look identical from the verdict and need harness fixes, not SDK pins.
- `expect_tools: [bash]` on an advisory question forces a restore or build and turns an answer into
Expand All @@ -87,10 +87,17 @@ See [references/eval-triage.md](references/eval-triage.md) for the full catalogu
- Unmatched trajectories, an errored trial, or a summary that disagrees make the comparison
**inconclusive**: the remaining matched trials are biased, so the record is not a measured null
and must not be read as a power or content problem.
- A generic YAML parser is not the production loader. If Vally rejects a skill eval, or the native
SDK lane rejects an agent eval's executable scenario fields, classify it as harness / spec-load
before changing content. The native agent parser ignores golden references, so validate those
separately with `check_eval_quality.py` and deterministic golden-workspace replay.
- A pass with one worker or an enlarged local timeout is not normal execution evidence. Reproduce
with the repository's normal concurrency and declared suite budget; a failure there is a
reliability defect.

### Step 4: Verify the fixtures before touching the skill

Run `python eng/eval-quality/check_eval_quality.py` — it blocks eleven defect classes that can
Run `python eng/eval-quality/check_eval_quality.py` — it blocks 22 defect classes that can
cost a real result here. Then confirm by hand:

- every fixture behaves as its stimulus assumes — a fixture meant to be healthy builds, and one
Expand All @@ -100,6 +107,10 @@ cost a real result here. Then confirm by hand:
- a fixture never states the same fact in two places that disagree — a Cobertura report whose
declared `line-rate`, summary totals and `<line>` elements differ is the canonical case — or the
two arms legitimately read different truths.
- every preservation or scope assertion covers the complete in-scope file set, not one
representative file.
- the golden trajectory and patch pass the deterministic graders, and a realistic broken mutation
fails the grader that is meant to protect the behavior.

### Step 5: Check whether the eval could ever have passed

Expand Down Expand Up @@ -175,6 +186,14 @@ python eng/eval-quality/check_eval_quality.py
./eng/run-skill-evals.sh <plugin> <skill>
```

Use the production path at normal worker concurrency and with the declared `defaults.timeout`:
Vally for skill evals, and `skill-validator evaluate` for agent evals. For an agent eval, separately
run `check_eval_quality.py`, apply each golden patch to its materialized fixture, and run the
applicable deterministic file, output, and command graders against the golden result. Do not use a
serial-only pass or a larger ad hoc budget as completion evidence. For broad routing or behavior
changes, collect separate GPT-family and Claude-family results. Do not pool model families into
extra stimulus votes.

Then request the official run by submitting a PR review containing `/evaluate` (Files changed →
Review changes), which binds the run to the reviewed commit. Before declaring a regression on the
result, confirm the skill payload actually changed — reruns on byte-identical content have shifted
Expand All @@ -185,16 +204,21 @@ result, confirm the skill payload actually changed — reruns on byte-identical
- [ ] For a content fix, a losing trial and the judge's stated reason are quoted in the PR description.
- [ ] The failure was classified before any content was edited.
- [ ] `check_eval_quality.py` and `skill-validator check` both pass.
- [ ] The production skill or agent runner accepts the executable spec.
- [ ] Golden references pass the standalone checker and deterministic replay.
- [ ] The eval completes under normal concurrency and its declared time budget.
- [ ] Golden acceptance and mutation rejection have been demonstrated.
- [ ] Distinct-stimulus count clears the power bar for the target effect and observed tie rate.
- [ ] Isolated **and** plugin activation are both reported.
- [ ] Broad changes have separate GPT-family and Claude-family evidence.
- [ ] The PR body records root cause, fix, and validation so the lesson is reusable.

## Common Pitfalls

| Pitfall | Solution |
|---------|----------|
| Rewriting skill prose in response to an underpowered verdict | Underpowered means too few distinct stimuli; add discriminating stimuli instead |
| Adding `defaults: runs:` to a spec that already has `config:` | Merge into a single `defaults:` block; vally rejects specs with both |
| Using the deprecated top-level `config:` alias | Rename it to `defaults:` and preserve its settings; the repository gate rejects the alias |
| Padding `runs` to clear the stimulus floor | Repeats measure reliability for one task; add stimuli |
| Treating an errored trial as fixture nondeterminism | Read the stderr first; judge-side auth failures need harness fixes |
| Fixing a "wrong" answer that the fixture actually made wrong | Check fixture self-consistency before blaming the response |
Expand All @@ -205,5 +229,5 @@ result, confirm the skill payload actually changed — reruns on byte-identical

- [references/writing-for-baseline-delta.md](references/writing-for-baseline-delta.md) — content patterns that beat the unskilled model
- [references/eval-triage.md](references/eval-triage.md) — symptom, cause and fix catalogue with PR citations
- [eng/eval-quality/README.md](../../../eng/eval-quality/README.md) — the eleven structural gate checks and why each exists
- [eng/eval-quality/README.md](../../../eng/eval-quality/README.md) — the 22 structural gate checks and why each exists
- [eng/vally-adapter/InvestigatingResults.md](../../../eng/vally-adapter/InvestigatingResults.md) — downloading artifacts and reading `results.json`. This is the current guide; the similarly-named `eng/skill-validator/src/docs/InvestigatingResults.md` documents the retired `skill-validator evaluate` schema and does not describe today's results.
Loading
Loading