Tier 5 · Contributor reference. Internal documentation for the
@testdouble/skillwalker-executionpackage — the test-run and test-eval pipelines, the SCIL and ACIL improvement loops, the error hierarchy, and path config. If you're a user looking to run or tune an evaluation, see Getting Started: Skill Trigger Accuracy or the SCIL Evals Guide.
This page is the orchestration reference for Skillwalker. It documents the four entry points (runEvals, runTestEval, runScilLoop, runAcilLoop), the numbered step files behind each pipeline, the SCIL and ACIL loop algorithms and their shared scoring/output/report modules, the path-parameter design, and the error hierarchy. This is where most pipeline changes land.
The @testdouble/skillwalker-execution package owns all test execution orchestration, the SCIL and ACIL improvement loops, and test evaluation pipelines — extracted from the CLI to keep the CLI as a thin Yargs wrapper.
- Last Updated: 2026-05-15
- Authors:
- River Bailey (river@testdouble.com)
- Four high-level orchestrators:
runEvals()for test execution,runTestEval()for result evaluation,runScilLoop()for iterative skill description improvement, andrunAcilLoop()for iterative agent description improvement - All filesystem paths (
outputDir,testsDir,repoRoot) are passed as parameters — the package never callsprocess.cwd()or reads environment variables - Owns the error hierarchy (
SkillwalkerError,ConfigNotFoundError,RunNotFoundError), path config factory, and the step-based pipelines that coordinate the other packages - Sits between the CLI (which parses args and resolves paths) and the lower-level packages (skillwalker-data, skillwalker-evals, claude-integration, sandbox-integration)
Key files:
packages/execution/index.ts— Barrel exports (public API surface)packages/execution/src/evals/run-evals.ts— Test execution orchestratorpackages/execution/src/test-eval/run-test-eval.ts— Evaluation orchestratorpackages/execution/src/scil/loop.ts— SCIL improvement loop orchestratorpackages/execution/src/scil/types.ts—ScilConfiginterface and re-exported SCIL typespackages/execution/src/acil/loop.ts— ACIL improvement loop orchestratorpackages/execution/src/acil/types.ts—AcilConfiginterface and re-exported ACIL types
flowchart TB
cli["CLI (thin wrapper)"]
entry1["runEvals()"]
entry2["runTestEval()"]
entry3["runScilLoop()"]
entry4["runAcilLoop()"]
subgraph pkg["@testdouble/skillwalker-execution"]
direction TB
runsteps["<b>test-runners/steps/</b><br>step-1 resolve · step-2 validate · step-3 read-config<br>step-4 generate-id · step-6 build-flags · step-7 init-totals<br>step-8 run-tests · step-9 print-totals · step-10 exit"]
runners["<b>test-runners/</b><br>prompt/ → runPromptTests()<br>skill-call/ → runSkillCallTests(), buildTempPlugin()<br>agent-call/ → buildTempAgentPlugin()"]
evalpipe["<b>test-eval/</b><br>run-test-eval · test-eval-steps/"]
loops["<b>scil/ · acil/</b><br>loop · steps 1-10"]
common["<b>common/</b><br>score · write-output · print-report"]
lib["<b>lib/</b><br>errors.ts · path-config · metrics.ts · output.ts"]
end
data["skillwalker-data"]
evals["skillwalker-evals"]
claude["claude-integration"]
sandbox["sandbox-integration"]
cli --> entry1 --> runsteps
cli --> entry2 --> evalpipe
cli --> entry3 --> loops
cli --> entry4 --> loops
runsteps --> runners
runners --> claude
evalpipe --> evals
loops --> common
pkg --> data
pkg --> evals
pkg --> claude
claude --> sandbox
| File | Purpose |
|---|---|
packages/execution/index.ts |
Barrel exports — public API for CLI consumption |
packages/execution/src/evals/run-evals.ts |
runEvals() — orchestrates the test-run pipeline |
packages/execution/src/test-eval/run-test-eval.ts |
runTestEval() — orchestrates eval pipeline, converts results |
packages/execution/src/scil/loop.ts |
runScilLoop() — orchestrates the iterative SCIL improvement loop |
packages/execution/src/scil/types.ts |
ScilConfig interface, re-exports ScilTestCase, QueryResult, IterationResult |
packages/execution/src/test-runners/prompt/index.ts |
Prompt test runner — full Claude sessions with output file extraction |
packages/execution/src/test-runners/skill-call/index.ts |
Skill-call test runner — trigger detection with temp plugins |
packages/execution/src/test-runners/agent-prompt/index.ts |
Agent-prompt test runner — agent delegation with output file extraction |
packages/execution/src/test-runners/agent-call/index.ts |
Agent-call test runner — trigger detection with temp agent plugins |
packages/execution/src/test-runners/skill-call/build-temp-plugin.ts |
Temp plugin builder — strips non-triggering fields |
packages/execution/src/test-runners/steps/ |
10 numbered step files for the test-run pipeline |
packages/execution/src/common/score.ts |
scoreResults(), selectBestIteration() — shared scoring logic |
packages/execution/src/common/write-output.ts |
writeIterationOutput(), writeSummaryOutput() — shared output writers with prefix parameter |
packages/execution/src/common/print-report.ts |
printIterationProgress(), printFinalSummary() — shared console reporting |
packages/execution/src/scil/step-1-resolve-and-load.ts |
Resolves target skill and loads skill-call tests |
packages/execution/src/scil/step-5-run-eval.ts |
Runs eval with concurrency pool and majority vote |
packages/execution/src/scil/step-6-score.ts |
Re-exports scoreResults, selectBestIteration from common/score.ts |
packages/execution/src/scil/step-7-improve-description.ts |
Generates improved descriptions via Claude |
packages/execution/src/scil/step-8-apply-description.ts |
Writes best description back to SKILL.md |
packages/execution/src/acil/loop.ts |
runAcilLoop() — orchestrates the iterative ACIL improvement loop |
packages/execution/src/acil/types.ts |
AcilConfig interface, re-exports AcilTestCase, AcilQueryResult, AcilIterationResult |
packages/execution/src/acil/step-1-resolve-and-load.ts |
Resolves target agent and loads agent-call tests |
packages/execution/src/acil/step-3-read-agent.ts |
Reads agent .md frontmatter (name, description, body) |
packages/execution/src/acil/step-5-run-eval.ts |
Runs agent-call eval with concurrency pool and majority vote |
packages/execution/src/acil/step-7-improve-description.ts |
Generates improved agent descriptions via Claude |
packages/execution/src/acil/step-8-apply-description.ts |
Writes best description back to agent .md |
packages/execution/src/lib/errors.ts |
SkillwalkerError, ConfigNotFoundError, RunNotFoundError |
packages/execution/src/lib/path-config.ts |
createPathConfig() — derives all paths from a root directory |
packages/execution/src/lib/metrics.ts |
accumulateTotals() — immutable token/duration accumulator |
packages/execution/src/lib/output.ts |
writeTestOutput() — writes test config and run events to JSONL |
// packages/execution/src/evals/run-evals.ts
interface RunEvalsOptions {
evals: string[] // Eval names to execute
testFilter?: string // Optional: filter to single test by name
debug: boolean // Show sandbox output in real time
outputDir: string // Where to write JSONL output (e.g., tests/output/)
testsDir: string // Root tests/ directory
repoRoot: string // Repository root (parent of testsDir)
}
interface RunEvalsResult {
testRunId: string // Generated timestamp ID (YYYYMMDDTHHmmss)
totalDurationMs: number // Total execution time across all tests
totalInputTokens: number // Total input tokens consumed
totalOutputTokens: number // Total output tokens consumed
failures: number // Number of failed tests
}
// packages/execution/src/test-eval/run-test-eval.ts
interface RunTestEvalOptions {
testRunId?: string // Specific run to evaluate (omit to evaluate all unevaluated)
debug: boolean // Enable debug output
outputDir: string // Where run output is stored
testsDir: string // Root tests/ directory (for rubric file resolution)
}
// packages/execution/src/scil/types.ts
interface ScilConfig {
eval: string // Eval name
skill?: string // Target skill in plugin:skill format (inferred if omitted)
maxIterations: number // Maximum improvement iterations
holdout: number // Fraction held out for validation (0-1)
concurrency: number // Parallel sandbox exec calls
runsPerQuery: number // Runs per test case for majority vote
model: string // Model for improvement prompt
debug: boolean // Show sandbox output in real time
apply: boolean // Auto-apply best description without prompting
outputDir: string // Where to write SCIL output
testsDir: string // Root tests/ directory
repoRoot: string // Repository root
}
// packages/execution/src/acil/types.ts
interface AcilConfig {
eval: string // Eval name
agent?: string // Target agent in plugin:agent format (inferred if omitted)
maxIterations: number // Maximum improvement iterations
holdout: number // Fraction held out for validation (0-1)
concurrency: number // Parallel sandbox exec calls
runsPerQuery: number // Runs per test case for majority vote
model: string // Model for improvement prompt
debug: boolean // Show sandbox output in real time
apply: boolean // Auto-apply best description without prompting
outputDir: string // Where to write ACIL output
testsDir: string // Root tests/ directory
repoRoot: string // Repository root
}
// packages/execution/src/lib/path-config.ts
interface PathConfig {
testsDir: string // = rootDir
skillwalkerDir: string // = rootDir/packages
repoRoot: string // = rootDir/..
outputDir: string // = rootDir/output
dataDir: string // = rootDir/analytics
}
// packages/execution/src/scil/step-3-read-skill.ts
interface SkillFileContent {
name: string // Skill name from frontmatter
description: string // Current description from frontmatter
frontmatterRaw: string // Raw YAML frontmatter
body: string // Markdown body after frontmatter
fullContent: string // Complete file content
}
// packages/execution/src/lib/errors.ts
class SkillwalkerError extends Error // Base error — caught at CLI for clean exit
class ConfigNotFoundError extends SkillwalkerError // tests.json not found
class RunNotFoundError extends SkillwalkerError // Test run directory not foundThe execution package never resolves paths from process.cwd(). The CLI owns path resolution via createPathConfig(process.cwd()) and passes the individual fields as parameters:
// CLI (packages/cli/src/paths.ts) — the only place process.cwd() is called
import { createPathConfig } from '@testdouble/skillwalker-execution'
const config = createPathConfig(process.cwd())
export const outputDir = config.outputDir
export const testsDir = config.testsDir
export const repoRoot = config.repoRoot
// CLI command (packages/cli/src/commands/test-run.ts) — passes paths explicitly
const result = await runEvals({
evals, testFilter, debug,
outputDir, testsDir, repoRoot,
})Paths flow through every layer — from orchestrator to step to runner — as explicit function parameters. This makes the package fully testable without filesystem mocks for path resolution.
The runEvals function orchestrates a 10-step pipeline for each eval:
| Step | File | Purpose |
|---|---|---|
| 1 | step-1-resolve-paths.ts |
Joins testsDir + evals/ + eval name |
| 2 | step-2-validate-config.ts |
Validates tests.json exists (throws ConfigNotFoundError) |
| 3 | step-3-read-config.ts |
Reads config, applies test filter, validates scaffolds |
| 4 | step-4-generate-run-id.ts |
Generates timestamp ID (YYYYMMDDTHHmmss) |
| 6 | step-6-build-flags.ts |
Resolves plugin directories from config + repoRoot |
| 7 | step-7-init-totals.ts |
Initializes zeroed accumulator |
| 8 | step-8-run-test-cases.ts |
Dispatches to prompt or skill-call runner by test type |
| 9 | step-9-print-totals.ts |
Prints run summary to stdout |
| 10 | step-10-exit.ts |
exitWithResult(failures) — exits 0 or 1 |
Step 5 is intentionally absent from the numbering. Step 8 splits tests by type:
// packages/execution/src/test-runners/steps/step-8-run-test-cases.ts
const promptTests = config.tests.filter(t => t.type === 'prompt' || t.type === undefined)
const skillCallTests = config.tests.filter(t => t.type === 'skill-call')Prompt runner (test-runners/prompt/index.ts): Reads the prompt file, runs Claude in Docker with all configured plugins, parses stream-JSON output, extracts metrics, extracts output files from the sandbox, and writes JSONL (including output-files.jsonl).
Skill-call runner (test-runners/skill-call/index.ts): Same flow, but first builds a temporary plugin via buildTempPlugin(skillFile, runDir, repoRoot). The temp plugin has a stripped SKILL.md with only name + description and a no-op body that responds "skill triggered."
Agent-prompt runner (test-runners/agent-prompt/index.ts): Same as the prompt runner but dispatches to an agent via delegation.
Agent-call runner (test-runners/agent-call/index.ts): Same as the skill-call runner but builds a temporary agent plugin.
All four runners share the same post-run output file extraction step: after Claude finishes, they call extractOutputFiles() from claude-integration to retrieve any files the skill/agent wrote inside the sandbox, then call appendOutputFiles() from skillwalker-data to persist them to output-files.jsonl in the run directory.
// packages/execution/src/test-runners/skill-call/build-temp-plugin.ts
function stripNonTriggeringFields(frontmatter: string): string {
return frontmatter
.replace(/^allowed-tools:.*$/m, '')
.replace(/^argument-hint:.*$/m, '')
.replace(/\n{2,}/g, '\n')
.trim()
}Two variants: buildTempPlugin (original description) and buildTempPluginWithDescription (substitutes a custom description for SCIL iterations). Both create a minimal .claude-plugin/plugin.json with version 0.0.0.
When testRunId is provided, evaluates that specific run. When omitted, scans outputDir for all directories and filters to unevaluated ones (those without a non-empty test-results.jsonl).
For each run:
- Resolves the run directory via
resolveRunDir(id, outputDir) - Reads
test-config.jsonlto determine the eval - Calls
evaluateTestRun()from@testdouble/skillwalker-evals - Converts
EvalResulttoTestResultRecord[]— boolean evals produce one record; LLM-judge evals produce per-criterion records plus an aggregate - Writes results to
test-results.jsonl - Marks forced re-evaluations for analytics reprocessing
The runScilLoop function orchestrates 10 steps iteratively:
- Resolve and load — Finds the target skill (explicit or inferred) and loads skill-call tests
- Split sets — Stratified train/test split based on holdout fraction (delegates to skillwalker-data)
- Read skill — Parses SKILL.md frontmatter and body
- Build temp plugin — Creates temp plugin with current (or improved) description
- Run eval — Concurrent sandbox execution with configurable concurrency and majority voting
- Score — Computes train/test accuracy; selects best iteration
- Improve description — Sends failure details and phase-specific instructions to Claude to generate a better description (max 1024 chars). During explore/transition phases, always generates a new description regardless of accuracy. During converge, only improves if accuracy is imperfect
- Apply description — Writes best description to SKILL.md (interactive prompt or
--apply) - Write output — Persists iteration JSONL (including phase) and summary JSON
- Print report — Iteration progress (with phase tag) and final comparison table (with phase column)
Each iteration is assigned a phase (explore, transition, or converge) via getPhase() from skillwalker-data. Early exit on perfect accuracy only occurs during or after the converge phase.
// packages/execution/src/scil/step-5-run-eval.ts — simplified
const pending = new Set<Promise<void>>()
for (const item of workItems) {
const task = runSingleQuery(item.test, item.runIndex, opts)
const tracked = task.then(() => { pending.delete(tracked) })
pending.add(tracked)
if (pending.size >= opts.concurrency) {
await Promise.race(pending)
}
}
await Promise.all(pending)Results are stored in a pre-sized array by work-item index for deterministic ordering. When runsPerQuery > 1, results are grouped by test name and aggregated via majority vote.
// packages/execution/src/common/score.ts
// Primary: test accuracy (when holdout > 0) or train accuracy (when holdout = 0)
// Tiebreaker: train accuracy
// Ties at both levels: earlier iteration winsrunAcilLoop mirrors the SCIL loop for agents: it iteratively improves an agent description by running agent-call evaluations and asking Claude for a better description. The ACIL pipeline consists of 10 numbered steps in packages/execution/src/acil/:
| Step | File | Function | Description |
|---|---|---|---|
| 1 | step-1-resolve-and-load.ts |
resolveAndLoad |
Filter agent-call tests, resolve agent .md path, validate identifier format |
| 2 | step-2-split-sets.ts |
splitSets |
Re-export from data package — deterministic stratified train/test split |
| 3 | step-3-read-agent.ts |
readAgent |
Parse agent frontmatter and body, return name/description/body |
| 4 | step-4-build-temp-plugin.ts |
buildTempPlugin |
Delegate to buildTempAgentPluginWithDescription |
| 5 | step-5-run-eval.ts |
runEval |
Execute tests using evaluateAgentCall, return AcilQueryResult[] |
| 6 | step-6-score.ts |
scoreResults |
Re-export from common/score.ts |
| 7 | step-7-improve-description.ts |
improveDescription |
Build ACIL improvement prompt, run Claude, validate result |
| 8 | step-8-apply-description.ts |
applyDescription |
Write improved description to agent .md file |
| 9 | step-9-write-output.ts |
writeOutput |
Delegate to common/write-output.ts with prefix: 'acil' |
| 10 | step-10-print-report.ts |
printReport |
Re-export from common/print-report.ts |
Steps 6, 9, and 10 are thin re-export wrappers around shared modules in common/, which are also used by the SCIL pipeline.
ACIL and SCIL share three modules in packages/execution/src/common/:
common/score.ts—scoreResults()andselectBestIteration()using genericScoreable/ScoredIterationinterfacescommon/write-output.ts—writeIterationOutput()andwriteSummaryOutput()usingWritableIterationinterface, parameterized byprefixcommon/print-report.ts—printIterationProgress()andprintFinalSummary()usingPrintableResult/PrintableIterationinterfaces
| Error Class | Thrown By | Trigger |
|---|---|---|
SkillwalkerError |
Multiple steps | General errors (prompt not found, config read failure, missing frontmatter) |
ConfigNotFoundError |
step-2-validate-config |
tests.json not found in eval directory |
RunNotFoundError |
step-1-resolve-run-dir |
Test run directory does not exist in outputDir |
The CLI catches SkillwalkerError at the top level and writes the message to stderr with exit code 1.
| Constant | Value | Location | Description |
|---|---|---|---|
NOOP_BODY |
'\nRespond with: "skill triggered" — nothing else.\n' |
build-temp-plugin.ts |
Body text for stripped temp plugins |
MAX_DESCRIPTION_LENGTH |
1024 |
step-7-improve-description.ts |
Maximum allowed skill description length |
packages/execution/src/lib/*.test.ts— Tests for errors, path-config, metrics, outputpackages/execution/src/common/*.test.ts— Tests for shared scoring, output writing, and reporting utilitiespackages/execution/src/test-runners/steps/*.test.ts— Tests for each pipeline steppackages/execution/src/test-eval-steps/*.test.ts— Tests for eval pipeline stepspackages/execution/src/scil/*.test.ts— Tests for each SCIL step and the loop orchestratorpackages/execution/src/acil/*.test.ts— Tests for each ACIL step and the loop orchestrator
Tests are co-located with source files. Tests that previously mocked paths.js singletons now pass path values as function parameters, eliminating the need for path mocking. Shared test fixtures are in packages/execution/src/test-runners/steps/fixtures.ts, which imports JSON fixtures from @testdouble/test-fixtures.
- Skillwalker Architecture — System-wide architecture, package boundaries, and dependency graph
- CLI Package — The thin CLI wrapper that delegates to this package
- Data Package — Shared data layer consumed by execution orchestrators
- Evals Package — Evaluation engine called by
runTestEval - Claude Integration — Claude CLI wrapper used for running prompts in sandbox
- Sandbox Integration — Test Sandbox API used for sandbox lifecycle
- Skill Call Improvement Loop — Detailed SCIL algorithm and design
- Step-Based Pipeline — Coding standard for the numbered-step architecture
- Custom Error Hierarchy — Error class conventions
- Agent Call Improvement Loop — ACIL mechanics: agent detection, temp plugin isolation, holdout splits, scoring
Next: CLI Package — the thin Yargs wrapper that resolves paths and calls into these orchestrators.
Related: Skill Call Improvement Loop — the SCIL algorithm and design behind runScilLoop; Agent Call Improvement Loop — the same for runAcilLoop.