From 45a8ee85828343db2fcd8837c6a97596a2273fbe Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Mats=20B=C3=B8e=20Bergmann?= Date: Wed, 7 Oct 2026 10:19:56 +0200 Subject: [PATCH 1/3] docs: inventory runner contracts for eval-engine extraction (#65) ## Scope Worker B for #58 (Task 1/9): post-PR-#45 runner contract/invariant inventory. This PR is intentionally documentation-only. It does **not** implement the eval engine, change `invoke`, migrate Loom, add a public `eval` CLI, or define the final Task 1 architecture. ## What it captures - one `invoke` = exactly one isolated invocation; - `opencode-eval-runner/v1` result boundary; - mandatory/authoritative `runtime_evidence/v1`; - assertion-scoped evidence readiness and field states; - product vs evidence vs infrastructure failure separation; - inner vs outer timeout behavior; - retry ownership above `invoke`, cross-checked against Loom's current narrow retry policy; - OpenCode vs Copilot transport capability differences; - host output validation and host/image compatibility; - CLI and GitHub Action interface limits; - explicit unsupported areas and trusted-checkout scope; - current Loom behaviors that must not accidentally become runner contracts. ## Validation Cross-checked against the landed post-#45 contract: - `main` / PR #45 merge commit `fd9da10cbe2a8182fc8910ec199a221250deb3ca` - final reviewed PR #45 head `25478106773969870931af3fb34e7b2586746606` - `runner/cli.py` - `container/invoke.py` - `docs/invocation-usage.md` - `docs/runtime-evidence-contract.md` - `action.yml` - Loom `functionality-anchor-requirements-coherance` - `scripts/run-evals.py` blob `30f2ae87764be180173a9e5b4b609ee79b2f063d` `eval-engine/01-architecture` now matches `main` at the PR #45 merge commit. The worker branch has been rebased onto that exact state, leaving a one-file documentation diff. Part of #58. --- docs/eval-engine-runner-contracts.md | 213 +++++++++++++++++++++++++++ 1 file changed, 213 insertions(+) create mode 100644 docs/eval-engine-runner-contracts.md diff --git a/docs/eval-engine-runner-contracts.md b/docs/eval-engine-runner-contracts.md new file mode 100644 index 0000000..c535768 --- /dev/null +++ b/docs/eval-engine-runner-contracts.md @@ -0,0 +1,213 @@ +# Eval-engine runner contract and invariant matrix + +Issue: #58 — Task 1 worker B +Scope: runner-contract inventory only; no eval-engine implementation or final architecture selection. + +## Source snapshot + +This inventory was validated against: + +- `bateau84/opencode-eval-runner` `main` at PR #45 merge commit `fd9da10cbe2a8182fc8910ec199a221250deb3ca` (final reviewed PR head `25478106773969870931af3fb34e7b2586746606`), including: + - `runner/cli.py` + - `container/invoke.py` + - `docs/invocation-usage.md` + - `docs/runtime-evidence-contract.md` + - `action.yml` +- `bateau84/loom` branch `functionality-anchor-requirements-coherance`, `scripts/run-evals.py` blob `30f2ae87764be180173a9e5b4b609ee79b2f063d`. + +PR #45 is merged. The integration branch `eval-engine/01-architecture` matches `main` at merge commit `fd9da10cbe2a8182fc8910ec199a221250deb3ca`, so the final Task 1 synthesis can consume this inventory against the landed runner contract. + +## Boundary in one sentence + +`opencode-eval-runner invoke` is a **single isolated invocation boundary**. It executes one target or judge call and returns one validated `opencode-eval-runner/v1` result containing the authoritative `runtime_evidence/v1` object. Cases, retries, iterations, target/judge sequencing, assertions, semantic grading, thresholds, aggregation, and final `pass | fail | non-evidence` policy belong above this boundary. + +## Contract / invariant matrix + +| Area | Post-#45 runner contract | Obligation on a future eval orchestrator | Must not be inferred or moved into `invoke` | +| --- | --- | --- | --- | +| Invocation cardinality | One `invoke` command performs one isolated transport invocation. | Treat every target attempt and every judge attempt as a separate invocation/result. | Hidden retry loops, multi-case execution, target+judge pairing, iterations. | +| Public command | Only `opencode-eval-runner invoke` exists. | Build orchestration on top of this boundary or an internal equivalent that preserves the same semantics. | A public `eval` CLI in Task 1. | +| Input API | No input-JSON request schema exists. Inputs are CLI/Action fields plus UTF-8 prompt/system/config seed files. | Normalize project cases into explicit invocation inputs. | An invented `invocation/v1` JSON request contract. | +| Supported transports | `opencode` and `github-copilot-cli`. | Select transport explicitly for each target/judge invocation. | Assuming equal runtime-observation capability across transports. | +| Isolation | Fresh OCI process, read-only root, tmpfs `/tmp`, dropped caps, no-new-privileges, explicit workspace/mount mode, fresh transport state. | Preserve runner-owned isolation by using the supported host boundary. | Direct transport-container entrypoint as a public contract. | +| Workspace default | `--workspace-mode ro` by default; `rw` is explicit. | Project/profile decides whether a case needs a writable workspace. | A global assumption that evals may mutate the source checkout. | +| Network default | No runner network override unless `--network` is supplied. | Project/profile opts into `host` or another mode when required. | Implicit host networking. | +| Result schema | Each official invocation emits one JSON object with `schema: "opencode-eval-runner/v1"`. | Persist the result as the low-level invocation result and preserve its provenance. | Treating model prose or workspace files as the result contract. | +| Runtime evidence | Every official result must include `runtime_evidence.schema == "opencode-eval-runner/runtime-evidence/v1"`. | Validate and use this object for runtime-evidence readiness. | Any second authoritative observer/evidence projection. | +| Evidence authority | `runtime_evidence/v1` is the only authoritative runtime-evidence object. | Make runtime assertions only from eligible facts in this object. | Backfilling from `tools`, `actions`, `tool_result_evidence`, stdout/stderr, Session/model text, or workspace files. | +| Evidence top-level status | `complete | incomplete | unsupported | invalid`. | Block PASS for all runtime assertions on overall `incomplete` or `invalid`; interpret `unsupported` by required scope. | Treating product success as proof that evidence is complete. | +| Evidence boundaries | `native`, `code_mode_execution`, `code_mode_finality`. | Each runtime assertion declares the boundary/boundaries it requires. | A single global boolean as sufficient evidence-readiness policy. | +| Exact field states | `available | redacted | omitted | unsupported`. | Exact-value assertions require `available`. | Treating redacted/omitted/unsupported values as negative facts or empty values. | +| Unknown coverage | Unknown is never converted to zero. | Keep unknown/unsupported distinct from observed absence. | Passing an absence assertion because a count is missing. | +| Native authority | Direct/native calls include authoritative identity, tool, actor, Session/message/CallID, executable input, ordering, terminal outcome, ancestry. | Consume only fields needed by project assertions. | Reconstructing invocation identity from order, tool name, input equality, or model output. | +| Code Mode execution | Inner call identity/tool/input/parent/outcome/ordering are supportable at `code_mode_execution`. | Assertions over these facts require that boundary to be `complete`. | Assuming exact final caller-visible value/error is known. | +| Code Mode finality | Exact script-visible final value/error is `unsupported` on stock OpenCode 2.0.23. | Mark an assertion requiring this boundary as unsupported/non-evidence. | Converting transformed-handler result/error into final caller evidence. | +| Product vs evidence outcome | `result.exit_code`, timeout, and behavioral/product success are separate from evidence eligibility. | Evaluate transport/product health and evidence readiness independently. | `exit_code == 0` => evidence complete, or evidence complete => behavioral PASS. | +| Inner timeout | `--timeout-seconds`, default 240s, limits the transport invocation inside the container. OpenCode timeout can still emit a structured result with `exit_code: 124`, `timed_out: true`, and timeout-aware runtime evidence. | Treat timeout as target/judge transport failure/non-evidence unless project policy explicitly says otherwise; retain the result artifact when available. | Retrying inside `invoke` or calling timeout a behavioral FAIL. | +| Outer timeout | `--container-timeout`, default 300s, limits the OCI process from the host wrapper. A host-side timeout makes the CLI return infrastructure failure and may leave no result file. | Orchestration owns whether an explicit retry policy applies; record the failed attempt. | Assuming every failed `invoke` has a result JSON. | +| CLI exit vs result exit | Host CLI exit reflects whether the isolated runner/container completed its contract; `result.exit_code` reflects the invoked product/transport result. They are not interchangeable. | Inspect both process outcome and result fields when classifying an attempt. | Using host exit code alone to decide target success. | +| Output parsing | Host requires container stdout to parse as a JSON object. | Treat missing/invalid JSON as infrastructure/non-evidence. | Parsing arbitrary stdout fragments as a successful result. | +| Output validation | Container validates `runtime_evidence` before stdout serialization; host validates it again before writing `--output`. | Reject missing/malformed authoritative evidence rather than degrading to diagnostics. | Assuming the host currently schema-validates every top-level `result/v1` field; its hard validator is specifically the object shape plus `runtime_evidence/v1`. | +| Host/image compatibility | Current host requires official/custom image output to include valid `runtime_evidence/v1`; legacy images without it are rejected. | Treat host+image as a compatibility pair and record selected image/transport configuration. | Silent fallback to an older result shape. | +| OpenCode reviewed runtime | Runtime-observer capability is reviewed against stock OpenCode 2.0.23. | An OpenCode upgrade requires capability/acceptance revalidation. | Assuming later OpenCode versions preserve the same observation boundaries. | +| Copilot runtime evidence | Copilot emits valid result shape but `runtime_evidence.status = unsupported` because no OpenCode runtime observer exists. | Copilot can be used for pure model/judge work where OpenCode runtime assertions are not required. | Runtime tool assertions from Copilot convenience output. | +| Expected plugin | OpenCode-only zero-inference preflight can fail before inference when required plugin activation is unavailable. | Treat preflight failure as infrastructure/non-evidence. | Behavioral FAIL for missing runner/runtime prerequisites. | +| Skills | `--skill` records skill-under-test intent; it does not force-load a skill. `skills_loaded` is observed convenience data. | Project-owned assertions decide whether the intended skill was used and which evidence they require. | Runner-owned skill PASS/FAIL semantics. | +| Reasoning | Optional explicit reasoning; otherwise provider default. OpenCode maps to model variant, Copilot to effort. | Preserve requested reasoning and result-reported source. | Guessing provider defaults or silently substituting effort. | +| GitHub Action modes | Three modes: repository `command` (highest precedence), direct single invocation when `model` is set, or setup-only. | Use command/setup mode when full orchestration or full CLI surface is needed. | Treating Action direct mode as a suite/eval engine. | +| Action surface | Direct mode exposes a subset of CLI options. It does not directly expose network, expected-plugin, seed/config paths, arbitrary env forwarding, or print-result. | Full-feature orchestration should use the host CLI after setup or Action `command` mode. | Designing engine semantics around Action-input limitations. | +| Behavioral semantics | Runner docs explicitly exclude cases, suites, assertions, judge semantics, thresholds, and final verdicts. | Project/profile extension code owns these semantics. | Universal assertion DSL or runner-owned domain rubric. | +| Trust scope | Normal profile is trusted-checkout with reviewed same-process instrumentation. | Preserve documented trust profile in artifacts/config. | Claiming protection from deliberately hostile same-process evaluated plugins. | + +## Evidence-readiness decision procedure + +The future eval layer must preserve the post-#45 consumer rule exactly enough that runtime evidence cannot be upgraded by orchestration: + +1. Require a valid `runtime_evidence/v1` object. +2. If overall status is `incomplete` or `invalid`, no runtime assertion may produce PASS. +3. For each runtime assertion, project/profile code declares the required boundary set. +4. Every required boundary must be `complete`. +5. If the assertion requires an exact field, that field must be `available`. +6. `redacted` or `omitted` means the exact-value assertion lacks complete evidence. +7. `unsupported` means the required fact/boundary is unsupported. +8. Never fill an authoritative gap from diagnostic/convenience fields. + +The eval layer may generalize the mechanics of applying a declared evidence requirement, but the **choice of required facts and what they mean behaviorally remains project-owned**. + +## Timeout and failure handling + +There are three distinct failure planes and they must remain separate. + +### 1. Invocation/product failure + +Examples: + +- model/provider returns non-zero; +- OpenCode invocation times out internally and returns `exit_code: 124`; +- assistant text is empty or unusable for a project-required judge/target step. + +This can produce a valid `result/v1` artifact. It is not automatically a behavioral failure. + +### 2. Evidence-readiness failure + +Examples: + +- overall runtime evidence is `incomplete` or `invalid`; +- a required boundary is not `complete`; +- a required exact field is redacted, omitted, or unsupported; +- Code Mode finality is required on stock 2.0.23. + +This is non-evidence for the affected assertion, even if the product invocation exited successfully. + +### 3. Runner/infrastructure failure + +Examples: + +- outer OCI timeout; +- invalid/non-object JSON from the container; +- invalid/missing `runtime_evidence/v1`; +- workspace/config/preflight failure. + +This is non-evidence. The orchestrator may apply an explicit retry policy, but `invoke` must not retry internally. + +## Retry invariant and Loom reference behavior + +The low-level runner has **no retry loop**. This is a hard boundary for the eval-engine extraction. + +Current Loom puts retry policy in `invoke_container_with_retry()`, above `invoke_container()`: + +- default: 2 retries (3 total attempts); +- accepted range: 0–5 retries; +- each attempt calls the runner boundary again; +- successful retries attach attempt count and prior errors to the returned result; +- exhausted retries retain all prior errors; +- backoff is bounded exponential (1s, then 2s); +- retries are only considered for the narrow provider-routing failure containing both `provider.no-route` and `model unavailable`; +- no retry is allowed after observable tools/actions, because replay could duplicate side effects. + +This Loom policy is a useful proven orchestration policy, **not** a low-level runner contract. Task 1's final architecture may generalize the policy mechanism, but any retry must remain explicit, artifact-visible, and outside `invoke`. + +## Public interface constraints + +### Host CLI + +The current host command is: + +```text +opencode-eval-runner invoke [options] +``` + +Important constraints for the future orchestrator: + +- model and prompt file are required; +- output path is required; +- one result file is written only after host validation; +- prompt/system content is file-based, not an input JSON request; +- direct container invocation is an implementation detail; +- no suite/case/iteration/judge command exists. + +### GitHub Action + +The Action is an adapter around the same boundary: + +1. `command` set: prepare runner/images, then execute repository-owned orchestration; +2. no `command`, `model` set: one direct invocation; +3. neither set: setup only. + +Therefore the future generic eval engine must not be designed as a special Action-only feature. Action `command` or setup mode can host it later without changing `invoke`. + +## Explicit unsupported / out-of-scope areas + +These are constraints to preserve, not gaps for Task 1 to solve: + +- exact Code Mode caller-final value/error on stock OpenCode 2.0.23; +- OpenCode runtime observation for the `github-copilot-cli` transport; +- resistance to a deliberately hostile evaluated plugin sharing the trusted OpenCode process; +- input JSON request API; +- public `eval` CLI; +- universal assertion language; +- runner-owned cases/suites/judge semantics/thresholds/final verdict policy. + +## Loom cross-check + +The current Loom script validates several extraction assumptions. + +### Correct behavior to preserve generically + +- Target and judge are separate invocations. +- Transport retry is above the invocation boundary. +- Retry attempts/errors are observable in result artifacts. +- Runtime cases have independent concurrency policy from non-runtime cases. +- Transport/infrastructure errors become `non-evidence`, not behavioral failure. +- Case/iteration artifacts keep target result, judge result, deterministic failures, semantic result, timing, and classification. + +### Existing Loom behavior that is not the post-#45 runner contract + +These are migration facts, not reasons to weaken the new runner boundary: + +- Loom still contains a direct OCI/container fallback when the `opencode-eval-runner` binary is unavailable. The post-#45 runner documentation says direct container entrypoints are implementation details, so this fallback must not become a generic eval-engine contract. +- Loom's current `target_scoring_evidence_error()` still checks legacy exactness for `text`, `tools`, `actions`, and `skills_loaded`. Post-#45 runtime assertions must instead apply `runtime_evidence/v1` boundary/field readiness. This migration belongs to later integration work, not this documentation-only worker. +- Loom has legacy observer/evidence-safety compatibility paths alongside the normal runner path. They are not additional authoritative runtime-evidence contracts after #45. +- Loom's behavioral classification names include `behavioral-fail`; issue #58's future generic engine will normalize final engine classification separately. This worker does not define that architecture. + +## Invariants handed to Task 1 synthesis + +The final architecture must not violate these runner-level facts: + +1. **One invoke = one isolated invocation.** +2. **No hidden retries in invoke.** +3. **Retries are orchestration policy and must be artifact-visible.** +4. **`opencode-eval-runner/v1` is the single invocation result envelope.** +5. **`runtime_evidence/v1` is mandatory and authoritative for runtime facts.** +6. **Evidence readiness is assertion-scoped by required boundaries and fields.** +7. **Diagnostics cannot repair missing authoritative evidence.** +8. **Product outcome, evidence outcome, and infrastructure outcome are separate.** +9. **Timeouts exist at inner invocation and outer OCI layers.** +10. **Project/profile code owns behavioral meaning, assertions, judge semantics, and thresholds.** +11. **No universal assertion DSL is implied by the runner contract.** +12. **No input JSON API or public eval CLI exists yet.** +13. **Copilot runtime evidence is explicitly unsupported.** +14. **Code Mode final caller value/error is explicitly unsupported on stock 2.0.23.** +15. **The normal trust profile is trusted-checkout, not hostile-plugin resistance.** + +This matrix is intentionally a constraint inventory. The sibling Loom-responsibility inventory and the Task 1 integration owner should use it as a non-negotiable boundary when defining the generic orchestration modules for Tasks 2–5. From bdcf9e3c0c0d7066f64b17d4a92fc036604dbed8 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Mats=20B=C3=B8e=20Bergmann?= Date: Wed, 7 Oct 2026 10:20:03 +0200 Subject: [PATCH 2/3] docs: inventory Loom eval runner responsibilities (#66) ## Summary Part of #58 (Task 1/9), parallel worker A. Adds a focused inventory of the current Loom eval harness from: - bateau84/loom branch functionality-anchor-requirements-coherance - scripts/run-evals.py at 30f2ae87764be180173a9e5b4b609ee79b2f063d The inventory traces and classifies: - case loading/normalization; - target/workspace setup; - invocation and orchestration retries; - runtime evidence and legacy observers; - judge construction/parsing; - deterministic checks; - classification; - artifacts; - iterations/concurrency; - skill ablation; - reporting. It separates each responsibility into generic orchestration, Loom/project semantics, low-level invoke behavior, or legacy/compatibility behavior that should not migrate. ## Post-#45 validation Cross-checked against the merged runner contracts on the integration base: - one invoke remains one isolated invocation; - retries remain outside invoke; - runtime_evidence/v1 is the only authoritative runtime-evidence object; - tools/actions/tool_result_evidence/stdout/stderr/model text are diagnostic only; - project code owns cases, assertions, judging, thresholds, and behavioral meaning. The document explicitly marks Loom's older direct-container fallback, observer/tool-result reconstruction, runner-safety adapter, and Loom-side evidence_safety eligibility machinery as non-migration paths. ## Scope Documentation only. No runtime changes, no Loom migration, no final engine API/module design, no universal assertion DSL, and no public eval CLI. PR target: eval-engine/01-architecture. --- ...om-eval-runner-responsibility-inventory.md | 686 ++++++++++++++++++ 1 file changed, 686 insertions(+) create mode 100644 docs/loom-eval-runner-responsibility-inventory.md diff --git a/docs/loom-eval-runner-responsibility-inventory.md b/docs/loom-eval-runner-responsibility-inventory.md new file mode 100644 index 0000000..5fd4927 --- /dev/null +++ b/docs/loom-eval-runner-responsibility-inventory.md @@ -0,0 +1,686 @@ +# Loom eval runner responsibility inventory + +Status: research inventory for issue #58, parallel worker A. + +This document is intentionally not the final generic eval-engine architecture. It records what the current Loom harness does, classifies ownership, and identifies behavior that must not be copied into the future engine. The integration owner can use this inventory together with the runner-contract inventory to synthesize the architecture. + +## Sources pinned for this inventory + +- Loom repository: bateau84/loom +- Loom branch: functionality-anchor-requirements-coherance +- Loom file: scripts/run-evals.py +- Loom file SHA: 30f2ae87764be180173a9e5b4b609ee79b2f063d +- Post-PR-#45 opencode-eval-runner integration base: fd9da10cbe2a8182fc8910ec199a221250deb3ca +- Runtime evidence contract file SHA: 01f7c72b57e468b8189927f148f9aaf8a2caac45 +- Invocation usage file SHA: 5e82aa91291fb12230c3ddba9eec11db7c190b5d +- Runner README result-contract file SHA: f93709bfe21116cfb7fad7a5e38b4f12b5ac9fa4 + +The post-#45 runner contract is treated as authoritative where it conflicts with historical Loom-side observation or evidence code. + +## Classification legend + +This inventory uses the four ownership classes required by issue #58: + +1. **Generic orchestration** — reusable sequencing or run-management behavior that is not Loom-specific. +2. **Project/profile** — Loom-owned case semantics, setup, assertions, judge policy, scoring, or execution-profile policy. +3. **Low-level invoke** — behavior already owned by opencode-eval-runner invoke and its result/runtime-evidence contracts. +4. **Legacy/compatibility — do not migrate** — historical Loom-side transport, evidence, or compatibility machinery superseded by the post-#45 runner boundary. + +Some current functions mix classes. Those functions are split by responsibility below instead of being assigned wholesale. + +## Current end-to-end lifecycle + +For a normal non-ablation case, scripts/run-evals.py currently performs this sequence: + +1. Discover Loom suites and skill-owned suites. +2. Select cases and normalize skill-owned legacy shapes. +3. Expand selected cases into a case × iteration job matrix. +4. Create an isolated target project and judge project. +5. Materialize Loom agents, skills, fixtures, and runtime plugin inputs. +6. Build target invocation inputs. +7. Invoke the target through opencode-eval-runner when available, otherwise through a historical direct-container fallback. +8. Apply Loom-side transport projection and historical observer/evidence processing. +9. Apply a Loom-side target evidence gate. +10. Run deterministic Loom assertions. +11. Build and invoke the semantic judge when the target is usable. +12. Parse and validate the Loom judge response. +13. Classify the run as pass, behavioral-fail, or non-evidence. +14. Write one per-case/per-iteration artifact, hash it, and verify the durable artifact before counting it. +15. Report per-case status and a final passed/total count. + +Skill-owned cases replace the single target/judge sequence with baseline target → baseline judge → candidate target → candidate judge, then calculate an ablation delta and skill-value classification. + +This lifecycle contains a reusable orchestration skeleton, but almost every input and semantic decision around that skeleton is Loom-owned. + +## Responsibility inventory + +### 1. Generic orchestration responsibilities + +These are the reusable orchestration behaviors demonstrated by Loom. + +| Current Loom behavior | Current location | Why it is generic | +| --- | --- | --- | +| Build a run identity and case/iteration identity | main, write_case_artifact | A multi-case engine needs stable run and job identity independent of project semantics. | +| Expand selected cases into case × iteration jobs | main: jobs construction | Iteration is orchestration, not a behavioral assertion. | +| Sequence target execution before judging | run_case | The project supplies target inputs and judge semantics; the engine controls phase order. | +| Do not run the judge when the target invocation is unusable | run_case | Phase gating is orchestration. The reason a target is unusable comes from invoke/evidence/project logic. | +| Track target, judge, and total timings | run_case, run_skill_ablation_case | Timing/provenance is generic run metadata. | +| Execute explicit retry attempts outside invoke | invoke_container_with_retry | The loop is orchestration. One invoke remains one isolated invocation. | +| Preserve retry attempt/error history in the invocation result | invoke_container_with_retry | Retry provenance must remain visible rather than hidden. | +| Produce a top-level run classification that distinguishes evidence failure from behavioral failure | classify_behavioral_result | The tri-state outcome concept is generic, although Loom currently names fail as behavioral-fail and supplies Loom semantic inputs. | +| Create and claim an artifact directory for one run | claim_artifact_directory | Preventing accidental cross-run artifact mixing is generic run management. | +| Write per-job artifacts atomically | write_case_artifact | Durable per-run evidence is generic orchestration. | +| Hash and verify the durable artifact before counting/reporting it | artifact_evidence_id, verify_case_artifact | Artifact integrity between in-memory result and durable result is generic. | +| Run jobs sequentially or through a bounded worker pool | execute_group in main | Bounded concurrency is generic orchestration. | +| Aggregate per-job success into process success/failure | report and main | Run-level completion/exit status is orchestration. | + +#### Important split: retry mechanism versus retry policy + +The current retry loop is reusable orchestration, but its hard-coded retryability rule is not a generic contract: + +- retries are limited to 0..5; +- only provider.no-route plus Model unavailable is considered transient; +- any observed tool/action disables retry because replay could duplicate side effects; +- backoff is capped at two seconds; +- exhausted retries remain non-evidence. + +The invariant to preserve is structural: retry happens outside invoke and every attempt is a separate invocation. The exact provider-error predicate and replay-safety evidence used by Loom are current policy, not a universal behavioral rule. + +#### Important split: concurrency mechanism versus Loom runtime policy + +Thread-pool execution and concurrency caps are generic. Loom's division into runtime and non-runtime groups is project/profile policy: + +- non-runtime work uses --parallel; +- runtime work defaults to serial execution; +- --runtime-parallel greater than one is explicitly treated as load/stress; +- non-runtime jobs are executed as one group before runtime jobs. + +That policy exists because one Loom runtime case can itself dispatch multiple model-backed subagents. It should be recorded as a Loom profile choice, not assumed to be universally correct. + +### 2. Loom/project/profile responsibilities + +These behaviors encode Loom's corpus, product topology, behavioral meaning, or benchmark policy and must stay project-owned. + +#### Case discovery and normalization + +| Function/behavior | Location | Project responsibility | +| --- | --- | --- | +| Discover evals/*.json | behavioral_eval_files | Loom suite layout. | +| Interpret suite/case default opt-in flags | load_cases | Loom corpus selection semantics. | +| Discover skills//evals/*.json | skill_eval_files | Loom skill repository layout. | +| Accept legacy skill eval shapes | skill_eval_raw_cases | Loom compatibility for existing case files. | +| Convert skill-owned cases to SKILL-- runtime cases | normalize_skill_eval_case | Loom-specific normalized case semantics. | +| Enforce skill folder/declaration agreement and unique skill case IDs | skill_eval_raw_cases, load_skill_owned_cases | Loom authoring contract. | +| Target kind/name and selector aliases | case_target_kind, case_target_name, case_selectors | Loom case selection UX. | +| Global duplicate case-ID rejection | main | Corpus validation belongs with project case normalization. | + +The central Loom suites are mostly consumed as authored. Skill-owned evals receive stronger runtime normalization and support several historical JSON shapes. + +#### Target and workspace preparation + +The following are Loom/profile behavior: + +- safe placement of authored fixture_files inside the isolated project; +- creation of target and judge OpenCode projects; +- copying the Loom skills tree; +- copying/promoting Loom agent files; +- wrapping role-decision and conversation-response agents with tool-denied evaluation boundaries; +- creating the eval-judge agent from Loom's JUDGE_AGENT prompt; +- choosing read-only versus read-write workspace mode from Loom execution mode; +- materializing the full Loom agent set for runtime cases so Loom subagent roles resolve; +- supplying the Loom repository as config-root and expecting the Loom plugin for runtime cases; +- mounting Loom node_modules for runtime cases; +- target prompt differences between runtime/conversation and role-decision cases; +- Copilot-specific insertion of Loom agent or skill guidance into system text. + +The generic orchestration layer may need a prepared invocation specification, but it must not know how Loom agents, skills, fixtures, or plugin trees are built. + +#### Judge construction and semantic interpretation + +All of these are Loom-owned semantics: + +- JUDGE_AGENT content and its execution-mode rules; +- positive expectations, forbidden behavior, and named trap meaning; +- judge_prompt layout and which target observations are shown to the judge; +- parse_judge strict JSON shape; +- judge_contract_error requirement that every expectation and violation is represented exactly once; +- semantic_pass rules; +- semantic_behavior_score; +- trap handling; +- skill-value labels and the 10 percentage-point material-improvement threshold. + +The engine can sequence a judge invocation, but it cannot own this rubric without turning Loom policy into a universal evaluation language. + +#### Deterministic assertions + +The entire deterministic assertion language in scripts/run-evals.py is Loom/project behavior: + +- tool requires/forbids; +- action requires/forbids/any_of; +- argument alias normalization; +- equals, ends_with, contains, contains_all matching; +- output contains/forbids; +- minimum source URL checks; +- native skill-load confirmation; +- tool-result requires/forbids; +- occurrence and after ordering; +- Loom-specific normalization of tool names and Code Mode call shapes. + +Key functions include normalize_tool, normalized_target_actions, action_matches, tool_result_matches, tool_result_call_state, result_evidence_events, describe_* helpers, and deterministic_failures. + +This is evidence-consuming project logic, not a candidate universal assertion DSL. + +#### Skill ablation + +run_skill_ablation_case is a Loom/project evaluation profile, not generic core behavior. + +It owns: + +- baseline versus candidate workspace construction; +- removal of the target skill from baseline; +- exposing only the target skill to the candidate; +- reasoning-only/no-persistence boundaries; +- transport-specific skill injection; +- baseline contamination detection; +- four invocations per iteration; +- baseline/candidate semantic scoring; +- delta_pp; +- trap_fixed/trap_regression; +- material-improvement/improvement/neutral/regression; +- candidate absolute PASS requirement. + +A generic engine may execute whatever phases a profile asks for, but this benchmark meaning belongs to Loom. + +#### Loom reporting and selection UX + +The following are project/CLI concerns rather than generic engine semantics: + +- --all, --cases, --suite, --target-kind, --target selection rules; +- the inference-spend safety rule requiring an explicit selection; +- native skill-routing restrictions for non-OpenCode target transport; +- Loom-oriented labels such as agent/execution and skill-ablation output; +- printing deterministic failure details and Loom action previews. + +Issue #58 explicitly excludes adding a public eval CLI now, so these current CLI choices are evidence of behavior, not an interface to copy. + +### 3. Low-level invoke responsibilities + +The post-#45 runner already owns the single-invocation execution boundary. These current Loom concerns must not be reimplemented by the generic eval engine. + +| Current Loom concern | Current functions | Post-#45 owner | +| --- | --- | --- | +| Resolve Podman/Docker | resolve_engine | opencode-eval-runner invoke | +| Select/override transport image | image_for_transport and DEFAULT_IMAGES | invoke configuration | +| Bind workspace and additional mounts | volume, prepare_node_modules_mount, invoke_container | invoke | +| Pass network mode | invoke_container | invoke | +| Forward explicit environment variables/provider credentials | host_environment_for_transport, pass_env | invoke | +| Seed auth/config/models/OpenCode DB/config-root | default_* helpers, sanitize_database_seed, resolve_optional_file | invoke | +| Enforce workspace ro/rw mount mode | invoke_container | invoke | +| Set model, agent, skill, reasoning, expected plugin | invoke_container | invoke | +| Enforce inner model timeout and outer container timeout | invoke_container | invoke | +| Execute exactly one isolated OCI transport invocation | invoke_container's runner path | invoke | +| Validate official result shape before persistence | historical Loom checks | host runner validates result/v1 and runtime_evidence/v1 | +| Build/validate authoritative runtime evidence | historical Loom observers/projections | runner runtime_evidence/v1 builder/validator | + +The orchestration layer chooses invocation inputs and time budgets. The low-level runner performs one invocation and returns one validated result. + +### 4. Legacy/compatibility behavior that should not migrate + +The largest non-portable section of run-evals.py is historical evidence/transport compatibility code. Post-#45 contracts supersede it. + +#### Direct OCI fallback + +invoke_container first prefers an opencode-eval-runner executable. If it is absent, Loom constructs and runs the transport container directly. + +That fallback duplicates runner responsibilities: + +- OCI hardening flags; +- workspace/input/seed mounts; +- provider environment forwarding; +- timeout handling; +- raw stdout JSON parsing; +- result projection. + +The future generic orchestration layer should not preserve a second execution implementation beside invoke. + +#### Candidate runner-safety mode and private-policy machinery + +The CredentialEntry/CredentialInventory block and associated RUNNER_SAFETY_* constants/functions implement an older candidate safety adapter. This includes: + +- immutable candidate image/hash checks; +- private credential inventory construction; +- policy-file generation; +- safe-result admission; +- evidence_safety acknowledgement/validation; +- custom safe-result schemas; +- disposable runtime-state checks; +- preflight result synthesis. + +Post-#45 runtime evidence moves authoritative observation safety and validation into the runner. This older Loom-side admission system is compatibility history, not generic eval-engine behavior. + +#### Loom-side tool-result reconstruction + +The following historical paths must not become a second evidence authority: + +- extracting tool results from raw target stdout; +- opencode-eval-runner/tool-results/v1 projection consumption; +- observed_tool_results; +- nested Code Mode call reconstruction; +- metadata alignment; +- truncation and omission accounting maintained by Loom; +- result matching against that reconstructed ledger. + +Post-#45 explicitly says tools, actions, tool_result_evidence, stdout/stderr, model text, and workspace files are diagnostic/convenience data and cannot fill an authoritative runtime-evidence gap. + +Project assertions may still consume non-authoritative fields for non-runtime/product semantics where appropriate, but they cannot treat those fields as substitutes for runtime_evidence/v1. + +#### In-process Loom observer plugin + +setup_projects can install eval-tool-observer.ts and attach_observer_capture can replace the projected tool-result ledger with data from .loom-eval-tool-observer.jsonl. + +That is historical Loom-side evidence collection. The post-#45 runner owns authoritative native and Code Mode observation and publishes only runtime_evidence/v1 as the authority. + +#### Normal-invoke diagnostic observer + +OPENCODE_EVAL_NORMAL_OBSERVATIONS enables loom-normal-invoke-observer.ts and stores normal_invoke_diagnostic. + +The code already marks this channel diagnostic_only and deliberately excludes it from scoring/evidence eligibility. It can remain a debugging aid in Loom if useful, but it is not an extraction input for the generic engine. + +#### Loom evidence_safety field disposition + +prepare_transport_result creates Loom's own evidence_safety map with field state exact/redacted/omitted and target_scoring_evidence_error requires exact values for text/tools/actions/skills_loaded. + +That is not the post-#45 authoritative evidence model. The public runtime-evidence field states are available/redacted/omitted/unsupported, and runtime readiness is decided per required runtime boundary/field. + +Do not migrate the Loom-side exact-field status engine as a competing evidence authority. + +## Detailed function-family map + +This map covers the major function blocks in the 5,215-line script. + +| Lines | Functions / behavior | Ownership | +| --- | --- | --- | +| 24-43 | strict_json_loads duplicate/non-finite rejection | Legacy helper used by old runner-safety admission; official runner now validates its result. | +| 53-219 | JUDGE_AGENT, agent wrappers, skill baseline/candidate wrappers | Project/profile | +| 222-354 | central and skill-owned case discovery/normalization | Project/profile | +| 357-404 | workspace mode, target identity, reasoning provenance, artifact default, selectors | Mixed: project/profile for target/workspace/selectors; run provenance is generic metadata. | +| 407-586 | safe fixtures, project config, observer attachment, target/judge workspace setup, ablation setup | Mostly project/profile; observer attachment is legacy. | +| 589-721 | engine resolution, default seed paths, DB sanitization, volumes, transport env | Low-level invoke duplication | +| 724-1675 | credential inventory, runner-safety admission, evidence projection/sanitization | Legacy/compatibility — do not migrate | +| 1676-2531 | redaction helpers, stdout/tool-result extraction, observer ledger, nested call reconstruction, prepare_transport_result | Legacy/compatibility — do not migrate as evidence authority | +| 2534-2601 | reasoning enforcement, node_modules mount prep, timeout/image helpers | Mixed: low-level invoke plus Loom profile defaults | +| 2602-3030 | invoke_container | Mixed: desired runner call is low-level invoke; direct-container fallback and Loom evidence projection are legacy | +| 3032-3449 | tool/action/result normalization and assertion matching helpers | Project/profile | +| 3451-3719 | deterministic_failures | Project/profile | +| 3722-3830 | judge prompt, parse/contract validation, semantic pass | Project/profile | +| 3833-3846 | classify_behavioral_result | Mixed: generic tri-state precedence fed by project semantic outcomes | +| 3849-3885 | skill score/value and target prompt | Project/profile | +| 3888-3994 | transport error extraction, retry predicate, old exact-field evidence gates | Mixed: orchestration error handling plus project/legacy policy | +| 3995-4032 | invoke_container_with_retry | Generic orchestration loop; hard-coded retry predicate is current policy | +| 4035-4210 | artifact naming, run claim, hash, atomic write, verification | Generic artifact plumbing with Loom-specific paths/names | +| 4213-4560 | run_skill_ablation_case | Project/profile workflow using generic phase sequencing | +| 4563-4814 | run_case | Mixed: generic target→judge→classification→artifact skeleton with Loom hooks/semantics | +| 4817-4858 | eval_job_concurrency | Generic scheduling mechanism plus Loom runtime/non-runtime policy | +| 4861-5210 | main selection, matrix, execution groups, reporting, exit | Mixed: generic run control plus Loom CLI/profile/reporting | + +## Trace details required by issue #58 + +### Case loading and normalization + +Central behavioral suites: + +- default discovery is direct JSON files under evals/; +- a suite or case can set default: false to be opt-in; +- explicitly selecting case IDs causes opt-in suites/cases to be included; +- central cases are mostly used in authored shape. + +Skill-owned suites: + +- every direct JSON file under skills//evals/ is discovered; +- accepted historical forms are top-level arrays, {skill,cases}, and {skill_name,evals}; +- the declared skill must match the containing folder; +- prompt and expectations are mandatory; +- negative_expectations or must_not becomes must_not; +- trap is normalized to a string and remembered as explicitly declared or not; +- the normalized ID is SKILL--; +- normalized skill-owned cases always use the skill-eval agent and runtime execution; +- duplicate normalized IDs are rejected. + +main later rejects duplicate IDs across the combined central and skill-owned corpus. + +### Target setup + +Normal cases use a fresh temporary root with separate target and judge projects. + +For Loom runtime cases, the target contains: + +- all Loom skills; +- the full Loom agent set; +- the selected target agent promoted to primary; +- authored fixture files; +- the Loom runtime plugin supplied from the repository config root; +- node_modules when runtime execution requires it; +- a writable workspace because Loom stores project-local runtime state. + +Role-decision and conversation-response cases use wrapped, tool-denied primary agents and a read-only workspace. + +The judge gets a separate project containing only the eval-judge agent. + +This isolation pattern is useful evidence for the generic lifecycle, but the contents and mutation policy are Loom profile decisions. + +### Target invocation + +run_case resolves target/judge models, reasoning provenance, images, timeouts, and seed files, then calls invoke_container_with_retry. + +For a Loom runtime target: + +- target transport must be OpenCode; +- config_root is the Loom repository; +- expected_plugin is loom; +- workspace is writable; +- node_modules are mounted; +- case-specific target timeout can override the global timeout. + +Judge invocations use a separate judge workspace and no Loom runtime plugin. + +Skill ablation uses separate baseline, candidate, and judge projects. + +### Retry behavior + +invoke_container_with_retry is the only retry loop in the Loom harness. + +Each attempt calls invoke_container once. + +The current predicate retries only when: + +- the invocation has an error; +- the error contains provider.no-route and Model unavailable; +- no tool/action has been observed; +- attempts remain. + +A successful retry returns transport_retry metadata with attempt count and prior errors. A terminal failed attempt after earlier retries records all errors. + +This is consistent with the issue invariant that retries belong to orchestration, not invoke. + +### Runtime evidence and legacy observers + +The current Loom script predates the final post-#45 evidence consumer shape. + +Its scoring path still uses: + +- Loom-side evidence_safety; +- observed_tool_results; +- tool_result_evidence/stdout reconstruction; +- optional eval-tool-observer.ts output; +- optional diagnostic normal-invoke observer output. + +Post-#45 changes the authority model: + +- runtime_evidence with schema opencode-eval-runner/runtime-evidence/v1 is mandatory; +- it is the only authoritative runtime-evidence object; +- overall status is complete/incomplete/unsupported/invalid; +- readiness is assertion-scoped by required boundary and required field availability; +- native and code_mode_execution may be complete while code_mode_finality is unsupported; +- tools/actions/tool_result_evidence/stdout/stderr/model text/workspace files are not substitutes; +- the host invoke CLI validates runtime_evidence before writing the result; +- Copilot reports runtime evidence as unsupported. + +Therefore the Loom observer/evidence stack is reference history for behavior and assertions, not code to extract into the generic engine. + +### Judge construction and parsing + +run_case invokes the judge whenever the target is usable, even if deterministic assertions have already found behavioral failures. + +The judge sees: + +- case identity, target, and execution mode; +- trap; +- scenario context; +- every positive expectation; +- every forbidden behavior; +- observed tools/actions; +- runtime tool results for normal Loom runtime cases; +- final assistant text. + +parse_judge accepts JSON, including JSON wrapped in common markdown fences, then requires: + +- boolean passed; +- expectations array; +- violations array; +- boolean trap_observed; +- string trap_evidence. + +judge_contract_error then requires the expectation/violation arrays to exactly match the number of authored rules and requires boolean met/violated values per entry. + +These are Loom's judge contract, not a generic judge schema. + +### Deterministic checks + +deterministic_failures evaluates authored Loom assertions separately from semantic judging. + +It distinguishes an observed behavior failure from inability to prove an assertion. Several evidence gaps are represented as strings prefixed with non-evidence:. + +Current examples include: + +- incomplete capture for required/forbidden tools; +- incomplete capture for required/forbidden actions; +- missing complete tool-result evidence; +- unavailable chronological ordering; +- malformed call identity/input; +- unavailable or indeterminate per-call result; +- missing nested Code Mode capture. + +This distinction is valuable, but the exact assertions and evidence mapping are Loom semantics. + +### Classification + +Normal cases currently classify as: + +- non-evidence when target_error or judge_error exists; +- non-evidence when every deterministic failure is a non-evidence failure; +- behavioral-fail when any deterministic behavioral failure exists, including a mix of behavioral and non-evidence failures; +- pass when deterministic checks are clear and semantic_pass is true; +- behavioral-fail otherwise. + +The issue #58 target vocabulary is pass | fail | non-evidence. This inventory records the current Loom name behavioral-fail without choosing the final engine representation. + +Skill ablation has related but distinct project logic: + +- artifact starts as non-evidence; +- target/judge evidence errors keep the run non-evidence; +- candidate absolute pass requires no candidate deterministic failures plus semantic_pass; +- candidate pass produces pass; otherwise behavioral-fail; +- skill-value classification is separate from absolute pass. + +### Artifacts + +For normal cases, the artifact records: + +- case and iteration; +- target identity/profile data; +- target/judge images and transports; +- model and reasoning provenance; +- target/judge/total timing; +- classification and passed boolean; +- full target result; +- target_error; +- normalized observed actions; +- observed_tool_results; +- deterministic failures; +- judge transport result; +- semantic judge object; +- judge_error. + +Skill ablation adds baseline/candidate phase records, scores, delta, trap change, and skill-value classification. + +Run-level artifact behavior: + +- a UUID eval_run_id is generated; +- default directory is .loom-evals/; +- .loom-eval-run-id claims the directory; +- a non-empty directory or conflicting owner fails closed; +- per-case files are written through a temporary file plus os.replace; +- artifact_evidence_id is SHA-256 over the serialized artifact excluding the ID itself; +- report re-reads and re-hashes the durable artifact before counting it. + +Current limitation to preserve as an observation, not a design decision: exceptions caught by main.execute before run_case writes an artifact are reported to the console but do not necessarily produce a per-job artifact. + +### Iterations and concurrency + +main creates one job for every selected case and every iteration from 1..N. + +Scheduling is split: + +- non-runtime jobs use --parallel; +- --parallel with no explicit number maps to the full non-runtime job count; +- runtime jobs use the separate --runtime-parallel cap; +- runtime default is one; +- non-runtime group executes before runtime group; +- parallel result reporting occurs in completion order. + +The Loom reason for separate runtime scheduling is product-specific: one runtime case may already dispatch many subagents. + +### Skill ablation + +Each skill-owned case performs four model invocations per iteration: + +1. baseline target; +2. baseline judge; +3. candidate target; +4. candidate judge. + +The same target reasoning is used for baseline and candidate. The benchmark separately records absolute candidate correctness and relative skill value. + +This is an important project extension use case for the final architecture: the generic engine must not assume every case is exactly one target plus one judge, but this inventory does not choose the extension API. + +### Reporting + +report first verifies durable artifact integrity. + +Console states are: + +- PASS for classification pass; +- ERROR for non-evidence; +- FAIL for behavioral-fail. + +It also reports: + +- total duration; +- artifact evidence ID prefix; +- ablation baseline/candidate delta; +- trap fixed/regression; +- deterministic failure details; +- observed tools/actions; +- semantic judge summary when false. + +main exits zero only if every job counted as pass. + +## Validation against post-PR-#45 runner contracts + +The following cross-checks were made against the merged runner documentation and host CLI behavior. + +### One invoke is exactly one isolated invocation + +Confirmed. + +The host CLI describes invoke as one isolated target or judge invocation and runner/cli.py executes one OCI subprocess for one invoke call. There is no runner retry loop. + +Loom's retries are outside invoke in invoke_container_with_retry, which preserves the invariant. + +### Runner does not own eval semantics + +Confirmed. + +docs/invocation-usage.md states that the runner does not own: + +- cases/suites; +- assertions; +- judge semantics; +- thresholds; +- target-versus-judge orchestration; +- iterations; +- aggregation; +- PASS/FAIL/non-evidence decisions. + +Those match the project/generic split observed in Loom. + +### runtime_evidence/v1 is authoritative + +Confirmed. + +docs/runtime-evidence-contract.md states: + +- runtime_evidence is the only authoritative runtime-evidence object; +- raw observers are internal adapter input; +- tools, actions, tool_result_evidence, stdout/stderr, model text, and workspace files are diagnostic/convenience only; +- incomplete or invalid overall capture makes runtime assertions ineligible; +- each assertion declares required runtime boundaries; +- each required boundary must be complete; +- exact required observation fields must be available. + +This directly supersedes the Loom-side historical evidence eligibility machinery. + +### Boundary support is partial, not all-or-nothing + +Confirmed. + +The post-#45 contract exposes: + +- native; +- code_mode_execution; +- code_mode_finality. + +native and code_mode_execution can be complete and eligible while code_mode_finality is unsupported. The generic evidence-readiness stage therefore cannot use only a single top-level complete flag, while project code still decides which boundaries/fields an assertion needs. + +### Result contract is validated by invoke + +Confirmed. + +Each invocation produces one opencode-eval-runner/v1 result with mandatory runtime_evidence. The host validates runtime_evidence before persisting the output file. + +The eval engine should consume that validated result rather than rebuild transport evidence. + +### Unsupported areas remain explicit + +Confirmed. + +Post-#45 explicitly reports: + +- exact Code Mode caller-final value/error as unsupported on stock OpenCode 2.0.23; +- runtime evidence as unsupported for github-copilot-cli; +- no hostile-plugin isolation guarantee for the trusted-checkout profile. + +The project must not fill those gaps from model text or diagnostic result fields. + +## Extraction facts for the integration owner + +These are inventory conclusions, not final module/API design: + +1. The proven reusable lifecycle is target execution → evidence readiness → project checks/judge → classification → durable artifact. +2. Runner invoke is below that lifecycle and remains one invocation per call. +3. Retry is above invoke and must remain visible as multiple attempts. +4. Case discovery, normalization, fixtures, workspace content, prompts, assertions, judge parsing, semantic pass, and ablation meaning are project/profile responsibilities. +5. The generic engine needs to carry project-produced metadata without interpreting it. +6. Iteration and bounded concurrency are generic run mechanics; Loom's runtime/non-runtime scheduling distinction is profile policy. +7. Artifact identity, atomic persistence, attempt/result provenance, and timing are reusable orchestration behavior. +8. The current Loom evidence stack is not an implementation source for the new authoritative evidence boundary. The post-#45 runtime_evidence/v1 contract is. +9. Diagnostic tools/actions/tool_result_evidence may still be useful project inputs, but never as replacements for missing authoritative runtime evidence. +10. Skill ablation proves the project extension surface must support more than one target/judge phase, but its benchmark rules remain Loom-owned. +11. No universal assertion DSL is implied by Loom's deterministic_failures function. +12. No Loom code migration or public eval CLI is performed by this worker. + +## Non-goals of this document + +This document does not: + +- define final engine modules or APIs; +- choose the normalized engine case schema; +- choose the final extension interface; +- migrate Loom cases or assertions; +- change runtime code; +- change invoke; +- add retries to invoke; +- add a public eval CLI; +- define a universal assertion DSL. + +Those decisions belong to the issue #58 integration synthesis after both worker inventories are available. From b7cddc3df30c73efb1da2655f4a886ae120a63d6 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Mats=20B=C3=B8e=20Bergmann?= Date: Wed, 7 Oct 2026 10:23:55 +0200 Subject: [PATCH 3/3] docs: define generic eval-engine extraction architecture --- docs/eval-engine-architecture.md | 976 +++++++++++++++++++++++++++++++ 1 file changed, 976 insertions(+) create mode 100644 docs/eval-engine-architecture.md diff --git a/docs/eval-engine-architecture.md b/docs/eval-engine-architecture.md new file mode 100644 index 0000000..f3f3e6c --- /dev/null +++ b/docs/eval-engine-architecture.md @@ -0,0 +1,976 @@ +# Generic eval-engine extraction architecture + +Issue: #58 — Task 1 of 9 +Status: architecture/extraction contract; no eval-engine implementation in this task. + +## 1. Purpose + +This document defines the internal boundary for extracting the proven generic orchestration mechanics from Loom's `scripts/run-evals.py` into `opencode-eval-runner`. + +It is constrained by two inventories merged into this branch: + +- [Loom eval runner responsibility inventory](loom-eval-runner-responsibility-inventory.md) +- [Runner contract and invariant matrix](eval-engine-runner-contracts.md) + +The design deliberately does **not** turn this repository into the owner of project-specific eval semantics. It provides reusable orchestration over the existing one-invocation runner boundary. + +The target ownership split is: + +```text +project/profile + cases / fixtures / prompts / assertions / judge semantics / thresholds + | + v +generic eval engine + selection / iterations / concurrency / attempts / readiness + target -> judge -> classification / artifacts / summaries + | + v +invoke + exactly one isolated invocation + | + v +opencode-eval-runner/v1 + + authoritative runtime_evidence/v1 +``` + +## 2. Non-negotiable invariants + +1. **One `invoke` means exactly one isolated invocation.** +2. **No hidden retry loop may be added to `invoke`.** +3. Retry is orchestration policy and every attempt must be visible in artifacts. +4. `opencode-eval-runner/v1` remains the low-level invocation result envelope. +5. `opencode-eval-runner/runtime-evidence/v1` is the only authoritative runtime-evidence object. +6. Runtime evidence readiness is scoped to the facts a project assertion requires. +7. Diagnostic/convenience fields cannot repair missing authoritative runtime evidence. +8. Product outcome, evidence outcome, and infrastructure outcome remain distinct. +9. Project/profile code owns what correct behavior means. +10. No universal assertion DSL is introduced. +11. No input-JSON invocation API is introduced. +12. No public `eval` CLI is introduced by Task 1. +13. Copilot runtime observation remains explicitly unsupported. +14. Code Mode exact caller-final value/error remains explicitly unsupported on stock OpenCode 2.0.23. +15. The normal trust profile remains trusted-checkout, not hostile-plugin isolation. +16. Loom's old direct-container fallback and old observer/evidence reconstruction paths are not migrated. + +## 3. Responsibility boundary + +### 3.1 Generic eval-engine responsibility + +The generic engine owns mechanics that do not depend on Loom's domain: + +- normalized case/run identity; +- selectors over already-normalized cases; +- explicit spend/selection safeguards; +- case x iteration job expansion; +- stable case/iteration labels; +- general and runtime execution lanes; +- bounded concurrency planning; +- target attempt orchestration; +- explicit retry policy and attempt accounting; +- runtime-evidence readiness mechanics; +- judge attempt orchestration; +- generic tri-state classification precedence; +- run/case timing; +- durable run/case artifacts; +- artifact integrity verification; +- run summary and process-level success/failure; +- preservation of transport/model/reasoning/image provenance. + +### 3.2 Project/profile responsibility + +A project integration such as Loom owns: + +- case discovery and source-file compatibility; +- case normalization into the generic case envelope; +- selector aliases beyond the canonical case ID; +- whether a case belongs to the standard or runtime concurrency lane; +- fixture construction; +- workspace contents and writable/read-only policy; +- target agent/skill/profile selection; +- target prompt construction; +- invocation-specific config/mount/network choices; +- which runtime-evidence boundaries a case/assertion requires; +- locating the observations/fields relevant to a project assertion; +- deterministic behavioral assertions; +- judge prompt/rubric construction; +- judge result parsing and validation; +- semantic pass/fail meaning; +- project-specific scores, traps, thresholds, metadata, and reporting details; +- skill-ablation meaning and comparison policy. + +### 3.3 Low-level `invoke` responsibility + +The existing invocation boundary continues to own: + +- OCI process construction and isolation; +- workspace/input/config/credential mounting; +- transport selection; +- provider authentication forwarding; +- inner invocation timeout; +- host/container timeout behavior; +- raw transport execution; +- result construction; +- result JSON parsing/validation; +- evidence sanitization before runner-controlled sinks; +- authoritative OpenCode runtime observation; +- `runtime_evidence/v1` construction and validation; +- output-file persistence for one invocation. + +The eval engine must reuse this implementation path rather than create another container runner. + +### 3.4 Legacy/compatibility behavior that must not migrate + +The following Loom-side mechanisms are explicitly excluded from the generic engine: + +- direct OCI/container execution fallback when the runner binary is unavailable; +- Loom-side `evidence_safety` eligibility as a competing authority; +- `observed_tool_results` as runtime evidence authority; +- reconstruction from `tool_result_evidence`, stdout/stderr, or model output; +- nested Code Mode reconstruction from diagnostic surfaces; +- `.loom-eval-tool-observer.jsonl` as scoring authority; +- the candidate runner-safety/private-policy adapter; +- disposable-state/candidate-image admission machinery from the superseded trust direction; +- old exact/redacted/omitted Loom evidence vocabulary as a competing runtime contract. + +Diagnostic compatibility output may remain in Loom during migration, but it cannot authorize PASS. + +## 4. Internal package/module decomposition + +The following decomposition is authoritative for Tasks 2-5 unless implementation detail forces a narrow rename. Responsibilities should not migrate between these modules without updating this architecture. + +```text +runner/ + eval_types.py shared internal dataclasses/enums/protocol values + eval_plan.py selection, iteration expansion, concurrency plan + eval_artifacts.py run/case persistence, schemas, integrity + eval_evidence.py canonical evidence-readiness mechanics + eval_execute.py target/judge attempt execution and retry plumbing + eval_judge.py judge lifecycle + generic classification + eval_engine.py composition/orchestration only +``` + +Task mapping: + +- **Task 2 / #59:** `eval_types.py` planning subset + `eval_plan.py` +- **Task 3 / #60:** artifact subset of `eval_types.py` + `eval_artifacts.py` +- **Task 4 / #61:** evidence/execution subset + `eval_evidence.py` + `eval_execute.py` +- **Task 5 / #62:** judge/classification subset + `eval_judge.py` + minimal `eval_engine.py` composition + +No task should need to redesign the ownership boundary to proceed. + +## 5. Core internal contracts + +These are internal Python contracts, not a new public wire protocol. + +### 5.1 JSON-compatible project metadata + +Generic artifacts may carry project-owned metadata, but the generic engine must not interpret it. + +```python +JsonValue = None | bool | int | float | str | list["JsonValue"] | dict[str, "JsonValue"] +``` + +Values written to durable artifacts must be strict JSON-compatible values. + +### 5.2 NormalizedCase + +```python +@dataclass(frozen=True) +class NormalizedCase: + id: str + selectors: tuple[str, ...] + lane: Literal["standard", "runtime"] + project_data: JsonValue + metadata: dict[str, JsonValue] +``` + +Rules: + +- `id` is globally unique within one normalized suite. +- `selectors` may contain project-defined aliases but always includes `id`. +- `lane` is scheduling metadata only. It does not imply behavioral meaning. +- `project_data` is opaque to the generic engine. +- the generic engine must not require Loom fields such as `agent`, `execution`, `requirements`, `trap`, or `expectations`. + +### 5.3 EvalJob + +```python +@dataclass(frozen=True) +class EvalJob: + case: NormalizedCase + iteration: int + label: str +``` + +`label` is stable for the same case/iteration plan: + +- one iteration: `` +- multiple iterations: `#` + +### 5.4 RunPlan + +```python +@dataclass(frozen=True) +class RunPlan: + run_id: str + jobs: tuple[EvalJob, ...] + standard_parallelism: int + runtime_parallelism: int +``` + +Task 2 may add normalized selector/planning metadata, but execution semantics remain: + +- case x iteration expansion is deterministic; +- standard and runtime jobs are planned separately; +- runtime parallelism is independently bounded; +- current Loom behavior of running the standard group before the runtime group is the initial generic ordering because it avoids mixing normal throughput with runtime stress; +- a future public interface may later expose another scheduling policy, but Tasks 2-5 should not invent one. + +### 5.5 InvocationSpec + +```python +@dataclass(frozen=True) +class InvocationSpec: + transport: Literal["opencode", "github-copilot-cli"] + model: str + reasoning: str | None + agent: str | None + skill: str | None + workspace: Path + workspace_mode: Literal["ro", "rw"] + prompt: str + system: str | None + expected_plugin: str | None + engine: str + network: str | None + image: str | None + auth: Path | None + database: Path | None + models_catalog: Path | None + config: Path | None + config_root: Path | None + env_names: tuple[str, ...] + timeout_seconds: int + container_timeout: int +``` + +This is a logical internal orchestration structure. It is **not** an input JSON API. + +The default invoker adapter must execute through the same host `invoke` implementation/semantics used by the public CLI. It may factor common Python code out of `runner.cli.invoke()`, but it must not duplicate OCI construction or invoke the transport container directly. + +### 5.6 AttemptFailure + +```python +@dataclass(frozen=True) +class AttemptFailure: + plane: Literal["infrastructure", "product", "evidence"] + code: str + message: str + retry_safe: bool +``` + +`retry_safe` is a fact established by the attempt classifier, not permission to retry. Policy still decides whether another attempt is allowed. + +Examples of infrastructure codes may include: + +- outer container timeout; +- missing result file; +- invalid/non-object result JSON; +- failed workspace/config/plugin preflight. + +Product failure remains distinct, for example: + +- transport/model non-zero result; +- inner model timeout represented by `result.exit_code == 124`. + +Evidence failure is scoped to a declared evidence requirement, not inferred from product status. + +### 5.7 AttemptRecord + +```python +@dataclass(frozen=True) +class AttemptRecord: + attempt: int + started_at: str + duration_seconds: float + host_exit_code: int | None + result: dict[str, JsonValue] | None + failure: AttemptFailure | None +``` + +Every actual call to `invoke` produces one AttemptRecord, including failed attempts where no valid result exists. + +### 5.8 RetryPolicy / RetryDecision + +Retry remains orchestration policy: + +```python +class RetryPolicy(Protocol): + def decide( + self, + attempts: tuple[AttemptRecord, ...], + latest: AttemptRecord, + ) -> "RetryDecision": ... + +@dataclass(frozen=True) +class RetryDecision: + retry: bool + reason: str + delay_seconds: float +``` + +Requirements: + +- default internal behavior is no hidden retry unless a policy is explicitly supplied; +- max attempts/backoff are bounded; +- every prior attempt is retained; +- a behavioral FAIL is never retried as infrastructure; +- replay safety must fail closed; +- convenience `tools`/`actions` must not be used as authoritative proof that no side effect occurred; +- if retry safety depends on whether execution reached a runtime tool boundary, use authoritative `runtime_evidence` when available; +- if safety cannot be established, do not retry. + +The narrow Loom `provider.no-route` + `Model unavailable` behavior is reference policy, not a hard-coded engine semantic. + +### 5.9 EvidenceRequirement + +The engine needs a generic way to decide capture/boundary readiness without owning behavioral assertions. + +```python +@dataclass(frozen=True) +class EvidenceRequirement: + boundaries: tuple[ + Literal["native", "code_mode_execution", "code_mode_finality"], + ... + ] +``` + +A profile chooses the required boundary set for a case or individual project check. + +The generic engine does **not** define a selector/assertion DSL for observations. Project checks locate the observation and exact field they care about. + +### 5.10 EvidenceReadiness + +```python +@dataclass(frozen=True) +class EvidenceReadiness: + status: Literal["ready", "incomplete", "unsupported", "invalid"] + reasons: tuple[str, ...] + required_boundaries: tuple[str, ...] +``` + +`eval_evidence.py` exposes two levels of reusable mechanics: + +1. **capture/boundary readiness** + +```python +check_evidence_readiness( + runtime_evidence: Mapping[str, Any], + requirement: EvidenceRequirement, +) -> EvidenceReadiness +``` + +2. **exact field readiness** + +```python +check_field_readiness(field: Mapping[str, Any]) -> EvidenceReadiness +``` + +Rules: + +- valid `runtime_evidence/v1` is mandatory; +- overall `incomplete` or `invalid` makes any non-empty runtime requirement not ready; +- every required boundary must be `complete`; +- a required boundary that is `unsupported` yields `unsupported`; +- exact field state `available` is ready; +- exact field `redacted` or `omitted` yields incomplete for that exact-value project check; +- exact field `unsupported` yields unsupported; +- an empty boundary requirement is allowed for cases/checks that do not depend on runtime evidence; +- diagnostics never backfill an unavailable authoritative field. + +This separation avoids a universal assertion DSL: the generic engine decides whether declared evidence is usable, while project code decides what fact to ask for and what it means. + +### 5.11 CheckOutcome + +Project deterministic checks normalize into: + +```python +@dataclass(frozen=True) +class CheckOutcome: + name: str + status: Literal["pass", "fail", "non-evidence"] + reason: str + metadata: dict[str, JsonValue] +``` + +The generic engine does not know how a check was computed. + +A profile may use authoritative runtime evidence, product text, repository data, or other project evidence as appropriate. It must not use diagnostic fields to substitute for missing authoritative runtime facts. + +### 5.12 SemanticDecision + +```python +@dataclass(frozen=True) +class SemanticDecision: + status: Literal["pass", "fail"] + summary: str + data: JsonValue +``` + +Judge parse/contract failure is not represented as a semantic FAIL. It is a judge attempt/infrastructure failure and therefore non-evidence. + +### 5.13 EvaluationResult + +```python +@dataclass(frozen=True) +class EvaluationResult: + classification: Literal["pass", "fail", "non-evidence"] + target_attempts: tuple[AttemptRecord, ...] + target_readiness: EvidenceReadiness + deterministic_checks: tuple[CheckOutcome, ...] + judge_attempts: tuple[AttemptRecord, ...] + semantic: SemanticDecision | None + project_metadata: dict[str, JsonValue] +``` + +The durable artifact adds case/run identity and timing. + +## 6. Project/profile extension protocol + +The project extension surface is intentionally small and imperative. It is not a declarative assertion language. + +Conceptually: + +```python +class EvalProfile(Protocol): + def discover_cases(self) -> Sequence[NormalizedCase]: ... + + @contextmanager + def prepare( + self, + case: NormalizedCase, + iteration: int, + ) -> Iterator[Any]: + ... + + def target_spec( + self, + case: NormalizedCase, + prepared: Any, + ) -> InvocationSpec: + ... + + def target_evidence_requirement( + self, + case: NormalizedCase, + prepared: Any, + ) -> EvidenceRequirement: + ... + + def deterministic_checks( + self, + case: NormalizedCase, + prepared: Any, + target: AttemptRecord, + readiness: EvidenceReadiness, + ) -> Sequence[CheckOutcome]: + ... + + def judge_spec( + self, + case: NormalizedCase, + prepared: Any, + target: AttemptRecord, + checks: Sequence[CheckOutcome], + ) -> InvocationSpec | None: + ... + + def parse_judge( + self, + case: NormalizedCase, + prepared: Any, + judge: AttemptRecord, + ) -> SemanticDecision: + ... + + def artifact_metadata( + self, + case: NormalizedCase, + prepared: Any, + ) -> Mapping[str, JsonValue]: + ... +``` + +Narrow implementation changes are allowed, but the ownership rules are fixed: + +- `prepare` owns project workspace/fixture lifecycle; +- target/judge specs are project-produced; +- project code declares runtime evidence needs; +- deterministic checks and semantic meaning remain project-owned; +- generic code owns attempts, retries, readiness mechanics, classification precedence, scheduling, and artifacts. + +A profile may return `None` from `judge_spec` for deterministic-only cases. + +Skill ablation is **not** implemented in Tasks 2-5. The normal target/judge lifecycle should be factored so Task 8 can compose the same phase primitive twice rather than duplicate execution. + +## 7. One-case lifecycle + +The generic normal-case lifecycle is: + +```text +NormalizedCase + iteration + | + v +profile.prepare() + | + v +profile.target_spec() + | + v +execute target attempts + one invoke per attempt + | + v +profile.target_evidence_requirement() + | + v +generic evidence readiness + | + v +profile.deterministic_checks() + | + +-------------------------------+ + | | + v | +profile.judge_spec() | + | | + | None | InvocationSpec + | v + | execute judge attempts + | | + | v + | profile.parse_judge() + | | + +-------------------------------+ + | + v + generic classification + | + v + durable artifact +``` + +### 7.1 Target invocation usability + +A target is unusable/non-evidence when the engine cannot obtain a valid target result required for the case, for example: + +- outer OCI timeout; +- invalid result JSON; +- failed required plugin/runtime preflight; +- exhausted retryable infrastructure failure; +- product/model timeout or error when the profile requires a completed target response; +- declared runtime evidence requirement is not ready. + +The profile may still have deterministic-only cases that do not require runtime evidence. An empty `EvidenceRequirement` means runtime status such as Copilot `unsupported` does not by itself block a purely semantic/textual case. + +### 7.2 Judge execution + +If the target is usable enough to judge, the engine should not automatically skip the judge merely because deterministic checks already contain a behavioral failure. Loom currently retains both deterministic and semantic evidence, and the generic engine should preserve that diagnostic value. + +The profile can explicitly return no judge for deterministic-only cases. + +### 7.3 Classification precedence + +Generic final classification is: + +1. required target execution/infrastructure failure -> **non-evidence** +2. required target evidence readiness failure -> **non-evidence** +3. required judge execution or judge-contract failure -> **non-evidence** +4. any deterministic `fail` -> **fail** +5. otherwise any deterministic `non-evidence` -> **non-evidence** +6. semantic decision `fail` -> **fail** +7. semantic decision `pass` (or no judge required) and all required deterministic checks pass -> **pass** + +This preserves the important distinction: + +- inability to prove required behavior is not a behavioral failure; +- an observed valid behavioral violation is a real FAIL; +- judge parser/contract failure is not evidence against the product. + +If a project needs a different domain scoring model, it should normalize its domain-specific results into `CheckOutcome` and `SemanticDecision`; it should not replace the generic infrastructure/evidence precedence. + +## 8. Run planning and concurrency contract + +Task 2 implements planning over normalized cases. + +### 8.1 Selection + +The planner accepts: + +- normalized cases from the profile; +- canonical case IDs and profile-provided selectors; +- explicit selection/all intent; +- iteration count; +- standard parallelism; +- runtime parallelism. + +It must reject: + +- duplicate normalized IDs; +- unknown selectors; +- iteration < 1; +- invalid concurrency values; +- live execution with no explicit selection/all intent. + +Listing cases is provider-free and does not require model configuration. + +### 8.2 Job expansion + +For selected cases: + +```python +jobs = [ + EvalJob(case, iteration) + for case in selected_cases + for iteration in range(1, iterations + 1) +] +``` + +The planner never executes jobs. + +### 8.3 Concurrency lanes + +Profiles assign each case to `standard` or `runtime`. + +Initial generic scheduling preserves Loom's proven separation: + +1. execute the standard group with standard parallelism; +2. execute the runtime group with independently bounded runtime parallelism. + +The generic reason is resource/load separation. Loom's reason that a runtime case may itself dispatch multiple subagents remains project context, not engine semantics. + +The planner records resolved concurrency in run metadata so artifacts explain how the run was scheduled. + +## 9. Retry contract + +Retry is performed by `eval_execute.py`, never by `invoke`. + +For each phase: + +```text +attempt 1 -> classify attempt -> RetryPolicy + | no + v + final attempt + | + yes + v + delay/backoff + | + v +attempt 2 -> ... +``` + +Required behavior: + +- attempt numbering starts at 1; +- max attempts are bounded; +- each attempt calls `invoke` once; +- all attempts survive in the final artifact; +- retry reasons and delay are recorded; +- a success after retry does not erase earlier failures; +- no retry occurs after a valid behavioral result merely because the result would fail the eval; +- replay safety is conservative. + +Task 4 may provide a built-in narrow transient-provider policy based on the proven Loom policy, but it must remain explicitly selected/configured rather than hidden in `invoke`. + +## 10. Artifact contract + +Task 3 implements durable artifacts before Task 4/5 execution is wired in. + +Two versioned on-disk schemas are chosen from the start: + +- `opencode-eval-runner/eval-run/v1` — run manifest/summary +- `opencode-eval-runner/eval-artifact/v1` — one case/iteration result + +These are versioned internal/on-disk contracts during Tasks 2-5. They are not promised as stable external/public API until Task 6 reviews the public `eval` interface. + +### 10.1 Run manifest minimum fields + +```json +{ + "schema": "opencode-eval-runner/eval-run/v1", + "run_id": "...", + "created_at": "...", + "selection": {}, + "iterations": 1, + "concurrency": { + "standard": 1, + "runtime": 1 + }, + "jobs": [], + "summary": {} +} +``` + +### 10.2 Per-job artifact minimum fields + +```json +{ + "schema": "opencode-eval-runner/eval-artifact/v1", + "run_id": "...", + "case": "...", + "iteration": 1, + "lane": "standard", + "classification": "pass", + "timing": {}, + "target": { + "attempts": [], + "evidence_readiness": {} + }, + "deterministic_checks": [], + "judge": { + "attempts": [], + "semantic": null + }, + "project_metadata": {}, + "artifact_evidence_id": "..." +} +``` + +Rules: + +- full validated low-level target/judge results are preserved inside their AttemptRecords; +- prior retry attempts are retained; +- runtime evidence is stored exactly as returned; artifact code does not rewrite it; +- project metadata is opaque JSON; +- artifacts use atomic replace where supported; +- an artifact evidence ID is calculated over a canonical serialized form excluding the evidence ID itself; +- reporting/counting re-reads and verifies durable artifacts before trusting them; +- run directories fail closed on conflicting ownership/non-empty incompatible contents. + +Task 3 may refine field names but must preserve these semantics. + +## 11. Timeout model + +The engine must distinguish: + +### Inner invocation timeout + +Owned by `invoke --timeout-seconds`. + +A valid result may exist with: + +- `result.exit_code == 124` +- `timed_out == true` +- timeout-aware `runtime_evidence` + +The orchestration layer records the result and classifies it according to phase requirements. It must not call it behavioral FAIL merely because time expired. + +### Outer container timeout + +Owned by `invoke --container-timeout`. + +A valid result may not exist. This is infrastructure/non-evidence and may be eligible for explicit orchestration retry only if replay safety is established. + +### Eval-engine scheduling timeout + +No new global wall-clock timeout is introduced in Tasks 2-5. A later public eval CLI may add one explicitly if needed. + +## 12. Judge contract + +The generic engine has no universal judge JSON schema. + +The profile owns: + +- judge system instructions; +- judge prompt; +- expected output schema; +- parser; +- schema/contract validation; +- semantic interpretation. + +The generic engine owns: + +- running the judge via `invoke`; +- retry/timeout/infrastructure handling; +- preserving judge result provenance; +- distinguishing judge transport/contract failure from semantic FAIL. + +Loom's current strict JSON judge contract is a project adapter reference, not the generic schema. + +## 13. Skill-ablation extension boundary + +Loom skill-owned evals prove a future case may require: + +```text +baseline target -> baseline judge +candidate target -> candidate judge +comparison +``` + +Tasks 2-5 implement only the normal single target/judge lifecycle. + +However, code must be factored so the phase primitive can be reused: + +```python +run_evaluation_phase( + invocation_spec, + evidence_requirement, + retry_policy, +) -> PhaseOutcome +``` + +Task 8 can compose two such phases and a project-owned comparison callback. + +Do not add baseline/candidate fields to `NormalizedCase` or hard-code skill/trap/score semantics now. + +## 14. Error taxonomy + +Generic error codes should be stable enough for artifacts/tests but not overfit Loom strings. + +Minimum categories: + +### Infrastructure + +- `invoke_outer_timeout` +- `invoke_no_result` +- `invoke_invalid_result` +- `invoke_preflight_failed` +- `profile_prepare_failed` +- `judge_contract_invalid` + +### Product + +- `product_error` +- `product_timeout` + +### Evidence + +- `evidence_incomplete` +- `evidence_invalid` +- `evidence_unsupported` +- `evidence_field_unavailable` + +Task implementations may add narrower codes but must preserve the three-plane taxonomy. + +## 15. Reporting and exit semantics + +Task 5 may provide internal summary helpers but no public `eval` CLI yet. + +Internal summary semantics: + +- PASS corresponds to classification `pass`; +- FAIL corresponds to classification `fail`; +- ERROR/non-evidence corresponds to `non-evidence`; +- a run succeeds only if every selected job is `pass`; +- reporting happens from verified durable artifacts where artifacts are enabled. + +Human-readable formatting remains replaceable; durable artifacts are the stronger resumption/debugging surface. + +## 16. Explicit migration mapping from Loom + +| Current Loom responsibility | Destination | +| --- | --- | +| behavioral/skill case discovery | Loom profile | +| normalization of historical skill formats | Loom profile | +| case IDs/selectors | profile produces `NormalizedCase`; generic selection uses them | +| setup_projects / setup_skill_ablation_projects | Loom profile `prepare` | +| target agent wrappers | Loom profile | +| runner/container invocation | existing low-level `invoke` | +| direct-container fallback | remove/do not migrate | +| Loom observer attachment | remove as authority/do not migrate | +| tool-result reconstruction | remove as authority/do not migrate | +| runner-safety candidate adapter | remove/do not migrate | +| invoke_container_with_retry loop | generic `eval_execute` with explicit policy | +| provider.no-route predicate | optional policy/reference, not hidden engine rule | +| target_scoring_evidence_error legacy checks | replace with runtime-evidence requirements + profile checks | +| deterministic_failures | Loom profile returning `CheckOutcome` | +| judge_prompt | Loom profile | +| parse_judge / judge_contract_error | Loom profile | +| semantic_pass | Loom profile returning `SemanticDecision` | +| classify_behavioral_result | generic classifier over normalized outcomes | +| artifact run claim/write/hash/verify | generic artifact module | +| case x iteration expansion | generic planner | +| standard/runtime concurrency | generic planner; profile assigns lane | +| skill ablation orchestration | Task 8 generic paired mode + Loom comparison policy | +| console formatting | thin project/public CLI layer; not core semantics | + +## 17. Task handoff + +### Task 2 / #59 may proceed with + +- `NormalizedCase` +- `EvalJob` +- `RunPlan` +- explicit-selection safeguards +- stable labels +- standard/runtime lanes +- deterministic case x iteration expansion +- separated concurrency resolution + +It must not execute models or create artifacts. + +### Task 3 / #60 may proceed with + +- run and per-job artifact schemas; +- directory claiming; +- stable artifact paths; +- atomic writes; +- canonical hash/integrity verification; +- opaque project metadata; +- preserving future target/judge/readiness/check data without interpreting it. + +### Task 4 / #61 may proceed with + +- `InvocationSpec` +- `AttemptRecord` +- three-plane `AttemptFailure` +- `RetryPolicy` +- `EvidenceRequirement` +- `EvidenceReadiness` +- boundary and exact-field readiness helpers; +- target phase execution over the existing `invoke` path. + +### Task 5 / #62 may proceed with + +- judge phase execution using the same attempt machinery; +- profile judge construction/parsing; +- `CheckOutcome` +- `SemanticDecision` +- generic classification precedence; +- minimal one-case `eval_engine` composition; +- provider-free complete-lifecycle test doubles. + +No one of Tasks 2-5 should need to choose a new ownership model. + +## 18. Deferred decisions + +These are deliberately deferred, not missing from Task 1: + +- public `eval` CLI syntax — Task 6; +- external suite/profile loading mechanism — Task 6; +- whether the internal artifact schemas become stable public schemas — Task 6; +- exact public retry defaults — Task 6; +- Loom migration adapter details — Task 7; +- generic baseline/candidate ablation mode — Task 8; +- final removal of Loom duplicated orchestration — Task 9; +- repository/product rename — separate product decision; +- input JSON request API — no current requirement. + +## 19. Architecture acceptance check + +This design satisfies issue #58 because: + +- every major Loom runner responsibility has an explicit owner; +- the generic target -> readiness -> judge -> classification -> artifact lifecycle is defined; +- retry ownership and replay-safety constraints are explicit; +- inner/outer timeout ownership is explicit; +- iteration/concurrency planning is explicit; +- artifacts have an implementable versioned internal shape; +- project extension points are explicit without creating an assertion DSL; +- `invoke` remains exactly one isolated invocation; +- `runtime_evidence/v1` remains the sole runtime authority; +- legacy Loom evidence and direct-container paths are explicitly excluded; +- Tasks 2-5 have concrete module/type handoffs and can proceed independently.