Skip to content

docs: inventory Loom eval runner responsibilities - #66

Merged
bateau84 merged 1 commit into
eval-engine/01-architecturefrom
eval-engine/01-loom-inventory
Oct 7, 2026
Merged

bateau84 merged 1 commit into
eval-engine/01-architecturefrom
eval-engine/01-loom-inventory

Conversation

@bateau84

@bateau84 bateau84 commented Oct 7, 2026

Copy link
Copy Markdown
Owner

Summary

Part of #58 (Task 1/9), parallel worker A.

Adds a focused inventory of the current Loom eval harness from:

  • bateau84/loom branch functionality-anchor-requirements-coherance
  • scripts/run-evals.py at 30f2ae87764be180173a9e5b4b609ee79b2f063d

The inventory traces and classifies:

  • case loading/normalization;
  • target/workspace setup;
  • invocation and orchestration retries;
  • runtime evidence and legacy observers;
  • judge construction/parsing;
  • deterministic checks;
  • classification;
  • artifacts;
  • iterations/concurrency;
  • skill ablation;
  • reporting.

It separates each responsibility into generic orchestration, Loom/project semantics, low-level invoke behavior, or legacy/compatibility behavior that should not migrate.

Post-#45 validation

Cross-checked against the merged runner contracts on the integration base:

  • one invoke remains one isolated invocation;
  • retries remain outside invoke;
  • runtime_evidence/v1 is the only authoritative runtime-evidence object;
  • tools/actions/tool_result_evidence/stdout/stderr/model text are diagnostic only;
  • project code owns cases, assertions, judging, thresholds, and behavioral meaning.

The document explicitly marks Loom's older direct-container fallback, observer/tool-result reconstruction, runner-safety adapter, and Loom-side evidence_safety eligibility machinery as non-migration paths.

Scope

Documentation only.

No runtime changes, no Loom migration, no final engine API/module design, no universal assertion DSL, and no public eval CLI.

PR target: eval-engine/01-architecture.

@bateau84
bateau84 merged commit bdcf9e3 into eval-engine/01-architecture Oct 7, 2026
1 check passed
@bateau84
bateau84 deleted the eval-engine/01-loom-inventory branch October 7, 2026 08:20
bateau84 added a commit that referenced this pull request Oct 7, 2026
## Task

Closes #58 — **Task 1 of 9: Define generic orchestration boundary and
extraction contract**.

This PR is architecture/documentation only. It does not implement the
eval engine.

## Inputs integrated

Parallel worker results were merged into the integration branch first:

- #65 — runner contract/invariant inventory
- #66 — Loom eval-runner responsibility inventory

Both are based on the post-#45 merged `main` contract.

## Architecture decisions

The final synthesis defines:

- one `invoke` = exactly one isolated invocation;
- retry ownership exclusively in orchestration;
- `runtime_evidence/v1` as the only authoritative runtime-evidence
source;
- explicit separation of product, evidence, and infrastructure outcomes;
- generic ownership for planning, iterations, concurrency, attempts,
evidence-readiness mechanics, judge lifecycle, tri-state classification,
artifacts, and summaries;
- project/profile ownership for case semantics, fixtures, prompts,
evidence requirements, deterministic assertions, judge meaning,
thresholds, and ablation policy;
- explicit non-migration of Loom's old direct-container fallback, legacy
observer/evidence reconstruction, and superseded runner-safety paths;
- concrete internal module decomposition for Tasks 2–5;
- concrete internal contracts for normalized cases/jobs, invocation
specs, attempt records, retry policy, evidence requirements/readiness,
check outcomes, semantic decisions, and durable artifacts;
- initial standard/runtime concurrency-lane behavior;
- versioned internal run/artifact schemas;
- explicit Task 2–5 handoff boundaries.

No universal assertion DSL, public `eval` CLI, input JSON API, Loom
migration, or skill-ablation implementation is introduced.

## Files

- `docs/eval-engine-architecture.md` — final Task 1 synthesis
- `docs/eval-engine-runner-contracts.md` — worker B inventory
- `docs/loom-eval-runner-responsibility-inventory.md` — worker A
inventory

## Branch topology

- integration branch: `eval-engine/01-architecture`
- base: `main`
- worker branches were merged into the integration branch before
synthesis.

## Acceptance

The architecture is intended to let Tasks #59–#62 proceed independently
without redesigning the ownership boundary.

Documentation only; no runtime behavior changes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant