Skip to content

[Eval engine 1/9] Define generic orchestration boundary and extraction contract #58

Description

@bateau84

Sequence

Task 1 of 9 — do this first.

Prerequisite: PR #45 should be merged before implementation begins so the new engine is designed on top of the final invoke / runtime_evidence/v1 contract.

Do not begin Task 2 until this task is merged.

Goal

Define the generic evaluation-orchestration boundary by extracting the architecture already proven in Loom's scripts/run-evals.py, without moving Loom-specific behavioral semantics into this repository.

Reference implementation:

  • repo: bateau84/loom
  • branch: functionality-anchor-requirements-coherance
  • file: scripts/run-evals.py

Required work

Create an architecture/extraction document in this repository that inventories the current Loom runner and classifies each responsibility as one of:

  1. generic eval-engine responsibility;
  2. project/profile responsibility;
  3. low-level invoke responsibility;
  4. legacy/compatibility behavior that should not be carried forward.

Define the minimal internal contracts needed by the generic engine, including at least:

  • normalized case/run identity;
  • target invocation specification;
  • target result;
  • evidence-readiness result;
  • judge invocation specification/result;
  • pass | fail | non-evidence classification;
  • artifact/run identity;
  • iteration and concurrency plan;
  • explicit retry policy.

Define extension points for project-owned behavior such as:

  • case discovery/normalization;
  • fixture/workspace preparation;
  • judge prompt construction;
  • deterministic assertions;
  • semantic judge parsing;
  • project-specific metadata.

Invariants

  • invoke remains exactly one isolated invocation.
  • No hidden retries may be added to invoke.
  • Any retry policy belongs to the eval orchestration layer and must be explicit in artifacts.
  • runtime_evidence/v1 is the authoritative runtime-evidence source.
  • The eval engine may decide evidence readiness, but project code still owns domain semantics and what constitutes correct behavior.
  • Do not design a universal assertion DSL in this task.
  • Do not migrate Loom code yet.
  • Do not add the public eval CLI yet.

Acceptance criteria

  • A concrete module/API decomposition exists.
  • Every major responsibility in Loom's current run-evals.py is mapped to generic/project/legacy ownership.
  • The target→evidence-readiness→judge→classification lifecycle is explicit.
  • Retry, timeout, iteration, concurrency, and artifact responsibilities are unambiguous.
  • The document is specific enough that Tasks 2–5 can implement independent modules without revisiting the architecture.

Implementation prompt

Work on bateau84/opencode-eval-runner issue #58: Define generic orchestration boundary and extraction contract.

This is Task 1 of 9. Do not implement the eval engine yet.

Use Loom's bateau84/loom branch functionality-anchor-requirements-coherance, especially scripts/run-evals.py, as the proven reference implementation. Inventory its responsibilities and separate them into generic eval-engine behavior, project/profile behavior, low-level invoke behavior, and legacy behavior that should not be migrated.

Preserve these invariants: one invoke means exactly one isolated invocation; retries exist only in orchestration; runtime_evidence/v1 is authoritative; project code owns behavioral semantics; no universal assertion DSL; no Loom migration; no public eval CLI yet.

Produce a concrete architecture/extraction document and internal API/module decomposition sufficient for Tasks 2–5 to implement without redesigning the boundary. Include target execution, evidence readiness, judge execution, classification, artifacts, iterations, concurrency, retries, and project extension points.

Validate the document against the actual current Loom script and the post-PR-#45 runner contracts. Keep the work architecture/documentation only. Open a focused PR and reference issue #58.

Branch and resume contract

Integration branch: eval-engine/01-architecture
Final PR target: main

All implementation work for this issue converges on the integration branch above. Parallel worker PRs must target the integration branch, never main directly.

If a chat/session is lost, resume by reading this issue and inspecting:

  1. the integration branch eval-engine/01-architecture;
  2. open PRs targeting that branch;
  3. merged worker PRs/commits on that branch;
  4. the issue acceptance criteria.

Do not invent a new branch name when resuming an existing task.

Worker branches

  • eval-engine/01-loom-inventory — Loom responsibility inventory
  • eval-engine/01-runner-contracts — Runner contract/invariant inventory

Parallel work plan

This issue can be researched in parallel before one owner synthesizes the final architecture. Parallel workers must not independently redefine the final shared contracts.

Parallel prompt A — Loom responsibility inventory

Branch: eval-engine/01-loom-inventory
PR target: eval-engine/01-architecture

Analyze bateau84/loom branch functionality-anchor-requirements-coherance, especially scripts/run-evals.py.

Produce a responsibility inventory of the current Loom eval runner. Group functions/behaviors into: generic orchestration, Loom/project semantics, low-level invocation concerns, and legacy/compatibility behavior that should not migrate.

Trace case loading/normalization, target setup, invocation, retry, runtime evidence/legacy observers, judge construction/parsing, deterministic checks, classification, artifacts, iterations, concurrency, skill ablation, and reporting.

Commit the inventory as a focused documentation artifact on this branch. Do not design the final architecture or modify runtime code.

Parallel prompt B — runner contract/invariant inventory

Branch: eval-engine/01-runner-contracts
PR target: eval-engine/01-architecture

Analyze the post-PR-#45 runner contracts.

Inventory the constraints the future eval engine must preserve: invoke semantics, result/v1, runtime_evidence/v1, evidence-readiness rules, timeouts, transport behavior, output validation, CLI/Action interfaces, and unsupported areas.

Commit a contract/invariant matrix as focused documentation. Do not redesign the engine or modify runtime behavior.

Integration order

Create eval-engine/01-architecture from main after PR #45 is merged. Create both worker branches from that integration branch. Merge both worker PRs into the integration branch, then the issue owner synthesizes the final architecture document and opens the issue PR from eval-engine/01-architecture to main.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions