Skip to content

[Eval engine 4/9] Add generic target execution, evidence readiness, and explicit retry policy #61

Description

@bateau84

Sequence

Task 4 of 9.

Depends on #60. Start only after #60 is merged.

Goal

Build the generic target-execution stage on top of the existing low-level invoke primitive.

Required lifecycle

normalized target job
  -> invoke exactly once
  -> validate result/v1
  -> decide runtime-evidence readiness
  -> optionally retry only when orchestration policy explicitly permits it
  -> record every attempt
  -> return target execution outcome

Reuse from Loom

Extract the generic parts of Loom's current handling for:

  • transport errors;
  • timeout/infrastructure classification;
  • explicit transient-provider retry policy;
  • target timing;
  • transport/model/reasoning provenance.

Do not copy Loom-specific observer compatibility code; PR #45's runtime_evidence/v1 replaces that authority.

Evidence-readiness procedure

Implement the documented decision procedure:

  1. validate runtime_evidence/v1;
  2. overall incomplete or invalid means no runtime assertion is PASS-eligible;
  3. required boundaries must be complete;
  4. exact required fields must be available;
  5. diagnostic fields may never backfill authoritative facts.

Expose this as a reusable API. It should return structured readiness/non-readiness reasons, not just a boolean.

Retry requirements

  • invoke itself must still perform one actual invocation.
  • Retry count defaults and retryable error classes must be orchestration configuration.
  • Every attempt must be represented in timing/artifacts.
  • Behavioral failure must never be retried as infrastructure failure.
  • Evidence incompleteness must not silently become a retry unless the configured policy explicitly says so.

Tests

Provider-free tests for:

  • successful target;
  • product failure with complete evidence;
  • timeout;
  • malformed/invalid result;
  • incomplete evidence;
  • unsupported boundary;
  • one retryable transport failure followed by success;
  • exhausted retries;
  • non-retryable behavioral/product failure.

Acceptance criteria

  • Generic target execution uses the existing public invoke behavior rather than duplicating container logic.
  • Evidence readiness has one canonical implementation reusable by Loom and the future eval CLI.

Branch and resume contract

Integration branch: eval-engine/04-target-execution
Final PR target: main

All worker PRs target the integration branch. Only the reconciled issue-level PR targets main.

If a chat/session is lost, resume from this issue plus the named branches/PRs. Do not invent replacement branch names.

Worker branches

  • eval-engine/04-evidence-readiness — A — canonical evidence readiness
  • eval-engine/04-target-retry — B — target attempt/retry orchestration

Copy/paste prompts

START / INTEGRATION OWNER — COPY THIS ENTIRE BLOCK

Work as integration owner for issue #61.

Integration branch: eval-engine/04-target-execution
Final target: main
Prerequisite: #60 merged.

Ensure integration branch exists from current main. Read issue #61 and Task 1 architecture. Establish only minimal shared execution/readiness types needed by workers.

Worker branches:
- eval-engine/04-evidence-readiness
- eval-engine/04-target-retry

Do not duplicate their scopes. If both PRs already exist, run the RECONCILIATION prompt.

Report integration SHA and dispatch readiness.

WORKER A — canonical evidence readiness — COPY THIS ENTIRE BLOCK

Work on bateau84/opencode-eval-runner issue #61, Task 4/9.

SUBTASK: canonical evidence-readiness API only.

Branch: eval-engine/04-evidence-readiness
PR target: eval-engine/04-target-execution
Do NOT target main.

Prerequisite: #60 merged. Branch from current integration head.

Read issue #61, docs/eval-engine-architecture.md, docs/runtime-evidence-contract.md.

Implement a pure reusable evidence-readiness evaluator over validated runtime_evidence/v1:
- EvidenceRequirement boundary set;
- EvidenceReadiness ready/incomplete/unsupported/invalid;
- overall incomplete/invalid fail closed;
- every required boundary must be complete;
- code_mode_finality unsupported remains unsupported;
- exact field readiness: available succeeds; redacted/omitted incomplete; unsupported unsupported;
- empty runtime requirement is allowed;
- no diagnostic/convenience field may backfill authoritative evidence.

Return structured reasons, not only bools.

Do NOT implement invoke/retry/attempt execution.
Add comprehensive provider-free unit tests.

Open focused PR to eval-engine/04-target-execution.
Report branch, commit, PR, tests, and API assumptions.

WORKER B — target attempt/retry orchestration — COPY THIS ENTIRE BLOCK

Work on bateau84/opencode-eval-runner issue #61, Task 4/9.

SUBTASK: target attempt and explicit retry orchestration.

Branch: eval-engine/04-target-retry
PR target: eval-engine/04-target-execution
Do NOT target main.

Prerequisite: #60 merged. Branch from current integration head.

Read issue #61 and docs/eval-engine-architecture.md.

Implement around the existing invoke boundary:
- InvocationSpec/AttemptRecord/AttemptFailure as approved;
- exactly one invoke per attempt;
- host vs product failure distinction;
- inner/outer timeout distinction;
- timing and transport/model/reasoning provenance;
- explicit RetryPolicy/RetryDecision;
- bounded attempt accounting and prior-error preservation;
- conservative replay safety;
- no retry of behavioral/product results as infrastructure;
- optional narrow transient-provider policy only if architecture allows, never hidden in invoke.

Do not duplicate OCI/container construction.
Do not reimplement evidence-readiness semantics. Depend on the sibling readiness API; if unavailable, use a narrow stub/interface and document it for reconciliation.

Add provider-free tests for success, product failure, timeout, invalid result, retry then success, exhausted retry, non-retryable failure, and visible attempt history.

Open focused PR to eval-engine/04-target-execution.
Report branch, commit, PR, tests, and assumptions.

RECONCILIATION — COPY THIS ENTIRE BLOCK AFTER THE REQUIRED WORKER PRs EXIST

Reconcile issue #61 after both worker PRs are ready.

Integration branch: eval-engine/04-target-execution
Workers:
- eval-engine/04-evidence-readiness
- eval-engine/04-target-retry
Final target: main

Review both PRs and CI. Merge them into integration, then replace any retry-worker readiness stub with the canonical sibling API.

Verify:
- one invoke per attempt;
- no hidden retry in invoke;
- structured evidence-readiness reasons;
- required boundary/field rules exactly match runtime-evidence contract;
- product/evidence/infrastructure planes remain separate;
- retries are explicit and artifact-ready;
- replay safety fails closed;
- all attempt/retry/readiness provider-free tests pass.

Resolve shared type duplication on integration. Update from current main if necessary. Open final PR to main referencing #61 and worker PRs. Do not start Task 5 or merge final PR unless explicitly tasked.

Report integration SHA, worker disposition, tests, final PR, and any remaining limitation.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions