Repository navigation
docs: define generic eval-engine extraction architecture - #68
Merged
Merged
Conversation
## Scope Worker B for #58 (Task 1/9): post-PR-#45 runner contract/invariant inventory. This PR is intentionally documentation-only. It does **not** implement the eval engine, change `invoke`, migrate Loom, add a public `eval` CLI, or define the final Task 1 architecture. ## What it captures - one `invoke` = exactly one isolated invocation; - `opencode-eval-runner/v1` result boundary; - mandatory/authoritative `runtime_evidence/v1`; - assertion-scoped evidence readiness and field states; - product vs evidence vs infrastructure failure separation; - inner vs outer timeout behavior; - retry ownership above `invoke`, cross-checked against Loom's current narrow retry policy; - OpenCode vs Copilot transport capability differences; - host output validation and host/image compatibility; - CLI and GitHub Action interface limits; - explicit unsupported areas and trusted-checkout scope; - current Loom behaviors that must not accidentally become runner contracts. ## Validation Cross-checked against the landed post-#45 contract: - `main` / PR #45 merge commit `fd9da10cbe2a8182fc8910ec199a221250deb3ca` - final reviewed PR #45 head `25478106773969870931af3fb34e7b2586746606` - `runner/cli.py` - `container/invoke.py` - `docs/invocation-usage.md` - `docs/runtime-evidence-contract.md` - `action.yml` - Loom `functionality-anchor-requirements-coherance` - `scripts/run-evals.py` blob `30f2ae87764be180173a9e5b4b609ee79b2f063d` `eval-engine/01-architecture` now matches `main` at the PR #45 merge commit. The worker branch has been rebased onto that exact state, leaving a one-file documentation diff. Part of #58.
## Summary Part of #58 (Task 1/9), parallel worker A. Adds a focused inventory of the current Loom eval harness from: - bateau84/loom branch functionality-anchor-requirements-coherance - scripts/run-evals.py at 30f2ae87764be180173a9e5b4b609ee79b2f063d The inventory traces and classifies: - case loading/normalization; - target/workspace setup; - invocation and orchestration retries; - runtime evidence and legacy observers; - judge construction/parsing; - deterministic checks; - classification; - artifacts; - iterations/concurrency; - skill ablation; - reporting. It separates each responsibility into generic orchestration, Loom/project semantics, low-level invoke behavior, or legacy/compatibility behavior that should not migrate. ## Post-#45 validation Cross-checked against the merged runner contracts on the integration base: - one invoke remains one isolated invocation; - retries remain outside invoke; - runtime_evidence/v1 is the only authoritative runtime-evidence object; - tools/actions/tool_result_evidence/stdout/stderr/model text are diagnostic only; - project code owns cases, assertions, judging, thresholds, and behavioral meaning. The document explicitly marks Loom's older direct-container fallback, observer/tool-result reconstruction, runner-safety adapter, and Loom-side evidence_safety eligibility machinery as non-migration paths. ## Scope Documentation only. No runtime changes, no Loom migration, no final engine API/module design, no universal assertion DSL, and no public eval CLI. PR target: eval-engine/01-architecture.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Task
Closes #58 — Task 1 of 9: Define generic orchestration boundary and extraction contract.
This PR is architecture/documentation only. It does not implement the eval engine.
Inputs integrated
Parallel worker results were merged into the integration branch first:
Both are based on the post-#45 merged
maincontract.Architecture decisions
The final synthesis defines:
invoke= exactly one isolated invocation;runtime_evidence/v1as the only authoritative runtime-evidence source;No universal assertion DSL, public
evalCLI, input JSON API, Loom migration, or skill-ablation implementation is introduced.Files
docs/eval-engine-architecture.md— final Task 1 synthesisdocs/eval-engine-runner-contracts.md— worker B inventorydocs/loom-eval-runner-responsibility-inventory.md— worker A inventoryBranch topology
eval-engine/01-architecturemainAcceptance
The architecture is intended to let Tasks #59–#62 proceed independently without redesigning the ownership boundary.
Documentation only; no runtime behavior changes.