diff --git a/README.md b/README.md index 08b8629..50d001d 100644 --- a/README.md +++ b/README.md @@ -285,7 +285,7 @@ Example: ## Result contract -Each invocation writes one JSON document: +Each invocation writes one `opencode-eval-runner/v1` JSON document. Every official result includes the canonical `runtime_evidence` object. ```json { @@ -299,17 +299,66 @@ Each invocation writes one JSON document: "exit_code": 0, "session_id": "...", "text": "...", - "tools": ["skill"], - "actions": [{"tool": "skill", "args": {"id": "architectural-design"}}], - "skills_loaded": ["architectural-design"], + "tools": [], + "actions": [], + "skills_loaded": [], "stderr": "", - "stdout": "..." + "stdout": "...", + "runtime_evidence": { + "schema": "opencode-eval-runner/runtime-evidence/v1", + "status": "complete", + "evidence_eligible": true, + "observations": [], + "coverage": { + "observation_closed": {"state": "available", "value": true}, + "process_state": "completed", + "starts": {"state": "available", "value": 0}, + "terminals": {"state": "available", "value": 0}, + "missing_terminals": {"state": "available", "value": 0}, + "observer_failures": {"state": "available", "value": 0}, + "callback_failures": {"state": "available", "value": 0}, + "losses": [], + "unsupported": ["stock_codemode_final_boundary_not_exposed"], + "boundaries": { + "native": { + "status": "complete", + "evidence_eligible": true, + "starts": {"state": "available", "value": 0}, + "terminals": {"state": "available", "value": 0}, + "missing_terminals": {"state": "available", "value": 0}, + "issues": [] + }, + "code_mode_execution": { + "status": "complete", + "evidence_eligible": true, + "starts": {"state": "available", "value": 0}, + "terminals": {"state": "available", "value": 0}, + "missing_terminals": {"state": "available", "value": 0}, + "issues": [] + }, + "code_mode_finality": { + "status": "unsupported", + "evidence_eligible": false, + "starts": {"state": "unsupported", "reason": "stock_codemode_final_boundary_not_exposed"}, + "terminals": {"state": "unsupported", "reason": "stock_codemode_final_boundary_not_exposed"}, + "missing_terminals": {"state": "unsupported", "reason": "stock_codemode_final_boundary_not_exposed"}, + "issues": ["stock_codemode_final_boundary_not_exposed"] + } + } + } + } } ``` -The container emits this object as a single JSON line on stdout. The host harness writes artifact files itself, so no writable bind mount is required for result transport. +The container emits this object as a single JSON line on stdout. The host harness re-validates `runtime_evidence` before it writes the result artifact, so a missing or malformed runtime-evidence object is rejected rather than silently downgraded. -The eval repository decides whether that observed behavior is PASS, FAIL, or non-evidence. +`runtime_evidence` is required for every official transport result. `github-copilot-cli` also emits the canonical object, but with `status: "unsupported"` because it has no OpenCode runtime observer. + +If you override `--image`, treat the host runner and image as one compatibility pair. Legacy or custom images that do not emit a valid `opencode-eval-runner/runtime-evidence/v1` object are rejected by this host version; upgrade the host executable and image together. + +For runtime verdicts, do not use top-level `evidence_eligible` by itself. An assertion may use evidence only when every boundary it requires is `complete` and every exact field it requires is `available`. `redacted`, `omitted`, or `unsupported` required fields are not PASS evidence. Existing `tools`, `actions`, `tool_result_evidence`, stdout/stderr, and model text are diagnostic/convenience data and must not fill an authoritative-evidence gap. See [Runtime evidence contract v1](docs/runtime-evidence-contract.md). + +The eval repository still owns the assertion semantics and decides PASS, FAIL, or non-evidence after applying those eligibility rules. ## Image versions diff --git a/docs/runtime-evidence-contract.md b/docs/runtime-evidence-contract.md index c4a9de1..01f7c72 100644 --- a/docs/runtime-evidence-contract.md +++ b/docs/runtime-evidence-contract.md @@ -53,6 +53,19 @@ Dynamic observation fields use exactly these public states: Unknown counts are never converted to \`0\`. +## Consumer decision rule + +A consumer must decide evidence eligibility per assertion, not from the top-level flag alone: + +1. Validate the `runtime_evidence` object against this schema before using it. +2. If overall status is `incomplete` or `invalid`, no runtime assertion is eligible for PASS. +3. Declare the boundary or boundaries required by the assertion. Every required boundary must be `complete`. +4. If the assertion depends on an exact observation field, that field must be `available`. `redacted` or `omitted` makes that value-dependent assertion incomplete; `unsupported` makes it unsupported. +5. Never fill a missing authoritative fact from `tools`, `actions`, `tool_result_evidence`, stdout/stderr, model text, or workspace files. + +For example, an assertion that a direct tool ran can depend on `native`. An assertion about a Code Mode inner tool identity/input can depend on `code_mode_execution`. An assertion about the exact final value or error seen by a Code Mode script depends on `code_mode_finality` and is therefore unsupported on stock OpenCode 2.0.23. + + ## Native observation The runner-owned stock OpenCode 2.0.23 observer records: @@ -151,6 +164,19 @@ Product outcome remains independent: - a successful product result can have incomplete evidence; - \`exit_code\` and timeout status do not become evidence eligibility. +## Operational bounds + +The current stock observer bounds authoritative capture to 8 MiB and at most 20,001 JSONL events, and bounds each projected dynamic field to 256 KiB. Crossing an aggregate capture bound makes the capture invalid/ineligible; the runner does not truncate it into apparently complete evidence. A field that exceeds its field bound is explicitly `omitted` with reason `size_limit`, so assertions requiring that exact value remain ineligible. + +These are current adapter safety limits, not provider or model guarantees. + + +## Compatibility and migration + +The outer result schema remains `opencode-eval-runner/v1`, but `runtime_evidence` is now mandatory. The host CLI validates it after container execution and rejects a result that omits it or violates this v1 contract. + +Official OpenCode and Copilot images in this revision emit the required object. A legacy or custom image built against the older result shape must be upgraded together with the host runner. This does not require any OpenCode modification: the supported OpenCode profile uses stock 2.0.23. Copilot results satisfy the result-shape requirement by reporting runtime evidence as explicitly `unsupported`. + ## Unsupported areas - \`code_mode_finality\`: stock OpenCode 2.0.23 does not expose the exact final value/error seen by each Code Mode script call.