Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
63 changes: 56 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -285,7 +285,7 @@ Example:

## Result contract

Each invocation writes one JSON document:
Each invocation writes one `opencode-eval-runner/v1` JSON document. Every official result includes the canonical `runtime_evidence` object.

```json
{
Expand All @@ -299,17 +299,66 @@ Each invocation writes one JSON document:
"exit_code": 0,
"session_id": "...",
"text": "...",
"tools": ["skill"],
"actions": [{"tool": "skill", "args": {"id": "architectural-design"}}],
"skills_loaded": ["architectural-design"],
"tools": [],
"actions": [],
"skills_loaded": [],
"stderr": "",
"stdout": "..."
"stdout": "...",
"runtime_evidence": {
"schema": "opencode-eval-runner/runtime-evidence/v1",
"status": "complete",
"evidence_eligible": true,
"observations": [],
"coverage": {
"observation_closed": {"state": "available", "value": true},
"process_state": "completed",
"starts": {"state": "available", "value": 0},
"terminals": {"state": "available", "value": 0},
"missing_terminals": {"state": "available", "value": 0},
"observer_failures": {"state": "available", "value": 0},
"callback_failures": {"state": "available", "value": 0},
"losses": [],
"unsupported": ["stock_codemode_final_boundary_not_exposed"],
"boundaries": {
"native": {
"status": "complete",
"evidence_eligible": true,
"starts": {"state": "available", "value": 0},
"terminals": {"state": "available", "value": 0},
"missing_terminals": {"state": "available", "value": 0},
"issues": []
},
"code_mode_execution": {
"status": "complete",
"evidence_eligible": true,
"starts": {"state": "available", "value": 0},
"terminals": {"state": "available", "value": 0},
"missing_terminals": {"state": "available", "value": 0},
"issues": []
},
"code_mode_finality": {
"status": "unsupported",
"evidence_eligible": false,
"starts": {"state": "unsupported", "reason": "stock_codemode_final_boundary_not_exposed"},
"terminals": {"state": "unsupported", "reason": "stock_codemode_final_boundary_not_exposed"},
"missing_terminals": {"state": "unsupported", "reason": "stock_codemode_final_boundary_not_exposed"},
"issues": ["stock_codemode_final_boundary_not_exposed"]
}
}
}
}
}
```

The container emits this object as a single JSON line on stdout. The host harness writes artifact files itself, so no writable bind mount is required for result transport.
The container emits this object as a single JSON line on stdout. The host harness re-validates `runtime_evidence` before it writes the result artifact, so a missing or malformed runtime-evidence object is rejected rather than silently downgraded.

The eval repository decides whether that observed behavior is PASS, FAIL, or non-evidence.
`runtime_evidence` is required for every official transport result. `github-copilot-cli` also emits the canonical object, but with `status: "unsupported"` because it has no OpenCode runtime observer.

If you override `--image`, treat the host runner and image as one compatibility pair. Legacy or custom images that do not emit a valid `opencode-eval-runner/runtime-evidence/v1` object are rejected by this host version; upgrade the host executable and image together.

For runtime verdicts, do not use top-level `evidence_eligible` by itself. An assertion may use evidence only when every boundary it requires is `complete` and every exact field it requires is `available`. `redacted`, `omitted`, or `unsupported` required fields are not PASS evidence. Existing `tools`, `actions`, `tool_result_evidence`, stdout/stderr, and model text are diagnostic/convenience data and must not fill an authoritative-evidence gap. See [Runtime evidence contract v1](docs/runtime-evidence-contract.md).

The eval repository still owns the assertion semantics and decides PASS, FAIL, or non-evidence after applying those eligibility rules.

## Image versions

Expand Down
26 changes: 26 additions & 0 deletions docs/runtime-evidence-contract.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,6 +53,19 @@ Dynamic observation fields use exactly these public states:

Unknown counts are never converted to \`0\`.

## Consumer decision rule

A consumer must decide evidence eligibility per assertion, not from the top-level flag alone:

1. Validate the `runtime_evidence` object against this schema before using it.
2. If overall status is `incomplete` or `invalid`, no runtime assertion is eligible for PASS.
3. Declare the boundary or boundaries required by the assertion. Every required boundary must be `complete`.
4. If the assertion depends on an exact observation field, that field must be `available`. `redacted` or `omitted` makes that value-dependent assertion incomplete; `unsupported` makes it unsupported.
5. Never fill a missing authoritative fact from `tools`, `actions`, `tool_result_evidence`, stdout/stderr, model text, or workspace files.

For example, an assertion that a direct tool ran can depend on `native`. An assertion about a Code Mode inner tool identity/input can depend on `code_mode_execution`. An assertion about the exact final value or error seen by a Code Mode script depends on `code_mode_finality` and is therefore unsupported on stock OpenCode 2.0.23.


## Native observation

The runner-owned stock OpenCode 2.0.23 observer records:
Expand Down Expand Up @@ -151,6 +164,19 @@ Product outcome remains independent:
- a successful product result can have incomplete evidence;
- \`exit_code\` and timeout status do not become evidence eligibility.

## Operational bounds

The current stock observer bounds authoritative capture to 8 MiB and at most 20,001 JSONL events, and bounds each projected dynamic field to 256 KiB. Crossing an aggregate capture bound makes the capture invalid/ineligible; the runner does not truncate it into apparently complete evidence. A field that exceeds its field bound is explicitly `omitted` with reason `size_limit`, so assertions requiring that exact value remain ineligible.

These are current adapter safety limits, not provider or model guarantees.


## Compatibility and migration

The outer result schema remains `opencode-eval-runner/v1`, but `runtime_evidence` is now mandatory. The host CLI validates it after container execution and rejects a result that omits it or violates this v1 contract.

Official OpenCode and Copilot images in this revision emit the required object. A legacy or custom image built against the older result shape must be upgraded together with the host runner. This does not require any OpenCode modification: the supported OpenCode profile uses stock 2.0.23. Copilot results satisfy the result-shape requirement by reporting runtime evidence as explicitly `unsupported`.

## Unsupported areas

- \`code_mode_finality\`: stock OpenCode 2.0.23 does not expose the exact final value/error seen by each Code Mode script call.
Expand Down
Loading