Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
47 changes: 47 additions & 0 deletions .github/workflows/runtime-evidence-acceptance.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
name: Provider-free runtime evidence acceptance

on:
pull_request:
paths:
- 'Containerfile'
- 'container/**'
- 'runner/**'
- 'tests/integration/**'
- 'tests/test_runtime_evidence_acceptance.py'
- '.github/workflows/runtime-evidence-acceptance.yml'
workflow_dispatch:

permissions:
contents: read

jobs:
provider-free-acceptance:
runs-on: ubuntu-latest
timeout-minutes: 15
steps:
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
with:
persist-credentials: false

- name: Check acceptance driver
run: |
python3 -m py_compile tests/integration/run_runtime_evidence_acceptance.py
python3 -m unittest discover -s tests -p 'test_runtime_evidence_acceptance.py' -v

- name: Build stock OpenCode 2.0.23 runner image
run: docker build -f Containerfile --target opencode -t opencode-eval-runner:runtime-evidence-acceptance .

- name: Run provider-free acceptance gate
run: |
python3 tests/integration/run_runtime_evidence_acceptance.py \
--image opencode-eval-runner:runtime-evidence-acceptance \
--output runtime-evidence-acceptance

- name: Preserve sanitized acceptance report
if: always()
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02
with:
name: runtime-evidence-acceptance-${{ github.event.pull_request.head.sha || github.sha }}
path: runtime-evidence-acceptance/
if-no-files-found: error
retention-days: 14
2 changes: 1 addition & 1 deletion Containerfile
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
FROM node:24-bookworm-slim@sha256:0e0ff40c39bc087845bfb27465a0df4ea419520094bc35842ff83dd8cbe6f9b6 AS opencode-builder
ARG OPENCODE_VERSION=2.0.18
ARG OPENCODE_VERSION=2.0.23
RUN npm install --global "@opencode/cli@${OPENCODE_VERSION}" \
&& resolved="$(readlink -f "$(command -v opencode)")" \
&& test -x "$resolved" \
Expand Down
18 changes: 15 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,7 +72,7 @@ Known API-key environment variables are passed when present:

Additional variables require explicit `--env NAME`.

Reasoning can be pinned explicitly with `--reasoning LEVEL`. The pinned OpenCode 2.0.18 CLI represents a model variant in the model reference, so the runner maps `--model provider/model --reasoning LEVEL` to `opencode run --model provider/model#LEVEL`. Supplying both a `#variant` in `--model` and `--reasoning` is rejected as ambiguous. If the model reference already contains a variant and `--reasoning` is omitted, the result records that variant with `"reasoning_source": "model-variant"`. If neither form supplies a level, the runner leaves OpenCode's provider/model default untouched and records `"reasoning": "provider-default"`.
Reasoning can be pinned explicitly with `--reasoning LEVEL`. The pinned stock OpenCode 2.0.23 CLI represents a model variant in the model reference, so the runner maps `--model provider/model --reasoning LEVEL` to `opencode run --model provider/model#LEVEL`. Supplying both a `#variant` in `--model` and `--reasoning` is rejected as ambiguous. If the model reference already contains a variant and `--reasoning` is omitted, the result records that variant with `"reasoning_source": "model-variant"`. If neither form supplies a level, the runner leaves OpenCode's provider/model default untouched and records `"reasoning": "provider-default"`.

### `github-copilot-cli`

Expand Down Expand Up @@ -173,7 +173,7 @@ opencode-eval-runner invoke \
...
```

OpenCode 2.0.18 does not expose the old singular `debug agent <id>` command that returned a resolved tool map. The runner therefore performs the strongest supported zero-inference preflight: it requires the expected plugin entrypoint to be materialized in the isolated OpenCode config, runs `opencode debug agents` to prove the configured location starts successfully with plugins active, and requires the selected agent to resolve. Missing plugin materialization, plugin/startup failure, or missing agent is infrastructure/non-evidence, never a behavioral FAIL. Actual tool use remains a repository-owned behavioral assertion in the eval corpus.
Stock OpenCode 2.0.23 does not expose the old singular `debug agent <id>` command that returned a resolved tool map. The runner therefore performs the strongest supported zero-inference preflight: it requires the expected plugin entrypoint to be materialized in the isolated OpenCode config, runs `opencode debug agents` to prove the configured location starts successfully with plugins active, and requires the selected agent to resolve. Missing plugin materialization, plugin/startup failure, or missing agent is infrastructure/non-evidence, never a behavioral FAIL. Actual tool use remains a repository-owned behavioral assertion in the eval corpus.

### Evaluating a skill

Expand Down Expand Up @@ -315,7 +315,7 @@ The eval repository decides whether that observed behavior is PASS, FAIL, or non

The transport images currently pin:

- OpenCode CLI `2.0.18`
- OpenCode CLI `2.0.23`
- GitHub Copilot CLI `1.0.83`

The two CLIs are not bundled together. OpenCode's npm package is used only as a build-time native-binary selector; GitHub Copilot CLI is installed from its native release installer. Node/npm are absent from the final runtime images.
Expand All @@ -340,6 +340,18 @@ OPENCODE_EVAL_RUNNER_COPILOT_IMAGE=...

Tags matching `v*` are published with `opencode-` and `copilot-` prefixes.

## Evaluation trust model

The normal evaluation profile is a **trusted-checkout** profile. It assumes the runner, pinned stock OpenCode runtime, reviewed instrumentation, and explicitly selected evaluated checkout/dependencies are trusted components of the evaluation environment.

They are not trusted merely because they produce data that looks like evidence. Model prose, tool-returned collector-shaped JSON, target-writable files, requested actions, inferred identities, and reconstructed results do not establish that an event occurred.

Authoritative runtime observations must come from reviewed instrumentation observing actual execution. Missing, partial, ambiguous, or unsupported required observations are non-evidence and must fail closed for the affected assertion.

This profile does **not** claim resistance to an evaluated plugin that deliberately compromises the trusted runtime or instrumentation. Hostile-plugin isolation is a separate optional profile, not a prerequisite for normal Loom evaluation.

See [Trusted-checkout runtime evidence](docs/trusted-checkout-evidence.md).

## Security boundary

The runner:
Expand Down
59 changes: 59 additions & 0 deletions docs/stock-opencode-2.0.23-observation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,59 @@
# Stock OpenCode 2.0.23 observation surface

Purpose: implementation reference for the trusted-checkout evidence profile.

Source checkpoint: stock OpenCode **v2.0.23** (`0fd7e2829449b052abf0078666669302923d77af`). This is distilled from the source assessment performed in superseded PR #43.

OpenCode remains stock and immutable. A missing observation boundary is reported as unsupported; it is not a reason to patch OpenCode or add a hostile-runtime broker.

## Useful stock surfaces

| Observation need | Stock surface | Status |
| --- | --- | --- |
| Live runtime events | `ctx.event.subscribe()` | supported source; ordering/drain must be proven by integration test |
| Session creation / ancestry | `session.created` + Session API | supported |
| Agent for a step | Session step/message events | supported |
| Native tool call identity/input | Session tool input/called events | supported source |
| Native terminal success/failure | Session tool success/failed events | supported source |
| Tool pre-execution hook | `ctx.tool.hook("execute.before")` | supported; occurs before tool decode/execution |
| Tool post-handler hook | `ctx.tool.hook("execute.after")` | supported; occurs after handler result but before later core normalization |
| Tool registration wrapping | `ctx.tool.transform(...)` | supported candidate for reviewed same-process instrumentation |
| Code Mode inner name/input/status | Code Mode metadata + tool hooks | supported source |
| Code Mode unique inner invocation + exact final caller value/error | no single public final boundary demonstrated | **must be proven or marked unsupported** |

## Native calls

Stock Session events are the preferred source for native terminal facts because they represent the runtime's own Session lifecycle rather than model or tool payload claims.

The observer must bind call identity, Session, agent/message context, input and terminal result/error without reconstructing them from prose or matching by value.

## Code Mode

Code Mode executes inner tools through the normal tool registry, so same-process reviewed instrumentation can observe real inner execution without isolating Loom.

The difficult part is not security; it is exact correlation and finality:

- inner calls share the outer `execute` context in stock OpenCode;
- public Code Mode metadata records name/input/status but not each inner returned value/error;
- `execute.after` is before later core normalization;
- concurrent identical inner calls must not be paired by FIFO, input equality, or completion order.

The first implementation should test a runner-owned observer plugin using supported tool transforms/hooks and runtime events. It must allocate a unique observation identity at an actual execution boundary and prove how that identity reaches the final inner value/error.

If that exact binding cannot be demonstrated for a case, the affected result field remains unavailable and the assertion cannot PASS.

## Ordering and completeness

A monotonic observer sequence is useful, but sequence alone is not completeness. The integration must also account for starts, terminals, observer loss, process interruption, and required descendant Sessions.

Absence assertions are eligible only when the relevant scope is complete. Missing capture is never interpreted as "did not happen".

## What is intentionally not required

- plugin/process isolation from the trusted Loom checkout;
- remote PluginHost or capability broker;
- evidence-channel peer authentication against same-authority attackers;
- patched/forked OpenCode;
- cryptographic evidence authenticity after runtime compromise.

Those belong only to a future optional untrusted-plugin profile.
127 changes: 127 additions & 0 deletions docs/trusted-checkout-evidence.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,127 @@
# Trusted-checkout runtime evidence

Status: **replacement direction for TRUST-001**.

This document supersedes the hostile-runtime direction explored in PR #41 and PR #43. Those PRs remain useful research/reference material, but normal Loom evaluation does not require the runner to defend itself from a deliberately malicious Loom checkout that shares its runtime authority.

## Contract

### TRUST-001 — authoritative runtime observation

For evaluation of an explicitly trusted checkout, evidence used for scoring MUST originate from reviewed runtime instrumentation observing actual execution.

The following MUST NOT independently establish that an event occurred:

- model assertions or generated prose;
- tool payloads shaped like collector/evidence records;
- requested or intended actions;
- inferred actor, parent, or execution identity;
- reconstructed results;
- target-writable evidence files.

Required observations MUST preserve enough runtime identity and ordering to evaluate the consumer contract, including actor/session/call identity, input, result or error, parent binding where applicable, and execution order.

Missing, partial, ambiguous, lost, or unsupported required observations MUST make the affected assertion ineligible for PASS. The runner MUST NOT fill gaps from model text, stdout, workspace files, or guessed correlations.

The trusted-checkout profile does not claim protection against malicious modification of the runner, stock OpenCode process, reviewed instrumentation, evaluated checkout, or their dependencies.

## Trust model

Trusted components:

- the selected `opencode-eval-runner` revision;
- pinned **stock OpenCode 2.0.23**;
- reviewed runtime instrumentation;
- the explicitly selected Loom checkout and its reviewed dependencies;
- host-side evidence projection/persistence code.

Not trusted as evidence authority:

- model output;
- agent claims;
- tool-returned collector-shaped data;
- normal product/session/workspace files;
- caller-supplied identity or completeness claims.

This is an evaluation-correctness boundary, not a hostile-code security boundary.

## Required evidence behavior

The target behavior remains strict even though the security scope is smaller:

- **Native calls:** observe the actual runtime call, actor/session/call identity, accepted/executable input, and terminal result/error.
- **Code Mode inner calls:** assign a unique runtime observation identity per actual inner invocation, bind it to the real outer `execute` call, and observe the final value/error that Code Mode exposes to the script.
- **Delegation:** derive child Session identity and ancestry from runtime facts, not a parent result payload.
- **Ordering:** preserve runtime observation order; do not correlate concurrent calls by FIFO or input equality.
- **Completeness:** explicitly report missing starts/terminals, capture loss, unsupported boundaries, and incomplete scope.
- **Confidentiality:** redact or omit credentials before the runner first persists, clips, logs, or exports evidence.
- **Noninterference:** observation must not add product retries or change normal Loom execution semantics.

If stock OpenCode's supported interfaces cannot expose an exact required boundary, the result is `unsupported`/ineligible for that assertion. The response is not to invent evidence and not to turn the normal profile into a hostile-code isolation project.

## Implementation direction

Keep the normal path:

```text
Loom eval harness
-> opencode-eval-runner invoke
-> stock OpenCode 2.0.23
+ reviewed runner-owned observation instrumentation
+ trusted Loom checkout
-> safe host projection
-> Loom judging
```

The preferred implementation is same-process reviewed instrumentation using supported stock OpenCode plugin/runtime surfaces. It may use a runner-owned observer plugin, tool/session hooks, live runtime events, and reviewed wrappers where those surfaces preserve the required boundary.

OpenCode source patches, forks, remote PluginHost isolation, evidence signing, and a capability broker are not requirements of this profile.

Provider-free integration tests must prove the exact observation/correlation behavior before a field becomes eligible evidence.

See [Stock OpenCode 2.0.23 observation surface](stock-opencode-2.0.23-observation.md) for the retained source/capability findings from PR #43.

## Reuse from PR #41

| Work | Disposition |
| --- | --- |
| Evidence-safety projection/redaction and fail-closed field handling | **Reuse/adapt**; keep the behavior, decouple it from hostile-runtime image/signing assumptions |
| Credential protection before host/file/print sinks | **Reuse** |
| Disposable OpenCode state/profile work | **Reuse where useful** for deterministic eval isolation |
| Normal `invoke` compatibility and provider-free integration probes | **Reuse/adapt** to stock 2.0.23 |
| Native/Code Mode observation schemas and concurrency tests | **Reuse as behavioral requirements/tests** |
| Delegated-session identity/ancestry probes | **Reuse** |
| Patched OpenCode runtime | **Drop** |
| HMAC observer/import trust boundary | **Drop** for the normal profile |
| protected-channel / remote tool service | **Drop** |
| plugin isolation / remote PluginHost work | **Drop** |
| Cosign evidence-authenticity machinery | **Drop** as a TRUST-001 prerequisite |
| adversarial same-authority attack tests | **Move to optional future untrusted profile** |

## Reuse from PR #43

| Work | Disposition |
| --- | --- |
| Stock OpenCode 2.0.23 source/capability assessment | **Reuse** |
| Identification of public Session/event/tool surfaces | **Reuse** |
| Scope/completeness rules that prevent false absence/PASS | **Reuse and simplify** |
| First-sink confidentiality inventory | **Reuse and simplify** |
| Loom callback/capability inventory | **Reference when needed for compatibility** |
| hostile-runtime TCB/authority model | **Drop** from the normal profile |
| isolated Loom execution domain | **Drop** |
| capability/evidence-channel peer-authentication requirements | **Drop** |
| OCI adversarial boundary experiment/gates | **Drop** |

## Implementation sequence

This is normal engineering work, not a multi-authorization security experiment:

1. Pin and verify stock OpenCode 2.0.23.
2. Add the smallest reviewed observation instrumentation that can capture native and Code Mode execution without changing product semantics.
3. Port the useful PR #41 evidence-safety projection so captured values are protected before persistence/export.
4. Add explicit completeness/loss fields and fail closed when required data is missing.
5. Exercise direct, Code Mode, delegation, error, timeout, and concurrent reverse-completion cases provider-free.
6. Compose through Loom's existing `eval:live -> run-evals.py -> invoke` path.
7. Only after those checks pass should Loom consume the new evidence schema for PASS/FAIL decisions.

A future **untrusted-plugin execution profile** may add isolation if there is a real need to evaluate hostile plugin code. It must remain optional and separate from the normal trusted-checkout path.
Loading
Loading