Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 7 additions & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ jobs:

- name: Python checks
run: |
python3 -m py_compile runner/cli.py container/invoke.py
python3 -m py_compile runner/cli.py container/invoke.py tests/integration/run_code_mode_observer_probe.py
python3 -m unittest discover -s tests -p 'test_*.py'

- name: Build fake Copilot transport for action smoke test
Expand Down Expand Up @@ -69,6 +69,12 @@ jobs:
- name: Build OpenCode transport image
run: docker build -f Containerfile --target opencode -t opencode-eval-runner:opencode-test .

- name: Probe stock Code Mode inner observation boundary
run: |
python3 tests/integration/run_code_mode_observer_probe.py \
--image opencode-eval-runner:opencode-test \
--output /tmp/code-mode-observer-probe

- name: Build Copilot transport image
run: docker build -f Containerfile --target copilot -t opencode-eval-runner:copilot-test .

Expand Down
2 changes: 1 addition & 1 deletion Containerfile
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
FROM node:24-bookworm-slim@sha256:0e0ff40c39bc087845bfb27465a0df4ea419520094bc35842ff83dd8cbe6f9b6 AS opencode-builder
ARG OPENCODE_VERSION=2.0.18
ARG OPENCODE_VERSION=2.0.23
RUN npm install --global "@opencode/cli@${OPENCODE_VERSION}" \
&& resolved="$(readlink -f "$(command -v opencode)")" \
&& test -x "$resolved" \
Expand Down
18 changes: 15 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,7 +72,7 @@ Known API-key environment variables are passed when present:

Additional variables require explicit `--env NAME`.

Reasoning can be pinned explicitly with `--reasoning LEVEL`. The pinned OpenCode 2.0.18 CLI represents a model variant in the model reference, so the runner maps `--model provider/model --reasoning LEVEL` to `opencode run --model provider/model#LEVEL`. Supplying both a `#variant` in `--model` and `--reasoning` is rejected as ambiguous. If the model reference already contains a variant and `--reasoning` is omitted, the result records that variant with `"reasoning_source": "model-variant"`. If neither form supplies a level, the runner leaves OpenCode's provider/model default untouched and records `"reasoning": "provider-default"`.
Reasoning can be pinned explicitly with `--reasoning LEVEL`. The pinned stock OpenCode 2.0.23 CLI represents a model variant in the model reference, so the runner maps `--model provider/model --reasoning LEVEL` to `opencode run --model provider/model#LEVEL`. Supplying both a `#variant` in `--model` and `--reasoning` is rejected as ambiguous. If the model reference already contains a variant and `--reasoning` is omitted, the result records that variant with `"reasoning_source": "model-variant"`. If neither form supplies a level, the runner leaves OpenCode's provider/model default untouched and records `"reasoning": "provider-default"`.

### `github-copilot-cli`

Expand Down Expand Up @@ -173,7 +173,7 @@ opencode-eval-runner invoke \
...
```

OpenCode 2.0.18 does not expose the old singular `debug agent <id>` command that returned a resolved tool map. The runner therefore performs the strongest supported zero-inference preflight: it requires the expected plugin entrypoint to be materialized in the isolated OpenCode config, runs `opencode debug agents` to prove the configured location starts successfully with plugins active, and requires the selected agent to resolve. Missing plugin materialization, plugin/startup failure, or missing agent is infrastructure/non-evidence, never a behavioral FAIL. Actual tool use remains a repository-owned behavioral assertion in the eval corpus.
Stock OpenCode 2.0.23 does not expose the old singular `debug agent <id>` command that returned a resolved tool map. The runner therefore performs the strongest supported zero-inference preflight: it requires the expected plugin entrypoint to be materialized in the isolated OpenCode config, runs `opencode debug agents` to prove the configured location starts successfully with plugins active, and requires the selected agent to resolve. Missing plugin materialization, plugin/startup failure, or missing agent is infrastructure/non-evidence, never a behavioral FAIL. Actual tool use remains a repository-owned behavioral assertion in the eval corpus.

### Evaluating a skill

Expand Down Expand Up @@ -315,7 +315,7 @@ The eval repository decides whether that observed behavior is PASS, FAIL, or non

The transport images currently pin:

- OpenCode CLI `2.0.18`
- OpenCode CLI `2.0.23`
- GitHub Copilot CLI `1.0.83`

The two CLIs are not bundled together. OpenCode's npm package is used only as a build-time native-binary selector; GitHub Copilot CLI is installed from its native release installer. Node/npm are absent from the final runtime images.
Expand All @@ -340,6 +340,18 @@ OPENCODE_EVAL_RUNNER_COPILOT_IMAGE=...

Tags matching `v*` are published with `opencode-` and `copilot-` prefixes.

## Evaluation trust model

The normal evaluation profile is a **trusted-checkout** profile. It assumes the runner, pinned stock OpenCode runtime, reviewed instrumentation, and explicitly selected evaluated checkout/dependencies are trusted components of the evaluation environment.

They are not trusted merely because they produce data that looks like evidence. Model prose, tool-returned collector-shaped JSON, target-writable files, requested actions, inferred identities, and reconstructed results do not establish that an event occurred.

Authoritative runtime observations must come from reviewed instrumentation observing actual execution. Missing, partial, ambiguous, or unsupported required observations are non-evidence and must fail closed for the affected assertion.

This profile does **not** claim resistance to an evaluated plugin that deliberately compromises the trusted runtime or instrumentation. Hostile-plugin isolation is a separate optional profile, not a prerequisite for normal Loom evaluation.

See [Trusted-checkout runtime evidence](docs/trusted-checkout-evidence.md).

## Security boundary

The runner:
Expand Down
188 changes: 188 additions & 0 deletions docs/code-mode-observer-experiment.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,188 @@
# Code Mode inner-call observation experiment

Status: **bounded negative result with partial stock support**

Parent: PR #45 trusted-checkout runtime evidence.

Runtime checkpoint: stock OpenCode **v2.0.23**, tag commit
`0fd7e2829449b052abf0078666669302923d77af`.

This experiment asks one question only: how much of an actual inner Code Mode
tool invocation can reviewed same-process runner instrumentation observe without
patching OpenCode?

## Result

Stock OpenCode 2.0.23 supports a useful partial observation path:

| Fact | Stock supported surface | Result |
| --- | --- | --- |
| unique inner invocation identity | runner-owned `tool.transform` wrapper allocates an ID when the decoded leaf handler is actually entered | **supported** |
| actual selected tool | wrapper is attached to the effective registered tool | **supported** |
| executable input | wrapper runs after core input decoding and receives the value passed to the leaf handler | **supported** |
| real outer `execute` binding | the real `Tool.Context` carries Session/message/outer CallID into each inner leaf | **supported** |
| start / handler-terminal ordering | runner observer sequence around the transformed leaf handler | **supported** |
| exact final value delivered to the Code Mode script | no supported public stock boundary exposes it with unique inner identity | **unsupported** |
| exact final error seen by the script catch path | no supported public stock boundary exposes it with unique inner identity | **unsupported** |

The partial path is useful for proving that an inner call really entered a
particular tool with a particular decoded input. It is **not** sufficient for
assertions about the value/error ultimately observed by Code Mode.

The experiment therefore reports the final-caller capability as:

```json
{
"status": "unsupported",
"reason": "stock_codemode_final_boundary_not_exposed"
}
```

No value is reconstructed from an earlier result.

## Exact missing boundary

There are three distinct boundaries in stock 2.0.23.

### 1. Public transformed leaf handler

`packages/core/src/tool/runtime.ts` decodes input and then calls the registered
tool's `execute(decoded, context)`.

A runner-owned `ctx.tool.transform(...)` wrapper can therefore allocate a fresh
per-call ID at a real execution boundary and observe:

- the effective tool registration;
- decoded/executable input;
- the real Session ID, agent, message ID and outer `execute` CallID carried in
`Tool.Context`;
- handler return or throw;
- start and handler completion order.

This works for identical concurrent calls because correlation is carried by the
wrapper's own per-invocation state. It does not pair calls by input, FIFO order,
tool name, or completion order.

But the handler result is still early. After it returns, core may:

1. encode/normalize the tool output;
2. run public `tool.execute.after` hooks, which may mutate the result;
3. normalize content;
4. let the Code Mode adapter choose structured output vs text/null fallback;
5. validate Code Mode output;
6. JSON stringify/parse the value before it crosses into the confined script.

So the transform wrapper cannot claim its handler terminal is the script-visible
terminal.

### 2. Public core `tool.execute.after`

`packages/core/src/tool.ts` exposes a later hook with the core result/error.

This is also insufficient for exact Code Mode correlation:

- every inner Code Mode call receives the same `Tool.Context.id` as the outer
model-visible `execute` call;
- therefore concurrent identical inner calls have the same public CallID;
- the hook runs before Code Mode's own final output decode/JSON round trip;
- failures that escape the core Tool.Error path need not produce this hook even
though Code Mode later converts the failure for the script.

Using input equality, FIFO order, object identity, or completion order to join
this hook back to wrapper records would invent a correlation contract that stock
OpenCode does not provide.

### 3. Private Code Mode terminal and catch materialization

The last success-value boundary exists inside `@opencode/codemode`.

In `packages/codemode/src/tool-runtime.ts`, `hooked(...)` runs
`hooks["tool.after"]` from `Effect.onExit` around the Code Mode execution body.
For a successful call, this happens after output decoding and the JSON
stringify/parse round trip, so its `CallResult.value` is the plain value that the
tool promise will deliver into the interpreter.

However, `packages/core/src/codemode/tool.ts` constructs Code Mode with only
OpenCode's private `progressHooks(record)`. That hook uses the internal call
object only to update UI rows and publishes name/input/status. It does **not**
publish the success value, and stock plugin APIs provide no supported way to add
another Code Mode hook there.

The error path is later still. A failed tool promise reaches the interpreter and
`packages/codemode/src/interpreter/interpreter.ts` materializes the failure into
the JavaScript error value bound by a `catch` clause. There is no plugin/runtime
hook at that materialization boundary either. The private Code Mode `tool.after`
can see the host-side failure before this conversion, but that is not the exact
JavaScript error object seen by the script.

So stock exposes neither the final success value with public unique correlation
nor the final catch-path error representation. Those are the missing boundaries.

## PR #41 reuse decision

PR #41 correctly demonstrated the semantics needed at this boundary, but it did
so by patching OpenCode and adding an internal Code Mode observer seam. That
implementation is not reused here.

Only the behavioral test ideas are retained:

- overlapping identical calls;
- reverse completion;
- caught inner errors;
- mutation after an earlier observation point;
- real parent Session/call binding;
- script output that resembles an observer record.

The stock experiment uses only supported plugin transforms/hooks.

## Provider-free integration probe

`tests/integration/run_code_mode_observer_probe.py` runs the actual
`opencode-eval-runner:opencode-test` image built from this checkout.

It starts a loopback OpenAI-compatible fixture provider inside the container, so
there is no external provider call and no credential use. The fixture makes one
real model-visible `execute` call whose Code Mode script performs:

1. one successful inner call;
2. one thrown inner error that the script catches;
3. two concurrent calls with identical input;
4. reverse completion of those two calls;
5. multiple inner calls under the same outer `execute`;
6. a returned collector-shaped fake record.

The test plugin also changes one tool result from `BEFORE-MUTATION` to
`AFTER-MUTATION` in a public `execute.after` hook. The script asserts that it
receives `AFTER-MUTATION`. The transformed handler observer records
`BEFORE-MUTATION`. This is the concrete counterexample proving that the handler
terminal is not the caller-final value.

The outer script finally returns JSON containing
`invocation_id: "fabricated-from-script"`. The runner event stream proves that
the outer script executed, while the same ID must be absent from observer records.
Script/model output therefore does not create an inner observation.

## Evidence status

The probe plugin writes diagnostics under `/tmp` only for the test. That file is
target-writable and is **not** an evidence authority.

The experiment summary always reports:

```json
{
"status": "unsupported",
"evidence_eligible": false
}
```

This does not mean all inner facts are unavailable. It means the requested
end-to-end Code Mode result/error capability is incomplete on stock supported
surfaces, so the partial records must not be promoted as proof of the final
caller-visible value/error.

A future stock OpenCode API could make this capability supported by exposing the
internal Code Mode per-call identity and post-conversion success value to plugins,
plus the materialized catch-path error value (or one supported terminal event that
carries both forms with the same invocation identity). Until then, the correct
runner result is `unsupported`.
73 changes: 73 additions & 0 deletions docs/stock-opencode-2.0.23-observation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
# Stock OpenCode 2.0.23 observation surface

Purpose: implementation reference for the trusted-checkout evidence profile.

Source checkpoint: stock OpenCode **v2.0.23** (`0fd7e2829449b052abf0078666669302923d77af`). This is distilled from the source assessment performed in superseded PR #43 and the bounded Code Mode experiment in this branch.

OpenCode remains stock and immutable. A missing observation boundary is reported as unsupported; it is not a reason to patch OpenCode or add a hostile-runtime broker.

## Useful stock surfaces

| Observation need | Stock surface | Status |
| --- | --- | --- |
| Live runtime events | `ctx.event.subscribe()` | supported source; ordering/drain must be proven by integration test |
| Session creation / ancestry | `session.created` + Session API | supported |
| Agent for a step | Session step/message events | supported |
| Native tool call identity/input | Session tool input/called events | supported source |
| Native terminal success/failure | Session tool success/failed events | supported source |
| Tool pre-execution hook | `ctx.tool.hook("execute.before")` | supported; occurs before tool decode/execution |
| Tool post-handler hook | `ctx.tool.hook("execute.after")` | supported; occurs after core tool execution but before Code Mode final conversion |
| Tool registration wrapping | `ctx.tool.transform(...)` | supported |
| Code Mode unique inner invocation/tool/input/outer binding | transformed leaf handler + real `Tool.Context` | **supported partial boundary** |
| Code Mode exact final caller value/error | internal `@opencode/codemode` `tool.after`; not exposed to plugins | **unsupported on stock public surfaces** |

## Native calls

Stock Session events are the preferred source for native terminal facts because they represent the runtime's own Session lifecycle rather than model or tool payload claims.

The observer must bind call identity, Session, agent/message context, input and terminal result/error without reconstructing them from prose or matching by value.

## Code Mode

Code Mode executes inner tools through the normal tool registry, so same-process reviewed instrumentation can observe real inner execution without isolating Loom.

The bounded stock-2.0.23 experiment establishes a useful partial path:

- `ctx.tool.transform(...)` can wrap the actual effective leaf registration;
- core decodes input before entering that wrapper, so the wrapper sees executable input;
- the real outer `execute` `Tool.Context` reaches each inner handler, providing actual Session, agent, message and outer CallID;
- the wrapper can allocate a unique per-inner observation ID at handler entry, so identical concurrent calls and reverse completion do not require input/FIFO correlation.

However, this wrapper is not the final Code Mode caller boundary. After it returns, core can encode the result, run mutating `tool.execute.after` hooks, normalize content, and Code Mode can select its return representation and perform output validation plus a JSON stringify/parse round trip.

The later public `tool.execute.after` hook is also insufficient for exact correlation: every inner call reuses the outer `execute` `Tool.Context.id`. Concurrent identical inner calls therefore have the same public CallID.

For successful calls, the last converted value exists internally: `packages/codemode/src/tool-runtime.ts` invokes Code Mode's `tool.after` after output validation and its JSON round trip. But `packages/core/src/codemode/tool.ts` supplies only private `progressHooks(record)` there. Those hooks expose name/input/status for UI progress and do not export the value. Stock plugin APIs do not provide a supported registration point for another Code Mode hook.

Errors have an additional private step. After the tool promise fails, `packages/codemode/src/interpreter/interpreter.ts` materializes that host failure into the JavaScript error value used by a `catch` clause. No public plugin/runtime hook observes that materialized error with a unique inner invocation identity.

Therefore:

- unique inner identity, actual tool, executable input, outer binding, and start/handler-terminal ordering are observable;
- exact final value delivered to the script and exact final error seen by its catch path are **unsupported**;
- no earlier result may be promoted, paired, or reconstructed to fill that gap.

See [Code Mode inner-call observation experiment](code-mode-observer-experiment.md) for the source trace and provider-free counterexamples.

## Ordering and completeness

A monotonic observer sequence is useful, but sequence alone is not completeness. The integration must also account for starts, terminals, observer loss, process interruption, and required descendant Sessions.

Absence assertions are eligible only when the relevant scope is complete. Missing capture is never interpreted as "did not happen".

For Code Mode specifically, a complete transformed-handler trace still does not make final caller-value assertions eligible: that capability is unsupported on the stock public surface.

## What is intentionally not required

- plugin/process isolation from the trusted Loom checkout;
- remote PluginHost or capability broker;
- evidence-channel peer authentication against same-authority attackers;
- patched/forked OpenCode;
- cryptographic evidence authenticity after runtime compromise.

Those belong only to a future optional untrusted-plugin profile.
Loading
Loading