Skip to content

experiment: prove stock Code Mode observation boundary - #50

Closed
bateau84 wants to merge 14 commits into
mainfrom
experiment/code-mode-observer-stock-2.0.23
Closed

bateau84 wants to merge 14 commits into
mainfrom
experiment/code-mode-observer-stock-2.0.23

Conversation

@bateau84

@bateau84 bateau84 commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Parent

Child of #45. This PR is intentionally limited to the remaining Code Mode inner-call observation question.

Result

Stock OpenCode 2.0.23 supports a useful partial same-process observation path without any OpenCode patch:

  • allocate a unique per-inner invocation ID in a supported tool.transform leaf wrapper;
  • capture the actual effective tool;
  • capture decoded/executable input;
  • bind the inner call to the real outer execute Session/message/CallID carried by Tool.Context;
  • preserve start and handler-terminal ordering across identical concurrent calls and reverse completion.

It does not expose enough supported public surface to capture the exact terminal value/error seen by the Code Mode script:

  • all public inner tool hooks reuse the outer execute CallID, so concurrent identical calls cannot be uniquely correlated there;
  • the transformed handler returns before later core hooks/normalization and Code Mode output conversion;
  • the post-conversion success value is visible only to the private @opencode/codemode tool.after hook wired to OpenCode's internal progress hook;
  • a failed tool promise is materialized into the JavaScript error bound by catch later in the interpreter, with no plugin hook at that boundary.

The final caller-terminal capability is therefore reported as unsupported. No earlier value is promoted or reconstructed.

Provider-free stock-runtime proof

The actual-image integration probe uses a loopback deterministic OpenAI-compatible fixture provider and stock 2.0.23. It covers:

  1. one successful inner call;
  2. an inner thrown error caught by Code Mode;
  3. two identical concurrent calls;
  4. reverse completion order;
  5. multiple calls under one outer execute;
  6. collector-shaped outer script output that must not fabricate an inner observation.

It also mutates one result after the transformed handler returns. The Code Mode script receives the mutated value, proving the handler terminal is not caller-final.

The probe remains diagnostic only and reports:

{
  "status": "unsupported",
  "evidence_eligible": false
}

PR #41 reuse

Only the useful test shapes and counterexamples were retained. None of PR #41's patched OpenCode observer seam is reused.

Scope

  • stock OpenCode 2.0.23 only;
  • Code Mode only;
  • no fork/patch;
  • no isolation/signing architecture;
  • no changes to Loom scoring;
  • no claim that target-writable probe diagnostics are authoritative evidence.

bateau84 added a commit that referenced this pull request Oct 7, 2026
## Integrated direction

This PR now integrates the completed child work from #46–#51 into one
trusted-checkout runtime-evidence implementation on stock OpenCode
**2.0.23**.

The normal Loom eval profile remains a trusted-checkout evaluation. This
PR does **not** add an OpenCode patch/fork, remote PluginHost, protected
channel, HMAC/signing boundary, capability broker, or hostile-plugin
isolation.

## Canonical public contract

The only authoritative public runtime-evidence object is:

`opencode-eval-runner/runtime-evidence/v1`

Raw observer records are internal adapter input only.

Existing `tools`, `actions`, `tool_result_evidence`, stdout/stderr,
Session/model text, and workspace files remain convenience/diagnostic
data and are not promoted into runtime evidence. Final cleanup #52 also
removes competing runtime `evidence_eligible` signals from the
diagnostic tool-result/safety projections.

There is one canonical builder/validator that owns:

- `complete | incomplete | unsupported | invalid`;
- capture closure and unknown coverage;
- missing terminals;
- observer/callback loss;
- timeout/interruption;
- duplicate/ambiguous invocation and sequence detection;
- identity-based concurrency;
- per-boundary support;
- assertion-scoped boundary/field eligibility.

Unknown coverage is never converted to zero.

## Boundary semantics

The v1 contract exposes:

- `native`
- `code_mode_execution`
- `code_mode_finality`

Overall `complete` means the supported capture boundaries are complete.
It does **not** mean every possible assertion is supported.

This allows:

```text
overall: complete / eligible
native: complete / eligible
code_mode_execution: complete / eligible
code_mode_finality: unsupported / ineligible
```

An assertion that only requires `native` evidence can therefore remain
eligible even though Code Mode finality is unsupported. Assertions that
require a redacted/omitted/unsupported exact field remain ineligible for
that field.

## Native observation

The integrated observer preserves #51's stock-2.0.23 boundary:

```text
decoded executable input -> transformed tool.execute wrapper

terminal success/error ->
  session.tool.success
  session.tool.failed
```

It records opaque invocation identity, actual tool, agent, Session,
message, real CallID, executable input, Session ancestry, terminal
result/error, and monotonic start/terminal ordering.

Correlation is identity-based. It does not use FIFO, input equality,
tool name, or completion order.

The raw former `native_tool_observations` projection is no longer a
public result.

## Code Mode

The #50 stock-runtime result is retained faithfully.

Supported facts:

- unique per-inner invocation identity;
- actual effective tool;
- decoded/executable input;
- Session/message/agent;
- actual outer `execute` CallID;
- binding to the outer invocation;
- start/handler-terminal ordering;
- identical concurrent calls and reverse completion.

The transformed handler's earlier value/error is **not** promoted as
caller-final evidence.

Exact final script-visible value/error is represented as:

```json
{
  "state": "unsupported",
  "reason": "stock_codemode_final_boundary_not_exposed"
}
```

> Stock OpenCode 2.0.23 does not expose a supported boundary that proves
the exact final value/error seen by a Code Mode script for each inner
call. That assertion is reported as unsupported.

## Evidence safety

The #47 safety behavior is applied to authoritative dynamic values
before their first observation persistence/output sink:

```text
raw value in observer memory
  -> sanitize/redact/omit
  -> size decision
  -> internal capture
  -> validate/account
  -> result serialization
  -> stdout/host-file persistence
```

Public field states are only:

- `available`
- `redacted`
- `omitted`
- `unsupported`

Product `exit_code`, timeout, and product success/failure remain
separate from evidence eligibility.

## Child PR disposition

| PR | Disposition |
| --- | --- |
| #46 fail-closed runtime evidence accounting | **Incorporated /
superseded as a separate implementation.** Its stronger accounting is
folded into the canonical v1 builder/validator. No second
public/accounting status engine remains. |
| #47 evidence safety | **Incorporated.** Pre-sink sanitization,
redaction/omission, size ordering, credential inventory, and output
guard behavior are retained. Internal `exact` terminology is not part of
the public runtime-evidence schema. |
| #48 provider-free acceptance gate | **Incorporated and tightened.** It
now requires the exact final v1 contract rather than rollout-compatible
alternate shapes. |
| #49 runtime evidence v1 contract | **Incorporated and evolved.** It
remains the public wire contract; overall vs boundary/assertion
eligibility was revised to preserve partial support. |
| #50 stock Code Mode experiment | **Incorporated as capability
proof/diagnostic probe.** Supported execution facts are used; exact
final caller value/error remains explicitly unsupported. |
| #51 stock native observer | **Incorporated / superseded as a public
projection.** Its stock observer boundary is retained as internal
adapter input to `runtime_evidence`; no public
`native_tool_observations` object remains. |

Child PRs #46–#51 have no remaining implementation authority after this
integration. Final cleanup #52 removes parallel-work duplication without
adding architecture. Explicit unsupported areas are: Code Mode
caller-final value/error on stock 2.0.23, OpenCode runtime observation
for the `github-copilot-cli` transport, and protection against a
deliberately hostile plugin sharing the trusted OpenCode process.

## Final integration cleanup

#52 is the final child PR targeting this branch. Its head
`a3f575d78f4b14d62b76d14dceea49b76496fd2d`:

- removes duplicate runtime-evidence constants and dead `truncated`
accounting;
- removes competing `evidence_eligible` signals from diagnostic
`tool_result_evidence` / safety projections;
- removes the unused legacy tool-result clipping helper;
- confirms PR #41 protected-runtime/signing/broker machinery is absent;
- documents Code Mode finality, Copilot runtime observation, and
hostile-plugin isolation as explicit unsupported areas;
- keeps stock OpenCode 2.0.23 and normal `invoke` behavior.

Validation on #52:
- CI #297 / run 37521327070: **PASS**, 93/93 unit tests;
- provider-free runtime evidence acceptance #11 / run 37521327281:
**PASS**, 6/6 helper tests and all 10 scenarios;
- stock native observer integration #23 / run 37521327205: **PASS**.

#52 has now been squash-merged into this branch as
`b10c52d50aaf0c038e77e80eaca3890347d5fa82`.

## Final-head validation

Head: `b10c52d50aaf0c038e77e80eaca3890347d5fa82`

- **CI #296 / run 37519197332 — PASS**
  - Python/unit suite: **92/92 PASS**
  - stock Code Mode probe: **PASS**
  - exact caller terminal capability reported unsupported: **PASS**
  - OpenCode plugin activation preflight: **PASS**
  - config-root plugin dependency preflight: **PASS**
  - pinned OpenCode remains **2.0.23**
- **Provider-free runtime evidence acceptance #10 / run 37519197280 —
PASS**
  - helper/unit checks: **5/5 PASS**
  - `native_success`: PASS
  - `native_error`: PASS
  - `code_success`: PASS
  - `code_caught_error`: PASS
  - `concurrent_reverse`: PASS
  - `delegation`: PASS
  - `timeout`: PASS
  - `interrupted`: PASS
  - `redaction`: PASS
  - `collector`: PASS
- **Stock native observer integration #22 / run 37519197390 — PASS**
  - stock 2.0.23: PASS
  - complete capture: PASS
  - provider-free probe: PASS

## Review state

The integration cleanup from #52 is merged. PR #45 now contains the
final integrated implementation and remains ready for final independent
review.

PR #45 remains open/draft and is **not merged**.


## Independent final review

Final review of head `b10c52d50aaf0c038e77e80eaca3890347d5fa82` found
one blocking correctness defect: stock Code Mode's synthetic
model-facing `execute` invocation was not emitted as an authoritative
native observation, so native coverage could falsely remain
complete/zero while Code Mode had run and inner parents resolved only to
an internal token.

Corrective child PR #54 observes the real outer call at stock
`execute.before`, binds inner parents to that observed invocation, and
makes dangling/wrong-identity parents invalid. #54 is green across CI,
the stock native probe, and the 10-scenario provider-free acceptance
gate.

**Review status: NOT READY until #54 is merged.** No threat-model
expansion is required.
Base automatically changed from refactor/trusted-checkout-evidence to main October 7, 2026 07:59
@bateau84 bateau84 closed this Oct 7, 2026
@bateau84
bateau84 deleted the experiment/code-mode-observer-stock-2.0.23 branch October 7, 2026 08:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant