Skip to content

refactor: tighten integrated runtime evidence authority - #52

Merged
bateau84 merged 9 commits into
refactor/trusted-checkout-evidencefrom
refactor/trusted-checkout-evidence-integration-pass
Oct 6, 2026
Merged

bateau84 merged 9 commits into
refactor/trusted-checkout-evidencefrom
refactor/trusted-checkout-evidence-integration-pass

Conversation

@bateau84

@bateau84 bateau84 commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Purpose

Final integration pass over PR #45 after the parallel runtime-evidence child work.

Findings and changes

  • opencode-eval-runner/runtime-evidence/v1 remains the only runtime-evidence authority.
  • Removed duplicate runtime-evidence constants and dead truncated accounting left by parallel work.
  • Removed evidence_eligible from tool_result_evidence and the evidence-safety diagnostic projection so convenience/safety surfaces cannot look authoritative.
  • Removed the unused legacy tool-result clipping helper.
  • Added unit and provider-free acceptance coverage that locks the single-authority rule.
  • Documented explicit unsupported areas: Code Mode caller finality, Copilot runtime observation, and hostile-plugin isolation.
  • Retained stock OpenCode 2.0.23 and the existing normal invoke convenience fields/flow.

PR #41 residue audit

The integrated tree contains none of PR #41's protected-runtime, patched-runtime, signing/Cosign, protected-channel, disposable-state, safe_invoke, or capability-broker machinery. No such security architecture was reintroduced. The dead remnants actually present were limited to the duplicate/dead accounting state and unused clipping helper removed here.

Architecture

No new architecture is introduced. Native and Code Mode observations continue through the same canonical builder/validator and public field-state contract. Dynamic evidence values are sanitized before the internal observation sink and are sanitized again before result export.

Explicit unsupported areas

  • Exact final success value/error seen by a Code Mode script remains unsupported on stock OpenCode 2.0.23.
  • github-copilot-cli has no OpenCode runtime observer; its canonical runtime_evidence status is unsupported.
  • The trusted-checkout profile does not defend same-process instrumentation from a deliberately hostile evaluated plugin.

Validation

Head: a3f575d78f4b14d62b76d14dceea49b76496fd2d

  • CI #297 / run 37521327070 — PASS
    • full Python/unit suite: 93/93 PASS
    • remaining image/integration checks: PASS
  • Provider-free runtime evidence acceptance Export tool call arguments in normalized eval results #11 / run 37521327281 — PASS
    • acceptance helper tests: 6/6 PASS
    • native_success, native_error, code_success, code_caught_error, concurrent_reverse, delegation, timeout, interrupted, redaction, and collector: all PASS
    • stock runtime confirmed as OpenCode 2.0.23
  • Stock native observer integration fix: resolve eval agents with opencode debug agents #23 / run 37521327205 — PASS

PR #52 targets refactor/trusted-checkout-evidence and is ready for review/merge.

@bateau84
bateau84 merged commit b10c52d into refactor/trusted-checkout-evidence Oct 6, 2026
3 checks passed
@bateau84
bateau84 deleted the refactor/trusted-checkout-evidence-integration-pass branch October 6, 2026 20:13
bateau84 added a commit that referenced this pull request Oct 6, 2026
## Purpose

Correct the blocking outer-`execute` evidence gap found during the
independent final review of #45.

This branch is based on the post-#52 target head
`b10c52d50aaf0c038e77e80eaca3890347d5fa82`.

## Defect

Stock OpenCode 2.0.23 creates the model-facing Code Mode `execute` tool
inside `Tool.snapshot`, after registration transforms have run. The
integrated observer therefore did not emit a native observation for that
real outer invocation. Inner records were linked only to an internal
opaque token.

That allowed the public native boundary to report complete/zero
outer-`execute` observations while Code Mode had actually run.

## Correction

- Observe the real model-facing outer `execute` invocation at stock
`execute.before`.
- Keep `session.tool.success` / `session.tool.failed` as terminal
authority.
- Bind Code Mode inner records to that observed outer invocation ID.
- Reject dangling, wrong-tool, or identity-mismatched parents in
canonical accounting/validation.
- Strengthen provider-free acceptance checks for outer observation,
exact parent resolution, identity agreement, non-zero native coverage,
concurrency, and reverse completion.
- Document that synthetic outer `execute` input is exact at
`execute.before` but pre-`CodeMode.Input` decode.

## Scope

No threat-model expansion. No OpenCode patch/fork, protected channel,
signing, broker, process isolation, or hostile-plugin machinery.

Targets `refactor/trusted-checkout-evidence`.

## Validation

Head: `221af580640360c16e263cc05136d3be0d090a0f`

- CI #307 / run 37525849151: **PASS**, 94/94 Python tests; stock Code
Mode diagnostic probe PASS; OpenCode **2.0.23**.
- Provider-free runtime evidence acceptance #21 / run 37525849081:
**PASS**, 6/6 helper tests and all 10 scenarios.
- Stock native observer integration #33 / run 37525849107: **PASS**.
- PR is mergeable against the current #45 branch.

#53 is closed as superseded; it became dirty only because #52 merged
while this review was in progress.
bateau84 added a commit that referenced this pull request Oct 7, 2026
## Integrated direction

This PR now integrates the completed child work from #46–#51 into one
trusted-checkout runtime-evidence implementation on stock OpenCode
**2.0.23**.

The normal Loom eval profile remains a trusted-checkout evaluation. This
PR does **not** add an OpenCode patch/fork, remote PluginHost, protected
channel, HMAC/signing boundary, capability broker, or hostile-plugin
isolation.

## Canonical public contract

The only authoritative public runtime-evidence object is:

`opencode-eval-runner/runtime-evidence/v1`

Raw observer records are internal adapter input only.

Existing `tools`, `actions`, `tool_result_evidence`, stdout/stderr,
Session/model text, and workspace files remain convenience/diagnostic
data and are not promoted into runtime evidence. Final cleanup #52 also
removes competing runtime `evidence_eligible` signals from the
diagnostic tool-result/safety projections.

There is one canonical builder/validator that owns:

- `complete | incomplete | unsupported | invalid`;
- capture closure and unknown coverage;
- missing terminals;
- observer/callback loss;
- timeout/interruption;
- duplicate/ambiguous invocation and sequence detection;
- identity-based concurrency;
- per-boundary support;
- assertion-scoped boundary/field eligibility.

Unknown coverage is never converted to zero.

## Boundary semantics

The v1 contract exposes:

- `native`
- `code_mode_execution`
- `code_mode_finality`

Overall `complete` means the supported capture boundaries are complete.
It does **not** mean every possible assertion is supported.

This allows:

```text
overall: complete / eligible
native: complete / eligible
code_mode_execution: complete / eligible
code_mode_finality: unsupported / ineligible
```

An assertion that only requires `native` evidence can therefore remain
eligible even though Code Mode finality is unsupported. Assertions that
require a redacted/omitted/unsupported exact field remain ineligible for
that field.

## Native observation

The integrated observer preserves #51's stock-2.0.23 boundary:

```text
decoded executable input -> transformed tool.execute wrapper

terminal success/error ->
  session.tool.success
  session.tool.failed
```

It records opaque invocation identity, actual tool, agent, Session,
message, real CallID, executable input, Session ancestry, terminal
result/error, and monotonic start/terminal ordering.

Correlation is identity-based. It does not use FIFO, input equality,
tool name, or completion order.

The raw former `native_tool_observations` projection is no longer a
public result.

## Code Mode

The #50 stock-runtime result is retained faithfully.

Supported facts:

- unique per-inner invocation identity;
- actual effective tool;
- decoded/executable input;
- Session/message/agent;
- actual outer `execute` CallID;
- binding to the outer invocation;
- start/handler-terminal ordering;
- identical concurrent calls and reverse completion.

The transformed handler's earlier value/error is **not** promoted as
caller-final evidence.

Exact final script-visible value/error is represented as:

```json
{
  "state": "unsupported",
  "reason": "stock_codemode_final_boundary_not_exposed"
}
```

> Stock OpenCode 2.0.23 does not expose a supported boundary that proves
the exact final value/error seen by a Code Mode script for each inner
call. That assertion is reported as unsupported.

## Evidence safety

The #47 safety behavior is applied to authoritative dynamic values
before their first observation persistence/output sink:

```text
raw value in observer memory
  -> sanitize/redact/omit
  -> size decision
  -> internal capture
  -> validate/account
  -> result serialization
  -> stdout/host-file persistence
```

Public field states are only:

- `available`
- `redacted`
- `omitted`
- `unsupported`

Product `exit_code`, timeout, and product success/failure remain
separate from evidence eligibility.

## Child PR disposition

| PR | Disposition |
| --- | --- |
| #46 fail-closed runtime evidence accounting | **Incorporated /
superseded as a separate implementation.** Its stronger accounting is
folded into the canonical v1 builder/validator. No second
public/accounting status engine remains. |
| #47 evidence safety | **Incorporated.** Pre-sink sanitization,
redaction/omission, size ordering, credential inventory, and output
guard behavior are retained. Internal `exact` terminology is not part of
the public runtime-evidence schema. |
| #48 provider-free acceptance gate | **Incorporated and tightened.** It
now requires the exact final v1 contract rather than rollout-compatible
alternate shapes. |
| #49 runtime evidence v1 contract | **Incorporated and evolved.** It
remains the public wire contract; overall vs boundary/assertion
eligibility was revised to preserve partial support. |
| #50 stock Code Mode experiment | **Incorporated as capability
proof/diagnostic probe.** Supported execution facts are used; exact
final caller value/error remains explicitly unsupported. |
| #51 stock native observer | **Incorporated / superseded as a public
projection.** Its stock observer boundary is retained as internal
adapter input to `runtime_evidence`; no public
`native_tool_observations` object remains. |

Child PRs #46–#51 have no remaining implementation authority after this
integration. Final cleanup #52 removes parallel-work duplication without
adding architecture. Explicit unsupported areas are: Code Mode
caller-final value/error on stock 2.0.23, OpenCode runtime observation
for the `github-copilot-cli` transport, and protection against a
deliberately hostile plugin sharing the trusted OpenCode process.

## Final integration cleanup

#52 is the final child PR targeting this branch. Its head
`a3f575d78f4b14d62b76d14dceea49b76496fd2d`:

- removes duplicate runtime-evidence constants and dead `truncated`
accounting;
- removes competing `evidence_eligible` signals from diagnostic
`tool_result_evidence` / safety projections;
- removes the unused legacy tool-result clipping helper;
- confirms PR #41 protected-runtime/signing/broker machinery is absent;
- documents Code Mode finality, Copilot runtime observation, and
hostile-plugin isolation as explicit unsupported areas;
- keeps stock OpenCode 2.0.23 and normal `invoke` behavior.

Validation on #52:
- CI #297 / run 37521327070: **PASS**, 93/93 unit tests;
- provider-free runtime evidence acceptance #11 / run 37521327281:
**PASS**, 6/6 helper tests and all 10 scenarios;
- stock native observer integration #23 / run 37521327205: **PASS**.

#52 has now been squash-merged into this branch as
`b10c52d50aaf0c038e77e80eaca3890347d5fa82`.

## Final-head validation

Head: `b10c52d50aaf0c038e77e80eaca3890347d5fa82`

- **CI #296 / run 37519197332 — PASS**
  - Python/unit suite: **92/92 PASS**
  - stock Code Mode probe: **PASS**
  - exact caller terminal capability reported unsupported: **PASS**
  - OpenCode plugin activation preflight: **PASS**
  - config-root plugin dependency preflight: **PASS**
  - pinned OpenCode remains **2.0.23**
- **Provider-free runtime evidence acceptance #10 / run 37519197280 —
PASS**
  - helper/unit checks: **5/5 PASS**
  - `native_success`: PASS
  - `native_error`: PASS
  - `code_success`: PASS
  - `code_caught_error`: PASS
  - `concurrent_reverse`: PASS
  - `delegation`: PASS
  - `timeout`: PASS
  - `interrupted`: PASS
  - `redaction`: PASS
  - `collector`: PASS
- **Stock native observer integration #22 / run 37519197390 — PASS**
  - stock 2.0.23: PASS
  - complete capture: PASS
  - provider-free probe: PASS

## Review state

The integration cleanup from #52 is merged. PR #45 now contains the
final integrated implementation and remains ready for final independent
review.

PR #45 remains open/draft and is **not merged**.


## Independent final review

Final review of head `b10c52d50aaf0c038e77e80eaca3890347d5fa82` found
one blocking correctness defect: stock Code Mode's synthetic
model-facing `execute` invocation was not emitted as an authoritative
native observation, so native coverage could falsely remain
complete/zero while Code Mode had run and inner parents resolved only to
an internal token.

Corrective child PR #54 observes the real outer call at stock
`execute.before`, binds inner parents to that observed invocation, and
makes dangling/wrong-identity parents invalid. #54 is green across CI,
the stock native probe, and the 10-scenario provider-free acceptance
gate.

**Review status: NOT READY until #54 is merged.** No threat-model
expansion is required.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant