Skip to content

feat: observe stock native tool executions - #51

Closed
bateau84 wants to merge 33 commits into
mainfrom
feat/stock-native-tool-observer
Closed

bateau84 wants to merge 33 commits into
mainfrom
feat/stock-native-tool-observer

Conversation

@bateau84

@bateau84 bateau84 commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Summary

Child of #45.

Implements the smallest reviewed same-process observer for native/direct tool calls on stock OpenCode 2.0.23.

Runtime boundary

  • runner-owned Promise plugin; no OpenCode patch/fork;
  • injected through stock OPENCODE_CONFIG_CONTENT so its tool transform is applied after project/global plugin transforms;
  • wraps direct tools at the decoded tool.execute boundary to capture the actual accepted input;
  • subscribes to stock Session runtime events and records canonical session.tool.success / session.tool.failed terminals;
  • correlates only by (sessionID, messageID, callID);
  • records resolved tool, Session, agent/message, real call ID, decoded input, terminal result/error, and monotonic ordering.

The Session terminal source is intentional. The first integration run exposed that a Promise-plugin rejection can bypass tool.execute.after; stock OpenCode still settles the real invocation through session.tool.failed. The observer therefore uses the Session-owned terminal rather than inventing or reconstructing one.

Fail-closed behavior

The projection is ineligible on missing capture/end, missing terminal, sequence loss, count mismatch, duplicate/ambiguous identity, observer failure, or unavailable required fields.

Observer I/O/snapshot/event-stream failure does not alter tool behavior and does not request product retries.

Model output and tool-returned collector-shaped JSON are never parsed as observation records.

Coverage

Adds unit coverage plus a provider-free integration probe using the actual repository image with stock OpenCode 2.0.23 under --network none and a loopback fake OpenAI-compatible provider.

The probe proves:

  1. native success;
  2. native failure;
  3. accepted/executable input after an earlier pre-hook rewrite;
  4. resolved tool + Session + agent/message + real call identity;
  5. start/terminal correlation;
  6. ordering across multiple calls;
  7. collector-shaped model/tool payload rejection;
  8. no observer-induced model retry.

Verification

Current head: 026e0c52551861933776644df6744eb42d91e893

  • CI #288: passed
  • Stock native observer integration feat(runner): support explicit OCI network mode #14: passed
  • Integration probe: 17/17 checks true
  • Observed coverage: 3 starts, 3 terminals, 0 missing terminals, 0 observer failures, 0 unavailable fields
  • Runtime: stock opencode v2.0.23

Scope

Native/direct calls only. This PR does not claim Code Mode inner-call finality.

The native terminal result is the canonical Session representation:

  • success: post-truncation Session content/metadata;
  • failure: canonical Session error.

Provenance

bateau84 added a commit that referenced this pull request Oct 7, 2026
## Integrated direction

This PR now integrates the completed child work from #46–#51 into one
trusted-checkout runtime-evidence implementation on stock OpenCode
**2.0.23**.

The normal Loom eval profile remains a trusted-checkout evaluation. This
PR does **not** add an OpenCode patch/fork, remote PluginHost, protected
channel, HMAC/signing boundary, capability broker, or hostile-plugin
isolation.

## Canonical public contract

The only authoritative public runtime-evidence object is:

`opencode-eval-runner/runtime-evidence/v1`

Raw observer records are internal adapter input only.

Existing `tools`, `actions`, `tool_result_evidence`, stdout/stderr,
Session/model text, and workspace files remain convenience/diagnostic
data and are not promoted into runtime evidence. Final cleanup #52 also
removes competing runtime `evidence_eligible` signals from the
diagnostic tool-result/safety projections.

There is one canonical builder/validator that owns:

- `complete | incomplete | unsupported | invalid`;
- capture closure and unknown coverage;
- missing terminals;
- observer/callback loss;
- timeout/interruption;
- duplicate/ambiguous invocation and sequence detection;
- identity-based concurrency;
- per-boundary support;
- assertion-scoped boundary/field eligibility.

Unknown coverage is never converted to zero.

## Boundary semantics

The v1 contract exposes:

- `native`
- `code_mode_execution`
- `code_mode_finality`

Overall `complete` means the supported capture boundaries are complete.
It does **not** mean every possible assertion is supported.

This allows:

```text
overall: complete / eligible
native: complete / eligible
code_mode_execution: complete / eligible
code_mode_finality: unsupported / ineligible
```

An assertion that only requires `native` evidence can therefore remain
eligible even though Code Mode finality is unsupported. Assertions that
require a redacted/omitted/unsupported exact field remain ineligible for
that field.

## Native observation

The integrated observer preserves #51's stock-2.0.23 boundary:

```text
decoded executable input -> transformed tool.execute wrapper

terminal success/error ->
  session.tool.success
  session.tool.failed
```

It records opaque invocation identity, actual tool, agent, Session,
message, real CallID, executable input, Session ancestry, terminal
result/error, and monotonic start/terminal ordering.

Correlation is identity-based. It does not use FIFO, input equality,
tool name, or completion order.

The raw former `native_tool_observations` projection is no longer a
public result.

## Code Mode

The #50 stock-runtime result is retained faithfully.

Supported facts:

- unique per-inner invocation identity;
- actual effective tool;
- decoded/executable input;
- Session/message/agent;
- actual outer `execute` CallID;
- binding to the outer invocation;
- start/handler-terminal ordering;
- identical concurrent calls and reverse completion.

The transformed handler's earlier value/error is **not** promoted as
caller-final evidence.

Exact final script-visible value/error is represented as:

```json
{
  "state": "unsupported",
  "reason": "stock_codemode_final_boundary_not_exposed"
}
```

> Stock OpenCode 2.0.23 does not expose a supported boundary that proves
the exact final value/error seen by a Code Mode script for each inner
call. That assertion is reported as unsupported.

## Evidence safety

The #47 safety behavior is applied to authoritative dynamic values
before their first observation persistence/output sink:

```text
raw value in observer memory
  -> sanitize/redact/omit
  -> size decision
  -> internal capture
  -> validate/account
  -> result serialization
  -> stdout/host-file persistence
```

Public field states are only:

- `available`
- `redacted`
- `omitted`
- `unsupported`

Product `exit_code`, timeout, and product success/failure remain
separate from evidence eligibility.

## Child PR disposition

| PR | Disposition |
| --- | --- |
| #46 fail-closed runtime evidence accounting | **Incorporated /
superseded as a separate implementation.** Its stronger accounting is
folded into the canonical v1 builder/validator. No second
public/accounting status engine remains. |
| #47 evidence safety | **Incorporated.** Pre-sink sanitization,
redaction/omission, size ordering, credential inventory, and output
guard behavior are retained. Internal `exact` terminology is not part of
the public runtime-evidence schema. |
| #48 provider-free acceptance gate | **Incorporated and tightened.** It
now requires the exact final v1 contract rather than rollout-compatible
alternate shapes. |
| #49 runtime evidence v1 contract | **Incorporated and evolved.** It
remains the public wire contract; overall vs boundary/assertion
eligibility was revised to preserve partial support. |
| #50 stock Code Mode experiment | **Incorporated as capability
proof/diagnostic probe.** Supported execution facts are used; exact
final caller value/error remains explicitly unsupported. |
| #51 stock native observer | **Incorporated / superseded as a public
projection.** Its stock observer boundary is retained as internal
adapter input to `runtime_evidence`; no public
`native_tool_observations` object remains. |

Child PRs #46–#51 have no remaining implementation authority after this
integration. Final cleanup #52 removes parallel-work duplication without
adding architecture. Explicit unsupported areas are: Code Mode
caller-final value/error on stock 2.0.23, OpenCode runtime observation
for the `github-copilot-cli` transport, and protection against a
deliberately hostile plugin sharing the trusted OpenCode process.

## Final integration cleanup

#52 is the final child PR targeting this branch. Its head
`a3f575d78f4b14d62b76d14dceea49b76496fd2d`:

- removes duplicate runtime-evidence constants and dead `truncated`
accounting;
- removes competing `evidence_eligible` signals from diagnostic
`tool_result_evidence` / safety projections;
- removes the unused legacy tool-result clipping helper;
- confirms PR #41 protected-runtime/signing/broker machinery is absent;
- documents Code Mode finality, Copilot runtime observation, and
hostile-plugin isolation as explicit unsupported areas;
- keeps stock OpenCode 2.0.23 and normal `invoke` behavior.

Validation on #52:
- CI #297 / run 37521327070: **PASS**, 93/93 unit tests;
- provider-free runtime evidence acceptance #11 / run 37521327281:
**PASS**, 6/6 helper tests and all 10 scenarios;
- stock native observer integration #23 / run 37521327205: **PASS**.

#52 has now been squash-merged into this branch as
`b10c52d50aaf0c038e77e80eaca3890347d5fa82`.

## Final-head validation

Head: `b10c52d50aaf0c038e77e80eaca3890347d5fa82`

- **CI #296 / run 37519197332 — PASS**
  - Python/unit suite: **92/92 PASS**
  - stock Code Mode probe: **PASS**
  - exact caller terminal capability reported unsupported: **PASS**
  - OpenCode plugin activation preflight: **PASS**
  - config-root plugin dependency preflight: **PASS**
  - pinned OpenCode remains **2.0.23**
- **Provider-free runtime evidence acceptance #10 / run 37519197280 —
PASS**
  - helper/unit checks: **5/5 PASS**
  - `native_success`: PASS
  - `native_error`: PASS
  - `code_success`: PASS
  - `code_caught_error`: PASS
  - `concurrent_reverse`: PASS
  - `delegation`: PASS
  - `timeout`: PASS
  - `interrupted`: PASS
  - `redaction`: PASS
  - `collector`: PASS
- **Stock native observer integration #22 / run 37519197390 — PASS**
  - stock 2.0.23: PASS
  - complete capture: PASS
  - provider-free probe: PASS

## Review state

The integration cleanup from #52 is merged. PR #45 now contains the
final integrated implementation and remains ready for final independent
review.

PR #45 remains open/draft and is **not merged**.


## Independent final review

Final review of head `b10c52d50aaf0c038e77e80eaca3890347d5fa82` found
one blocking correctness defect: stock Code Mode's synthetic
model-facing `execute` invocation was not emitted as an authoritative
native observation, so native coverage could falsely remain
complete/zero while Code Mode had run and inner parents resolved only to
an internal token.

Corrective child PR #54 observes the real outer call at stock
`execute.before`, binds inner parents to that observed invocation, and
makes dangling/wrong-identity parents invalid. #54 is green across CI,
the stock native probe, and the 10-scenario provider-free acceptance
gate.

**Review status: NOT READY until #54 is merged.** No threat-model
expansion is required.
Base automatically changed from refactor/trusted-checkout-evidence to main October 7, 2026 07:59
@bateau84 bateau84 closed this Oct 7, 2026
@bateau84
bateau84 deleted the feat/stock-native-tool-observer branch October 7, 2026 08:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant