Skip to content

feat: add fail-closed runtime evidence accounting - #46

Closed
bateau84 wants to merge 8 commits into
mainfrom
feat/runtime-evidence-accounting
Closed

bateau84 wants to merge 8 commits into
mainfrom
feat/runtime-evidence-accounting

Conversation

@bateau84

@bateau84 bateau84 commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Purpose

Child PR of #45 for Prompt 5: explicit completeness/loss accounting for trusted runtime evidence.

What this adds

  • observer-agnostic normalized start/terminal accounting;
  • explicit complete | incomplete | unsupported | invalid status;
  • counts for starts, terminals, missing terminals, observer failures, callback failures, losses, malformed observations, field loss, and identity ambiguity;
  • an explicit observation_closed correctness gate so missing observations cannot be treated as proof of absence before the adapter closes its scope;
  • per-boundary status plus assertion_status(...), allowing an unsupported Code Mode boundary to remain local to assertions that need it;
  • identity-based correlation for concurrent/reverse-completion calls;
  • timeout/interruption fail-closed behavior;
  • required omitted/truncated fields => incomplete, required unsupported fields => unsupported;
  • product outcome kept separate from evidence completeness.

Scope

This deliberately does not add generation sealing, authenticated channels, hostile-runtime isolation, or observer-specific capture code.

The native and Code Mode observer child PRs can adapt their actual runtime events into this small normalized accounting input.

Tests

Added 18 unit tests covering balanced capture, absence/closure, missing terminals, observer/callback loss, malformed input, unsupported boundaries, timeout/interruption, required field loss, duplicate/ambiguous identity, duplicate sequence, concurrent reverse completion, and product-failure separation.

Locally exercised with:

PYTHONPATH=. python3 -m unittest discover -s tests -p 'test_*.py' -v

for the new isolated test module: 18/18 passed.

bateau84 added a commit that referenced this pull request Oct 7, 2026
## Integrated direction

This PR now integrates the completed child work from #46–#51 into one
trusted-checkout runtime-evidence implementation on stock OpenCode
**2.0.23**.

The normal Loom eval profile remains a trusted-checkout evaluation. This
PR does **not** add an OpenCode patch/fork, remote PluginHost, protected
channel, HMAC/signing boundary, capability broker, or hostile-plugin
isolation.

## Canonical public contract

The only authoritative public runtime-evidence object is:

`opencode-eval-runner/runtime-evidence/v1`

Raw observer records are internal adapter input only.

Existing `tools`, `actions`, `tool_result_evidence`, stdout/stderr,
Session/model text, and workspace files remain convenience/diagnostic
data and are not promoted into runtime evidence. Final cleanup #52 also
removes competing runtime `evidence_eligible` signals from the
diagnostic tool-result/safety projections.

There is one canonical builder/validator that owns:

- `complete | incomplete | unsupported | invalid`;
- capture closure and unknown coverage;
- missing terminals;
- observer/callback loss;
- timeout/interruption;
- duplicate/ambiguous invocation and sequence detection;
- identity-based concurrency;
- per-boundary support;
- assertion-scoped boundary/field eligibility.

Unknown coverage is never converted to zero.

## Boundary semantics

The v1 contract exposes:

- `native`
- `code_mode_execution`
- `code_mode_finality`

Overall `complete` means the supported capture boundaries are complete.
It does **not** mean every possible assertion is supported.

This allows:

```text
overall: complete / eligible
native: complete / eligible
code_mode_execution: complete / eligible
code_mode_finality: unsupported / ineligible
```

An assertion that only requires `native` evidence can therefore remain
eligible even though Code Mode finality is unsupported. Assertions that
require a redacted/omitted/unsupported exact field remain ineligible for
that field.

## Native observation

The integrated observer preserves #51's stock-2.0.23 boundary:

```text
decoded executable input -> transformed tool.execute wrapper

terminal success/error ->
  session.tool.success
  session.tool.failed
```

It records opaque invocation identity, actual tool, agent, Session,
message, real CallID, executable input, Session ancestry, terminal
result/error, and monotonic start/terminal ordering.

Correlation is identity-based. It does not use FIFO, input equality,
tool name, or completion order.

The raw former `native_tool_observations` projection is no longer a
public result.

## Code Mode

The #50 stock-runtime result is retained faithfully.

Supported facts:

- unique per-inner invocation identity;
- actual effective tool;
- decoded/executable input;
- Session/message/agent;
- actual outer `execute` CallID;
- binding to the outer invocation;
- start/handler-terminal ordering;
- identical concurrent calls and reverse completion.

The transformed handler's earlier value/error is **not** promoted as
caller-final evidence.

Exact final script-visible value/error is represented as:

```json
{
  "state": "unsupported",
  "reason": "stock_codemode_final_boundary_not_exposed"
}
```

> Stock OpenCode 2.0.23 does not expose a supported boundary that proves
the exact final value/error seen by a Code Mode script for each inner
call. That assertion is reported as unsupported.

## Evidence safety

The #47 safety behavior is applied to authoritative dynamic values
before their first observation persistence/output sink:

```text
raw value in observer memory
  -> sanitize/redact/omit
  -> size decision
  -> internal capture
  -> validate/account
  -> result serialization
  -> stdout/host-file persistence
```

Public field states are only:

- `available`
- `redacted`
- `omitted`
- `unsupported`

Product `exit_code`, timeout, and product success/failure remain
separate from evidence eligibility.

## Child PR disposition

| PR | Disposition |
| --- | --- |
| #46 fail-closed runtime evidence accounting | **Incorporated /
superseded as a separate implementation.** Its stronger accounting is
folded into the canonical v1 builder/validator. No second
public/accounting status engine remains. |
| #47 evidence safety | **Incorporated.** Pre-sink sanitization,
redaction/omission, size ordering, credential inventory, and output
guard behavior are retained. Internal `exact` terminology is not part of
the public runtime-evidence schema. |
| #48 provider-free acceptance gate | **Incorporated and tightened.** It
now requires the exact final v1 contract rather than rollout-compatible
alternate shapes. |
| #49 runtime evidence v1 contract | **Incorporated and evolved.** It
remains the public wire contract; overall vs boundary/assertion
eligibility was revised to preserve partial support. |
| #50 stock Code Mode experiment | **Incorporated as capability
proof/diagnostic probe.** Supported execution facts are used; exact
final caller value/error remains explicitly unsupported. |
| #51 stock native observer | **Incorporated / superseded as a public
projection.** Its stock observer boundary is retained as internal
adapter input to `runtime_evidence`; no public
`native_tool_observations` object remains. |

Child PRs #46–#51 have no remaining implementation authority after this
integration. Final cleanup #52 removes parallel-work duplication without
adding architecture. Explicit unsupported areas are: Code Mode
caller-final value/error on stock 2.0.23, OpenCode runtime observation
for the `github-copilot-cli` transport, and protection against a
deliberately hostile plugin sharing the trusted OpenCode process.

## Final integration cleanup

#52 is the final child PR targeting this branch. Its head
`a3f575d78f4b14d62b76d14dceea49b76496fd2d`:

- removes duplicate runtime-evidence constants and dead `truncated`
accounting;
- removes competing `evidence_eligible` signals from diagnostic
`tool_result_evidence` / safety projections;
- removes the unused legacy tool-result clipping helper;
- confirms PR #41 protected-runtime/signing/broker machinery is absent;
- documents Code Mode finality, Copilot runtime observation, and
hostile-plugin isolation as explicit unsupported areas;
- keeps stock OpenCode 2.0.23 and normal `invoke` behavior.

Validation on #52:
- CI #297 / run 37521327070: **PASS**, 93/93 unit tests;
- provider-free runtime evidence acceptance #11 / run 37521327281:
**PASS**, 6/6 helper tests and all 10 scenarios;
- stock native observer integration #23 / run 37521327205: **PASS**.

#52 has now been squash-merged into this branch as
`b10c52d50aaf0c038e77e80eaca3890347d5fa82`.

## Final-head validation

Head: `b10c52d50aaf0c038e77e80eaca3890347d5fa82`

- **CI #296 / run 37519197332 — PASS**
  - Python/unit suite: **92/92 PASS**
  - stock Code Mode probe: **PASS**
  - exact caller terminal capability reported unsupported: **PASS**
  - OpenCode plugin activation preflight: **PASS**
  - config-root plugin dependency preflight: **PASS**
  - pinned OpenCode remains **2.0.23**
- **Provider-free runtime evidence acceptance #10 / run 37519197280 —
PASS**
  - helper/unit checks: **5/5 PASS**
  - `native_success`: PASS
  - `native_error`: PASS
  - `code_success`: PASS
  - `code_caught_error`: PASS
  - `concurrent_reverse`: PASS
  - `delegation`: PASS
  - `timeout`: PASS
  - `interrupted`: PASS
  - `redaction`: PASS
  - `collector`: PASS
- **Stock native observer integration #22 / run 37519197390 — PASS**
  - stock 2.0.23: PASS
  - complete capture: PASS
  - provider-free probe: PASS

## Review state

The integration cleanup from #52 is merged. PR #45 now contains the
final integrated implementation and remains ready for final independent
review.

PR #45 remains open/draft and is **not merged**.


## Independent final review

Final review of head `b10c52d50aaf0c038e77e80eaca3890347d5fa82` found
one blocking correctness defect: stock Code Mode's synthetic
model-facing `execute` invocation was not emitted as an authoritative
native observation, so native coverage could falsely remain
complete/zero while Code Mode had run and inner parents resolved only to
an internal token.

Corrective child PR #54 observes the real outer call at stock
`execute.before`, binds inner parents to that observed invocation, and
makes dangling/wrong-identity parents invalid. #54 is green across CI,
the stock native probe, and the 10-scenario provider-free acceptance
gate.

**Review status: NOT READY until #54 is merged.** No threat-model
expansion is required.
Base automatically changed from refactor/trusted-checkout-evidence to main October 7, 2026 07:59
@bateau84 bateau84 closed this Oct 7, 2026
@bateau84
bateau84 deleted the feat/runtime-evidence-accounting branch October 7, 2026 08:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant