Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
42e7253
fix(evidence): observe outer Code Mode execute call
bateau84 Oct 6, 2026
c1b91ad
test(evidence): accept outer execute observation boundary
bateau84 Oct 6, 2026
1eaa116
test(evidence): accept outer execute observation boundary
bateau84 Oct 6, 2026
4ef34b4
fix(evidence): require observed outer execute parent
bateau84 Oct 6, 2026
f6b69a7
test(evidence): reject unobserved Code Mode parents
bateau84 Oct 6, 2026
03e63cf
test(evidence): model observed outer execute in helper capture
bateau84 Oct 6, 2026
0b4b960
test(evidence): gate on observed outer execute parent
bateau84 Oct 6, 2026
4278b79
docs(evidence): document observed outer execute boundary
bateau84 Oct 6, 2026
bce3ea1
docs(evidence): make outer execute semantics explicit
bateau84 Oct 6, 2026
8963ae9
docs(evidence): document stock outer execute observation
bateau84 Oct 6, 2026
e2e7d12
docs(evidence): tie inner records to observed outer call
bateau84 Oct 6, 2026
4c69dbc
docs(evidence): fix outer execute boundary wording
bateau84 Oct 6, 2026
fba55b3
docs(evidence): fix native and parent boundary wording
bateau84 Oct 6, 2026
0fe41f0
docs(evidence): fix observer paragraph formatting
bateau84 Oct 6, 2026
8d40db6
docs(evidence): document synthetic execute start surface
bateau84 Oct 6, 2026
5f98787
chore(evidence): integrate final authority cleanup
bateau84 Oct 6, 2026
328d75a
chore(evidence): integrate final authority cleanup
bateau84 Oct 6, 2026
7961055
chore(evidence): integrate final authority cleanup
bateau84 Oct 6, 2026
c92a66b
chore(evidence): integrate final authority cleanup
bateau84 Oct 6, 2026
77c5a84
chore(evidence): integrate final authority cleanup
bateau84 Oct 6, 2026
1f75b5e
refactor(evidence): keep one canonical accounting model
bateau84 Oct 6, 2026
4fc7a32
docs(evidence): clarify sole authority and unsupported areas
bateau84 Oct 6, 2026
31b5946
test(evidence): enforce single authority surface
bateau84 Oct 6, 2026
463279b
test(evidence): reject competing eligibility surfaces
bateau84 Oct 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -352,10 +352,12 @@ This profile does **not** claim resistance to an evaluated plugin that deliberat

See [Trusted-checkout runtime evidence](docs/trusted-checkout-evidence.md) and the [versioned runtime-evidence result contract](docs/runtime-evidence-contract.md).

OpenCode results now expose `opencode-eval-runner/runtime-evidence/v1` as the single authoritative runtime-evidence object. Native calls are observed through the stock-2.0.23 decoded-execution and Session terminal boundaries. Code Mode inner identity/input/ordering is observable, while exact final script-visible value/error remains explicitly `unsupported`. Existing `tools`, `actions`, `tool_result_evidence`, stdout/stderr, and model text remain convenience/diagnostic data only.
OpenCode results now expose `opencode-eval-runner/runtime-evidence/v1` as the single authoritative runtime-evidence object. Native calls are observed through the stock-2.0.23 decoded-execution and Session terminal boundaries. Code Mode inner identity/input/ordering is observable, while exact final script-visible value/error remains explicitly `unsupported`. Existing `tools`, `actions`, `tool_result_evidence`, stdout/stderr, and model text remain convenience/diagnostic data only; they do not expose runtime-evidence eligibility and are never substitutes for `runtime_evidence`.

Overall evidence eligibility is separate from assertion eligibility: an unsupported Code Mode finality boundary does not invalidate an unrelated complete native assertion, and redacted/omitted fields only block assertions that require those exact values.

The `github-copilot-cli` transport has no OpenCode runtime observer. It still emits the canonical `runtime_evidence` object, but with status `unsupported`.

> Stock OpenCode 2.0.23 does not expose a supported boundary that proves the exact final value/error seen by a Code Mode script for each inner call. That assertion is reported as unsupported.

## Security boundary
Expand Down
1 change: 0 additions & 1 deletion container/evidence_safety.py
Original file line number Diff line number Diff line change
Expand Up @@ -400,7 +400,6 @@ def summary(self) -> dict[str, Any]:
return {
"schema": SCHEMA,
"inventory_complete": self.sanitizer.inventory_complete,
"evidence_eligible": self.sanitizer.inventory_complete and not losses,
"fields": self.fields,
"losses": losses,
}
12 changes: 0 additions & 12 deletions container/invoke.py
Original file line number Diff line number Diff line change
Expand Up @@ -167,16 +167,6 @@ def extract_actions(events: list[dict[str, Any]]) -> list[dict[str, Any]]:
STDERR_CAPTURE_LIMIT = 20000


def _tool_result_text(value: Any, limit: int) -> tuple[str, bool]:
text = value if isinstance(value, str) else json.dumps(value, ensure_ascii=False, sort_keys=True)
if len(text) <= limit:
return text, False
marker = "\n[... tool-result field truncated ...]\n"
retained = limit - len(marker)
head = retained // 2
return text[:head] + marker + text[-(retained - head):], True


def extract_tool_result_evidence(
events: list[dict[str, Any]],
sanitizer: Sanitizer | None = None,
Expand Down Expand Up @@ -316,7 +306,6 @@ def drop_event_fields(sequence: int) -> None:

evidence["events"] = recent
evidence["safety"] = projection.summary()
evidence["evidence_eligible"] = evidence["safety"]["evidence_eligible"]

while (
len(json.dumps(evidence, ensure_ascii=False, separators=(",", ":")))
Expand All @@ -328,7 +317,6 @@ def drop_event_fields(sequence: int) -> None:
evidence["omitted_events"] += 1
projection.loss("size_limit")
evidence["safety"] = projection.summary()
evidence["evidence_eligible"] = False

return evidence

Expand Down
2 changes: 1 addition & 1 deletion container/native_observer.py
Original file line number Diff line number Diff line change
Expand Up @@ -132,7 +132,7 @@ def load_runtime_observations(path: Path = OBSERVATION_PATH) -> dict[str, Any]:
"kind": "capture_start",
"version": 1,
"source": "stock-opencode-2.0.23-plugin",
"native_input_boundary": "decoded-tool-execute",
"native_input_boundary": "decoded-tool-execute+outer-execute-before",
"native_terminal_boundary": "session.tool.success+session.tool.failed",
"code_input_boundary": "decoded-code-tool-handler",
"code_terminal_boundary": "tool-handler-return+tool-handler-throw",
Expand Down
60 changes: 44 additions & 16 deletions container/native_observer.ts
Original file line number Diff line number Diff line change
Expand Up @@ -258,7 +258,7 @@ export default {
kind: "capture_start",
version: 1,
source: "stock-opencode-2.0.23-plugin",
native_input_boundary: "decoded-tool-execute",
native_input_boundary: "decoded-tool-execute+outer-execute-before",
native_terminal_boundary: "session.tool.success+session.tool.failed",
code_input_boundary: "decoded-code-tool-handler",
code_terminal_boundary: "tool-handler-return+tool-handler-throw",
Expand Down Expand Up @@ -395,25 +395,53 @@ export default {
}
})

// Public hooks are used only to bind inner calls to the actual outer
// execute CallID. They are not Code Mode final-result evidence.
await ctx.tool.hook("execute.before", (event: any) => {
// The synthetic Code Mode execute tool is created inside Tool.snapshot,
// after registration transforms have run. Observe that real model-facing
// invocation at the stock execute.before runtime hook so it cannot vanish
// from the native boundary. This is an exact observed effective input, but
// unlike transformed registered tools it is before CodeMode.Input decode.
await ctx.tool.hook("execute.before", async (event: any) => {
if (event?.tool !== "execute") return
if (
typeof event.sessionID === "string" &&
typeof event.messageID === "string" &&
typeof event.id === "string"
typeof event.sessionID !== "string" ||
typeof event.agent !== "string" ||
typeof event.messageID !== "string" ||
typeof event.id !== "string"
) {
const key = identityKey({
sessionID: event.sessionID,
messageID: event.messageID,
callID: event.id,
})
activeOuter.set(
key,
nativeActive.get(key)?.invocationID ?? "native-outer-unobserved:" + randomUUID(),
)
callbackFailures += 1
return
}

const identity: Identity = {
invocationID: nativeInvocationID(),
tool: "execute",
sessionID: event.sessionID,
agent: event.agent,
messageID: event.messageID,
callID: event.id,
}
const key = identityKey(identity)
const existing = nativeActive.get(key)
if (existing) {
activeOuter.set(key, existing.invocationID)
return
}

nativeStarts += 1
nativeActive.set(key, identity)
activeOuter.set(key, identity.invocationID)
write({
kind: "native_start",
invocation_id: identity.invocationID,
tool: project(identity.tool),
session_id: project(identity.sessionID),
agent: project(identity.agent),
message_id: project(identity.messageID),
call_id: project(identity.callID),
parent_session_id: await parentSession(ctx, identity.sessionID),
input: project(event.input),
boundary: "tool-execute-before",
})
})
await ctx.tool.hook("execute.after", (event: any) => {
if (event?.tool !== "execute") return
Expand Down
89 changes: 73 additions & 16 deletions container/runtime_evidence.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,11 +6,12 @@

RUNTIME_EVIDENCE_SCHEMA = "opencode-eval-runner/runtime-evidence/v1"

STATUSES = {"complete", "incomplete", "unsupported", "invalid"}
FIELD_STATES = {"available", "redacted", "omitted", "unsupported"}
MODES = {"native", "code_mode"}
OUTCOMES = {"success", "error", "missing"}
PROCESS_STATES = {"completed", "timeout", "interrupted", "unsupported"}
STATUSES = frozenset({"complete", "incomplete", "unsupported", "invalid"})
FIELD_STATES = frozenset({"available", "redacted", "omitted", "unsupported"})
MODES = frozenset({"native", "code_mode"})
OUTCOMES = frozenset({"success", "error", "missing"})
PROCESS_STATES = frozenset({"completed", "timeout", "interrupted", "unsupported"})
EVENT_KINDS = frozenset({"start", "terminal"})

BOUNDARY_NATIVE = "native"
BOUNDARY_CODE_MODE_EXECUTION = "code_mode_execution"
Expand Down Expand Up @@ -136,11 +137,6 @@ def _codes(raw: Any, where: str) -> list[str]:
return values


STATUSES = ("complete", "incomplete", "unsupported", "invalid")
FIELD_STATES = frozenset({"available", "redacted", "omitted", "unsupported"})
PROCESS_STATES = frozenset({"completed", "timeout", "interrupted", "unsupported"})
EVENT_KINDS = frozenset({"start", "terminal"})

_GLOBAL_INVALID = frozenset({
"invalid_accounting_input",
"malformed_observation",
Expand Down Expand Up @@ -178,7 +174,6 @@ def _new_boundary(*, declared_supported: bool = False, declared_unsupported: boo
"terminals": 0,
"missing_terminals": 0,
"required_fields_omitted": 0,
"required_fields_truncated": 0,
"required_fields_unsupported": 0,
"issues": ["unsupported_boundary"] if declared_unsupported else [],
}
Expand Down Expand Up @@ -221,8 +216,8 @@ def account_runtime_evidence(

Each observation must contain kind, sequence, invocation_id,
boundary, and required_fields. required_fields maps semantic
field names to available, redacted, omitted, truncated, or
unsupported. Adapters decide which fields are required; this layer only
field names to available, redacted, omitted, or unsupported.
Adapters decide which fields are required; this layer only
accounts for their explicit states.

observation_closed is an ordinary correctness signal from the capture
Expand All @@ -242,7 +237,6 @@ def account_runtime_evidence(
"ambiguous_invocations": 0,
"duplicate_sequences": 0,
"required_fields_omitted": 0,
"required_fields_truncated": 0,
"required_fields_unsupported": 0,
"process_state": process_state if process_state in PROCESS_STATES else "invalid",
"supported_boundaries": [],
Expand Down Expand Up @@ -416,8 +410,6 @@ def account_runtime_evidence(
_add_issue(issues, "missing_terminal")
if coverage["required_fields_omitted"]:
_add_issue(issues, "required_field_omitted")
if coverage["required_fields_truncated"]:
_add_issue(issues, "required_field_truncated")
if coverage["required_fields_unsupported"]:
_add_issue(issues, "required_field_unsupported")
if unsupported:
Expand Down Expand Up @@ -559,6 +551,43 @@ def _parent(start: Mapping[str, Any]) -> dict[str, Any]:
return field_unavailable("omitted", "session_parent_unavailable")


def _available_string(raw: Any) -> str | None:
if (
isinstance(raw, Mapping)
and raw.get("state") == "available"
and set(raw) == {"state", "value"}
and type(raw.get("value")) is str
and raw["value"]
):
return raw["value"]
return None


def _code_parent_is_observed(
start: Mapping[str, Any],
starts: Mapping[str, Mapping[str, Any]],
) -> bool:
parent_id = start.get("parent_invocation_id")
if type(parent_id) is not str or not parent_id:
return False
outer = starts.get(parent_id)
if not isinstance(outer, Mapping) or outer.get("kind") != "native_start":
return False
if outer.get("boundary") != "tool-execute-before":
return False
outer_sequence = outer.get("sequence")
child_sequence = start.get("sequence")
if type(outer_sequence) is not int or type(child_sequence) is not int or outer_sequence >= child_sequence:
return False
outer_tool = _available_string(outer.get("tool"))
if outer_tool is not None and outer_tool != "execute":
return False
return all(
outer.get(name) == start.get(name)
for name in ("session_id", "agent", "message_id", "call_id")
)


def _capture_issue_kind(code: str) -> str:
if code in {
"malformed_capture",
Expand Down Expand Up @@ -643,6 +672,11 @@ def build_runtime_evidence(
for invocation_id, start in sorted(starts.items(), key=lambda item: item[1]["sequence"]):
mode = "native" if start.get("kind") == "native_start" else "code_mode"
boundary = BOUNDARY_NATIVE if mode == "native" else BOUNDARY_CODE_MODE_EXECUTION
if mode == "code_mode" and not _code_parent_is_observed(start, starts):
# A Code Mode parent is authoritative only when it resolves to the
# observed model-facing outer execute invocation with the same
# runtime identity. A dangling opaque token is invalid evidence.
accounting_events.append({})
terminal = terminals.get(invocation_id)
expected_terminal = "native_terminal" if mode == "native" else "code_terminal"
if terminal is not None and terminal.get("kind") != expected_terminal:
Expand Down Expand Up @@ -904,6 +938,7 @@ def validate_runtime_evidence(raw: Any) -> dict[str, Any]:
previous_start = -1
terminal_count = 0
reconstructed: list[dict[str, Any]] = []
validated_by_id: dict[str, Mapping[str, Any]] = {}

for index, observation in enumerate(evidence["observations"]):
where = f"runtime_evidence.observations[{index}]"
Expand All @@ -924,6 +959,26 @@ def validate_runtime_evidence(raw: Any) -> dict[str, Any]:
_string(parent["id"], f"{where}.parent.value.id")
if mode == "code_mode":
required["parent"] = parent_state
if status != "invalid":
_require(
parent_state == "available"
and isinstance(parent, dict)
and parent.get("kind") == "invocation",
f"{where}.parent must identify an observed outer invocation",
)
outer = validated_by_id.get(parent["id"])
_require(
isinstance(outer, Mapping) and outer.get("mode") == "native",
f"{where}.parent must reference an earlier native observation",
)
outer_tool_state, outer_tool = _field(outer.get("tool"), f"{where}.parent.outer.tool")
if outer_tool_state == "available":
_require(outer_tool == "execute", f"{where}.parent must reference outer execute")
for identity_name in ("actor", "session_id", "message_id", "call_id"):
_require(
outer.get(identity_name) == observation.get(identity_name),
f"{where}.parent outer identity does not match {identity_name}",
)

start_sequence = observation["start_sequence"]
_require(type(start_sequence) is int and start_sequence >= 0, f"{where}.start_sequence must be >= 0")
Expand All @@ -948,6 +1003,7 @@ def validate_runtime_evidence(raw: Any) -> dict[str, Any]:
)
if outcome == "missing":
_require(terminal_state == result_state == error_state == "omitted", f"{where} missing terminal must be explicit")
validated_by_id[invocation_id] = observation
continue
_require(terminal_state == "available" and terminal_sequence > start_sequence, f"{where}.terminal_sequence must follow start")
_require(terminal_sequence not in seen_sequences, f"{where}.terminal_sequence must be unique")
Expand All @@ -972,6 +1028,7 @@ def validate_runtime_evidence(raw: Any) -> dict[str, Any]:
"boundary": boundary,
"required_fields": terminal_required,
})
validated_by_id[invocation_id] = observation

if aggregate_available:
_require(starts == len(evidence["observations"]), "coverage starts does not match observations")
Expand Down
4 changes: 3 additions & 1 deletion docs/code-mode-observer-experiment.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ Stock OpenCode 2.0.23 supports a useful partial observation path:
| unique inner invocation identity | runner-owned `tool.transform` wrapper allocates an ID when the decoded leaf handler is actually entered | **supported** |
| actual selected tool | wrapper is attached to the effective registered tool | **supported** |
| executable input | wrapper runs after core input decoding and receives the value passed to the leaf handler | **supported** |
| real outer `execute` binding | the real `Tool.Context` carries Session/message/outer CallID into each inner leaf | **supported** |
| real outer `execute` binding | the real `Tool.Context` carries Session/message/outer CallID into each inner leaf; production evidence also records the outer call at `execute.before` | **supported** |
| start / handler-terminal ordering | runner observer sequence around the transformed leaf handler | **supported** |
| exact final value delivered to the Code Mode script | no supported public stock boundary exposes it with unique inner identity | **unsupported** |
| exact final error seen by the script catch path | no supported public stock boundary exposes it with unique inner identity | **unsupported** |
Expand Down Expand Up @@ -56,6 +56,8 @@ per-call ID at a real execution boundary and observe:
- decoded/executable input;
- the real Session ID, agent, message ID and outer `execute` CallID carried in
`Tool.Context`;
- a parent ID that resolves to the separately observed model-facing outer `execute`
invocation rather than to an internal correlation-only token;
- handler return or throw;
- start and handler completion order.

Expand Down
6 changes: 4 additions & 2 deletions docs/native-tool-observer.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,13 +8,15 @@ This observer is the production native/direct-call adapter for the trusted-check

The runner injects a reviewed Promise plugin through stock \`OPENCODE_CONFIG_CONTENT\` so its transform is applied after project/global transforms.

For native/direct tools it observes:
For registered native/direct tools it observes:

1. decoded/executable input at the transformed \`tool.execute\` boundary;
2. Session-owned terminal events:
- \`session.tool.success\`
- \`session.tool.failed\`.

Stock Code Mode's model-facing `execute` tool is synthetic: OpenCode creates it inside `Tool.snapshot` after registration transforms have run. The observer records that outer invocation at the stock `execute.before` hook. Its input is the exact effective value seen by that hook, before `CodeMode.Input` decode. The same Session terminal events settle it.

The Session terminal is intentional. A rejected transformed handler can bypass \`tool.execute.after\` while stock OpenCode still settles the invocation through \`session.tool.failed\`.

## Identity and ordering
Expand All @@ -25,7 +27,7 @@ The adapter correlates using the real runtime identity tuple:

It never correlates by FIFO, input equality, tool name, or completion order.

The public invocation ID is opaque. Dynamic identity/input/result/error fields are sanitized before the internal capture file is written.
The public invocation ID is opaque. Code Mode inner records reference the invocation ID of the observed outer `execute` record; a dangling or identity-mismatched parent is invalid evidence. Dynamic identity/input/result/error fields are sanitized before the internal capture file is written.

The canonical \`runtime_evidence\` builder then validates:

Expand Down
Loading