Repository navigation
refactor: scope eval trust to reviewed runtime observations - #45
Conversation
Supersession checkpoint
CI #265 is green on Verified on this replacement head:
The old #41 implementation is intentionally not merged wholesale. Evidence-safety, disposable-state, observation semantics, delegation/concurrency probes, and other useful pieces will be ported only where they fit the new stock-runtime trust model. Patched-runtime, protected-channel, remote PluginHost/isolation, HMAC trust-boundary, and signing-as-TRUST-prerequisite work stay out. |
bateau84
left a comment
There was a problem hiding this comment.
Final independent review — NOT READY
BLOCKING — outer Code Mode execute invocation is missing from authoritative evidence
container/native_observer.ts observes registered direct tools through tool.transform, but stock OpenCode creates the synthetic Code Mode execute tool later inside Tool.snapshot; it is therefore not transform-wrapped.
The observer's execute.before hook only creates/updates an internal activeOuter correlation token. It does not emit a public/native start record for the actual outer execute invocation, and Session terminal handling ignores the outer terminal because there is no corresponding nativeActive entry.
Consequences:
- a Code Mode run can have an actual model-facing
executetool invocation while the publicnativeboundary reports zero such calls and remainscomplete; - an absence assertion over native/direct tools can therefore be false-positive for
execute; - Code Mode
parentpoints to an internal opaque token that has no corresponding outer observation inruntime_evidence; - the acceptance gate currently only checks that this parent ID is non-empty/shared; it does not require a matching observed outer invocation.
This conflicts with the v1 contract's description of native as direct/native execution and its claim of parent binding to the actual outer invocation.
Required correction
Observe the model-facing outer execute call through a supported stock-2.0.23 runtime surface and include it in authoritative evidence, or explicitly add a separate unsupported/partial outer-execute boundary so absence cannot be claimed.
Do not infer it from product stdout/model data and do not patch OpenCode. Stock Session tool input/called/terminal surfaces or another reviewed supported boundary should be investigated.
Add provider-free regression coverage that proves:
- one Code Mode run contains an authoritative outer
executeobservation; - inner records bind to that exact observed outer invocation;
- a no-native/absence assertion cannot report zero when
executeactually ran; - concurrent inner calls retain the same observed outer parent.
Other review results
The scoped threat model is otherwise preserved: I found no reintroduction of protected-channel, signing, broker, OpenCode patching, or hostile-plugin isolation. The native Session terminal boundary, fail-closed accounting, explicit Code Mode finality limitation, pre-sink evidence projection, host re-validation, and the final-head CI/acceptance runs are coherent.
Verdict: NOT READY until the outer-execute coverage hole is closed or explicitly represented as unsupported.
Final integration cleanup for PR #45: keep runtime_evidence/v1 as the sole evidence authority, remove duplicate/dead accounting and competing eligibility signals, retain stock OpenCode 2.0.23 behavior, and document unsupported areas.
## Purpose Correct the blocking outer-`execute` evidence gap found during the independent final review of #45. This branch is based on the post-#52 target head `b10c52d50aaf0c038e77e80eaca3890347d5fa82`. ## Defect Stock OpenCode 2.0.23 creates the model-facing Code Mode `execute` tool inside `Tool.snapshot`, after registration transforms have run. The integrated observer therefore did not emit a native observation for that real outer invocation. Inner records were linked only to an internal opaque token. That allowed the public native boundary to report complete/zero outer-`execute` observations while Code Mode had actually run. ## Correction - Observe the real model-facing outer `execute` invocation at stock `execute.before`. - Keep `session.tool.success` / `session.tool.failed` as terminal authority. - Bind Code Mode inner records to that observed outer invocation ID. - Reject dangling, wrong-tool, or identity-mismatched parents in canonical accounting/validation. - Strengthen provider-free acceptance checks for outer observation, exact parent resolution, identity agreement, non-zero native coverage, concurrency, and reverse completion. - Document that synthetic outer `execute` input is exact at `execute.before` but pre-`CodeMode.Input` decode. ## Scope No threat-model expansion. No OpenCode patch/fork, protected channel, signing, broker, process isolation, or hostile-plugin machinery. Targets `refactor/trusted-checkout-evidence`. ## Validation Head: `221af580640360c16e263cc05136d3be0d090a0f` - CI #307 / run 37525849151: **PASS**, 94/94 Python tests; stock Code Mode diagnostic probe PASS; OpenCode **2.0.23**. - Provider-free runtime evidence acceptance #21 / run 37525849081: **PASS**, 6/6 helper tests and all 10 scenarios. - Stock native observer integration #33 / run 37525849107: **PASS**. - PR is mergeable against the current #45 branch. #53 is closed as superseded; it became dirty only because #52 merged while this review was in progress.
bateau84
left a comment
There was a problem hiding this comment.
Final re-review — NOT READY
#54 is merged and closes the outer Code Mode execute coverage defect. The integrated head is green across CI, provider-free acceptance, and the stock native observer gate.
BLOCKING — authoritative capture is still target-writable
The observer still writes authoritative records to:
/tmp/runtime/runtime-observer.jsonl
runner/cli.py mounts all of /tmp as a writable tmpfs for the evaluated OpenCode process. Normal model-driven tool/shell subprocesses execute under that same runtime user and can therefore potentially read, truncate, replace, or append to this file.
This conflicts with the trusted-checkout contract already documented in this PR: target-writable evidence files do not independently establish that an event occurred.
This is not hostile-plugin resistance and does not require the old TRUST-001 security architecture. It is ordinary eval correctness: model/tool execution must not be able to manufacture or alter the evidence used to score itself.
Required correction
Move authoritative observer transport off the evaluated process's writable filesystem. A small runner-owned pipe/socketpair/inherited-FD style transport is sufficient if normal tool subprocesses cannot write to it. No signing, HMAC, broker, remote PluginHost, or OpenCode patch is required.
Add one provider-free tamper regression proving an evaluated tool/subprocess cannot create, replace, truncate, or append authoritative records.
Status
- outer
executeauthoritative observation: FIXED - inner-to-outer binding: FIXED
- Code Mode finality remains correctly
unsupported - stock OpenCode 2.0.23: PASS
- CI / acceptance / native observer gates: PASS
- authoritative transport integrity against normal evaluated tool writes: BLOCKING
Verdict: NOT READY until the target-writable capture file is removed from the authoritative path.
|
Final blocker follow-up is now green in #55.
Current #55 validation:
Both final review blockers are closed by the combined base + #55 state. PR #45 remains open/unmerged pending integration/review of #55. |
## Scope Focused corrective child of #45, targeting `refactor/trusted-checkout-evidence`. The base now includes merged #54 (`b10c8e9`), which closes the authoritative outer Code Mode `execute` blocker. This PR is rebased directly on that commit and contains only the remaining authoritative-capture transport correction plus its regression coverage. Stock OpenCode remains **2.0.23**. No threat-model expansion. ## Blocker 1 — outer Code Mode `execute` **Closed in the base by #54 and revalidated here.** The current integrated runtime evidence: - observes the real synthetic outer `execute` at the supported stock `execute.before` boundary; - records Session, agent, message, real CallID, supported hook input, and a unique runtime invocation ID; - uses stock `session.tool.success` / `session.tool.failed` as terminal authority; - binds every inner Code Mode call to that exact observed outer invocation; - rejects dangling or identity-mismatched parents; - keeps exact Code Mode caller-finality explicitly unsupported. The Code Mode probe remains green after the capture-transport change. ## Blocker 2 — remove target-writable authoritative capture The production observer no longer uses: `/tmp/runtime/runtime-observer.jsonl` Authoritative observer records now cross the process boundary through a **runner-owned one-connection loopback stream**: 1. the runner binds an ephemeral listener on `127.0.0.1` immediately before the real OpenCode invocation; 2. the trusted observer connects during plugin startup; 3. the listener closes after accepting that one connection; 4. the observer removes the endpoint from `process.env` before evaluated tool subprocesses run; 5. sanitized observer records flow over the established stream; 6. the runner drains and validates those bytes through the existing canonical builder/validator/accounting path. A direct inherited memfd/FD was tried first, but stock `@opencode/cli@2.0.23` crosses an internal process boundary that does not preserve arbitrary extra descriptors. The one-connection stream keeps stock OpenCode unchanged and avoids filesystem authority without introducing security-platform machinery. Protection against a deliberately malicious same-process plugin remains explicitly out of scope. ## Tamper regression Provider-free acceptance includes an evaluated tool that spawns `/bin/sh` and: - creates the old capture path; - deletes it; - recreates it; - appends a forged observer record. The scenario passes only if the shell tamper completes, authoritative evidence remains complete/eligible, the real tamper tool is observed, and the forged invocation/tool never appears in `runtime_evidence`. ## Preserved - `opencode-eval-runner/runtime-evidence/v1`; - canonical builder / validator / accounting; - pre-sink credential sanitization; - stock Session terminal authority; - Code Mode exact caller-finality remains unsupported; - normal `invoke` behavior; - assertion-scoped eligibility; - stock OpenCode 2.0.23 only. No signing/HMAC, protected channel, remote PluginHost, capability broker, hostile-plugin isolation, or OpenCode patch is introduced. ## Validation Current clean head is based directly on merged #54 and is mergeable. - **CI #329 / run 37532684965 — PASS** - **95/95 Python tests** - stock Code Mode probe: **PASS** (`diagnostics_passed: true`) - OpenCode **2.0.23** - **Provider-free runtime evidence acceptance #43 / run 37532684985 — PASS** - all **11/11** scenarios: - native_success - native_error - code_success - code_caught_error - concurrent_reverse - delegation - timeout - interrupted - redaction - collector - capture_tamper - **Stock native observer integration #55 / run 37532684878 — PASS** - stock OpenCode 2.0.23 - capture complete - provider-free native observation ## Merge scope This PR is ready for review against `refactor/trusted-checkout-evidence`. Do **not** merge PR #45 as part of this change.
bateau84
left a comment
There was a problem hiding this comment.
Final independent review — READY TO MERGE
Reviewed integrated head:
e3eb017ff070f3956119472a7d0fb74faa247102
Both previously blocking correctness defects are closed:
- Outer Code Mode
executeauthority — the real model-facing syntheticexecuteinvocation is observed at stockexecute.before, settled by Session terminal events, and inner Code Mode calls bind to that exact observed outer invocation. - Target-writable capture authority — production no longer consumes
/tmp/runtime/runtime-observer.jsonl; authoritative records cross a runner-owned one-connection loopback stream, the listener closes after the trusted observer connects, and evaluated subprocess writes to the old path cannot affect evidence.
Final-head validation:
- CI #330: PASS
- Provider-free runtime evidence acceptance #44: PASS, including
capture_tamper - Stock native observer integration #56: PASS
- stock OpenCode 2.0.23 retained
The canonical runtime_evidence/v1 contract remains the sole eligibility authority. Missing/lost/malformed capture fails closed. Product outcome remains separate from evidence outcome. Exact Code Mode caller-final value/error remains explicitly unsupported. No OpenCode patch, signing/HMAC, protected channel, broker, remote PluginHost, or hostile-plugin isolation architecture has been reintroduced.
No BLOCKING, IMPORTANT, or MINOR correctness findings requiring changes remain.
Verdict: READY TO MERGE #45.
## Product Readiness finding PR #45 makes `runtime_evidence` mandatory and the host CLI rejects a result that omits or violates `opencode-eval-runner/runtime-evidence/v1`, but the main README's Result contract example still showed the older shape without that object. The compatibility consequence for legacy/custom images was also not stated. That leaves the primary consumer-facing contract contradictory even though the runtime implementation is correct. ## Changes - make `runtime_evidence` explicit in the README Result contract example; - state that every official transport result carries it, with Copilot reporting `unsupported`; - document that host runner + overridden/custom image are a compatibility pair and must be upgraded together; - add the assertion-level consumer rule: required boundaries must be `complete` and required exact fields `available`; diagnostic convenience fields cannot fill gaps; - document the current aggregate and per-field capture bounds and their fail-closed meaning. ## Scope Documentation only. No runtime behavior, OpenCode patching, evidence channel redesign, signing, broker, PluginHost, or hostile-plugin isolation changes. Targets `refactor/trusted-checkout-evidence` and is based directly on reviewed PR #45 head `e3eb017ff070f3956119472a7d0fb74faa247102`.
bateau84
left a comment
There was a problem hiding this comment.
Focused Product Readiness re-pass — PRODUCT READY
Reviewed integrated head:
a67c931696d1ae2b2d16f49beab6c01c912aa02b
PR #56 is merged and closes the only Product Readiness blocker identified in the previous gate.
Evidence-readiness decision procedure
The public contract now gives consumers an explicit assertion-level procedure:
- validate
runtime_evidence/v1; - overall
incompleteorinvalidmakes every runtime assertion ineligible for PASS; - declare the boundary/boundaries required by the assertion and require each to be
complete; - if an assertion needs an exact field, require that field to be
available;redacted/omittedare incomplete for that assertion andunsupportedis unsupported; - never backfill authoritative facts from
tools,actions,tool_result_evidence, stdout/stderr, model text, or workspace files.
This is sufficient for Loom to make evidence-readiness decisions without moving judging semantics into the runner.
Product Readiness findings
BLOCKING: none.
IMPORTANT non-blocking follow-ups:
- Loom must migrate its verdict path to consume
runtime_evidencebefore using this capability to authorize PASS. - Aggregate capture-limit failures currently diagnose generically; dedicated
capture_size_limit/capture_record_limitreason codes would improve operations but cannot create false PASS. - Any OpenCode upgrade beyond reviewed stock 2.0.23 should re-run the observer capability probes as an upgrade gate.
Accepted release limitations
- exact Code Mode script-visible final value/error remains
unsupported; github-copilot-cliruntime observation remainsunsupported;- deliberately malicious same-process evaluated plugins remain outside the trusted-checkout profile;
- stock OpenCode 2.0.23 is the reviewed runtime.
Compatibility / migration
The README now states that runtime_evidence is mandatory, host and image form a compatibility pair, and legacy/custom images without valid v1 evidence are rejected. The primary Result contract therefore matches actual host behavior.
Exact-head validation
- CI #332: PASS — 95/95 tests
- Provider-free acceptance #45: PASS — all 11 scenarios including
capture_tamper - Stock native observer #57: PASS
- stock OpenCode 2.0.23 confirmed
No additional Product Readiness pass is required for PR #45 unless the runtime or public contract changes again. Loom should perform its own integration/acceptance pass when its scoring path is migrated.
PRODUCT READY — PR #45 may merge
## Purpose Close the remaining user-facing documentation gaps before PR #45 merges. The runner is an isolated **invocation execution boundary**, not an eval-suite engine. This PR makes that contract explicit and documents every supported public invocation surface. ## Changes - add `docs/invocation-usage.md` as the authoritative usage/interface reference; - explicitly state that there is no input-JSON request API today; - document the invocation input contract: CLI/Action arguments plus prompt/system/config seed files; - document all **24/24 CLI options**, including defaults and transport applicability; - document all runner-specific environment overrides and recognized provider credentials; - document all **21/21 GitHub Action inputs** and defaults; - document Action precedence: repository `command` mode, direct invocation mode, and setup-only mode; - document which CLI capabilities are not exposed as direct Action inputs; - add a complete "ways to run" matrix; - clarify that eval cases/assertions/judging remain owned by the calling harness; - correct the stale README description of the current zero-inference plugin activation preflight; - link the complete reference prominently from README local and Action usage sections. ## Validation Cross-checked the docs against the current implementation: - CLI flags documented: **24/24** - Action inputs documented: **21/21** - runner-specific environment overrides documented: **8/8** Documentation only; no runtime or contract behavior changes. Targets `refactor/trusted-checkout-evidence` / PR #45.
## Scope Worker B for #58 (Task 1/9): post-PR-#45 runner contract/invariant inventory. This PR is intentionally documentation-only. It does **not** implement the eval engine, change `invoke`, migrate Loom, add a public `eval` CLI, or define the final Task 1 architecture. ## What it captures - one `invoke` = exactly one isolated invocation; - `opencode-eval-runner/v1` result boundary; - mandatory/authoritative `runtime_evidence/v1`; - assertion-scoped evidence readiness and field states; - product vs evidence vs infrastructure failure separation; - inner vs outer timeout behavior; - retry ownership above `invoke`, cross-checked against Loom's current narrow retry policy; - OpenCode vs Copilot transport capability differences; - host output validation and host/image compatibility; - CLI and GitHub Action interface limits; - explicit unsupported areas and trusted-checkout scope; - current Loom behaviors that must not accidentally become runner contracts. ## Validation Cross-checked against the landed post-#45 contract: - `main` / PR #45 merge commit `fd9da10cbe2a8182fc8910ec199a221250deb3ca` - final reviewed PR #45 head `25478106773969870931af3fb34e7b2586746606` - `runner/cli.py` - `container/invoke.py` - `docs/invocation-usage.md` - `docs/runtime-evidence-contract.md` - `action.yml` - Loom `functionality-anchor-requirements-coherance` - `scripts/run-evals.py` blob `30f2ae87764be180173a9e5b4b609ee79b2f063d` `eval-engine/01-architecture` now matches `main` at the PR #45 merge commit. The worker branch has been rebased onto that exact state, leaving a one-file documentation diff. Part of #58.
## Summary Part of #58 (Task 1/9), parallel worker A. Adds a focused inventory of the current Loom eval harness from: - bateau84/loom branch functionality-anchor-requirements-coherance - scripts/run-evals.py at 30f2ae87764be180173a9e5b4b609ee79b2f063d The inventory traces and classifies: - case loading/normalization; - target/workspace setup; - invocation and orchestration retries; - runtime evidence and legacy observers; - judge construction/parsing; - deterministic checks; - classification; - artifacts; - iterations/concurrency; - skill ablation; - reporting. It separates each responsibility into generic orchestration, Loom/project semantics, low-level invoke behavior, or legacy/compatibility behavior that should not migrate. ## Post-#45 validation Cross-checked against the merged runner contracts on the integration base: - one invoke remains one isolated invocation; - retries remain outside invoke; - runtime_evidence/v1 is the only authoritative runtime-evidence object; - tools/actions/tool_result_evidence/stdout/stderr/model text are diagnostic only; - project code owns cases, assertions, judging, thresholds, and behavioral meaning. The document explicitly marks Loom's older direct-container fallback, observer/tool-result reconstruction, runner-safety adapter, and Loom-side evidence_safety eligibility machinery as non-migration paths. ## Scope Documentation only. No runtime changes, no Loom migration, no final engine API/module design, no universal assertion DSL, and no public eval CLI. PR target: eval-engine/01-architecture.
## Task Closes #58 — **Task 1 of 9: Define generic orchestration boundary and extraction contract**. This PR is architecture/documentation only. It does not implement the eval engine. ## Inputs integrated Parallel worker results were merged into the integration branch first: - #65 — runner contract/invariant inventory - #66 — Loom eval-runner responsibility inventory Both are based on the post-#45 merged `main` contract. ## Architecture decisions The final synthesis defines: - one `invoke` = exactly one isolated invocation; - retry ownership exclusively in orchestration; - `runtime_evidence/v1` as the only authoritative runtime-evidence source; - explicit separation of product, evidence, and infrastructure outcomes; - generic ownership for planning, iterations, concurrency, attempts, evidence-readiness mechanics, judge lifecycle, tri-state classification, artifacts, and summaries; - project/profile ownership for case semantics, fixtures, prompts, evidence requirements, deterministic assertions, judge meaning, thresholds, and ablation policy; - explicit non-migration of Loom's old direct-container fallback, legacy observer/evidence reconstruction, and superseded runner-safety paths; - concrete internal module decomposition for Tasks 2–5; - concrete internal contracts for normalized cases/jobs, invocation specs, attempt records, retry policy, evidence requirements/readiness, check outcomes, semantic decisions, and durable artifacts; - initial standard/runtime concurrency-lane behavior; - versioned internal run/artifact schemas; - explicit Task 2–5 handoff boundaries. No universal assertion DSL, public `eval` CLI, input JSON API, Loom migration, or skill-ablation implementation is introduced. ## Files - `docs/eval-engine-architecture.md` — final Task 1 synthesis - `docs/eval-engine-runner-contracts.md` — worker B inventory - `docs/loom-eval-runner-responsibility-inventory.md` — worker A inventory ## Branch topology - integration branch: `eval-engine/01-architecture` - base: `main` - worker branches were merged into the integration branch before synthesis. ## Acceptance The architecture is intended to let Tasks #59–#62 proceed independently without redesigning the ownership boundary. Documentation only; no runtime behavior changes.
## Summary Consolidates PR CI after merged #45 without changing runtime/evidence semantics. ### Workflow cleanup - keeps `.github/workflows/ci.yml` and `.github/workflows/publish.yml`; - removes `.github/workflows/native-observer-integration.yml`; - removes `.github/workflows/runtime-evidence-acceptance.yml`; - keeps stable PR job names `test` and `runtime-integration`. ### `test` Runs on every PR and keeps fast/general checks only: - Python compile checks; - full unit-test discovery; - fake Copilot image + GitHub Action token/sanitization smoke test. It no longer builds the real OpenCode or Copilot runtime images. ### `runtime-integration` The job exists on every PR. It decides applicability internally, so documentation-only/unrelated PRs still receive a successful `runtime-integration` check instead of a missing required check. For runtime-relevant changes it: 1. builds the stock OpenCode image once; 2. asserts OpenCode remains 2.0.23 and checks the pinned model-variant interface; 3. runs the provider-free native observer probe; 4. runs the stock Code Mode capability probe; 5. runs the full provider-free runtime-evidence acceptance suite; 6. runs the OpenCode plugin activation preflight; 7. runs the config-root plugin dependency-resolution preflight; 8. builds/checks the real Copilot image only when shared/Copilot runtime inputs changed; 9. uploads the consolidated integration reports. This removes the prior three-way rebuild of the same stock OpenCode image on runtime PRs. ### Dead code cleanup Confirmed `loaded_skills_from_export` and `assistant_from_export` have no production/runtime/runner/documented caller on current `main`; their only remaining references were their own definitions and the two obsolete Session-export unit tests. Removed: - `loaded_skills_from_export`; - `assistant_from_export`; - `test_failed_exported_skill_call_is_not_reported_as_loaded`; - `test_session_export_preserves_tool_inputs_as_actions`. Observable timing compatibility fields were left untouched because they are output data and were not proven dead. ### Required-check recommendation After this PR is merged, the repository ruleset should require exactly: - `test` - `runtime-integration` Keep strict/up-to-date required status checks enabled. `Publish images` must not be required. The ruleset is **not changed in this PR**: the available repository tools do not expose ruleset mutation, and enabling the new required check before the workflow lands on `main` could block other PRs that cannot yet produce it. Follow-up in GitHub UI after merge: `Settings → Rules → Rulesets → <active main ruleset> → Require status checks to pass` Keep `test`, add `runtime-integration`, keep strict/up-to-date enabled, and do not add `Publish images`. ### Validation Final head: `a5bb16e9bbe55efde1ab7a7f5eb9de8b07336495` CI run **#341** / **37593353163**: **PASS** - `test`: **PASS** - Python compile checks: PASS - unit tests: **93/93 PASS** - fake Copilot image + direct Action smoke: PASS - repository-command token sanitization smoke: PASS - `runtime-integration`: **PASS** - stock OpenCode image built **once**: PASS - pinned OpenCode: **opencode v2.0.23** - native observer probe: PASS - Code Mode capability probe: PASS, including explicit unsupported caller-final value/error behavior - provider-free runtime-evidence acceptance: **11/11 scenarios PASS** - native success - native error - Code Mode success - caught Code Mode error - concurrent reverse completion - delegated Session ancestry - timeout - interrupted execution - credential redaction - forged collector-shaped payload rejection - capture tamper regression - OpenCode plugin activation preflight: PASS - config-root dependency-resolution preflight: PASS - real Copilot image build/version/interface check: PASS The internal path decision excludes documentation-only changes such as `README.md` / `docs/**`; for those PRs the `runtime-integration` job still starts and succeeds through its explicit not-applicable path, so a future required check will not hang. No `runtime_evidence/v1`, trusted-checkout threat-model, or `invoke` semantic changes were made.
Integrated direction
This PR now integrates the completed child work from #46–#51 into one trusted-checkout runtime-evidence implementation on stock OpenCode 2.0.23.
The normal Loom eval profile remains a trusted-checkout evaluation. This PR does not add an OpenCode patch/fork, remote PluginHost, protected channel, HMAC/signing boundary, capability broker, or hostile-plugin isolation.
Canonical public contract
The only authoritative public runtime-evidence object is:
opencode-eval-runner/runtime-evidence/v1Raw observer records are internal adapter input only.
Existing
tools,actions,tool_result_evidence, stdout/stderr, Session/model text, and workspace files remain convenience/diagnostic data and are not promoted into runtime evidence. Final cleanup #52 also removes competing runtimeevidence_eligiblesignals from the diagnostic tool-result/safety projections.There is one canonical builder/validator that owns:
complete | incomplete | unsupported | invalid;Unknown coverage is never converted to zero.
Boundary semantics
The v1 contract exposes:
nativecode_mode_executioncode_mode_finalityOverall
completemeans the supported capture boundaries are complete. It does not mean every possible assertion is supported.This allows:
An assertion that only requires
nativeevidence can therefore remain eligible even though Code Mode finality is unsupported. Assertions that require a redacted/omitted/unsupported exact field remain ineligible for that field.Native observation
The integrated observer preserves #51's stock-2.0.23 boundary:
It records opaque invocation identity, actual tool, agent, Session, message, real CallID, executable input, Session ancestry, terminal result/error, and monotonic start/terminal ordering.
Correlation is identity-based. It does not use FIFO, input equality, tool name, or completion order.
The raw former
native_tool_observationsprojection is no longer a public result.Code Mode
The #50 stock-runtime result is retained faithfully.
Supported facts:
executeCallID;The transformed handler's earlier value/error is not promoted as caller-final evidence.
Exact final script-visible value/error is represented as:
{ "state": "unsupported", "reason": "stock_codemode_final_boundary_not_exposed" }Evidence safety
The #47 safety behavior is applied to authoritative dynamic values before their first observation persistence/output sink:
Public field states are only:
availableredactedomittedunsupportedProduct
exit_code, timeout, and product success/failure remain separate from evidence eligibility.Child PR disposition
exactterminology is not part of the public runtime-evidence schema.runtime_evidence; no publicnative_tool_observationsobject remains.Child PRs #46–#51 have no remaining implementation authority after this integration. Final cleanup #52 removes parallel-work duplication without adding architecture. Explicit unsupported areas are: Code Mode caller-final value/error on stock 2.0.23, OpenCode runtime observation for the
github-copilot-clitransport, and protection against a deliberately hostile plugin sharing the trusted OpenCode process.Final integration cleanup
#52 is the final child PR targeting this branch. Its head
a3f575d78f4b14d62b76d14dceea49b76496fd2d:truncatedaccounting;evidence_eligiblesignals from diagnostictool_result_evidence/ safety projections;invokebehavior.Validation on #52:
#52 has now been squash-merged into this branch as
b10c52d50aaf0c038e77e80eaca3890347d5fa82.Final-head validation
Head:
b10c52d50aaf0c038e77e80eaca3890347d5fa82native_success: PASSnative_error: PASScode_success: PASScode_caught_error: PASSconcurrent_reverse: PASSdelegation: PASStimeout: PASSinterrupted: PASSredaction: PASScollector: PASSReview state
The integration cleanup from #52 is merged. PR #45 now contains the final integrated implementation and remains ready for final independent review.
PR #45 remains open/draft and is not merged.
Independent final review
Final review of head
b10c52d50aaf0c038e77e80eaca3890347d5fa82found one blocking correctness defect: stock Code Mode's synthetic model-facingexecuteinvocation was not emitted as an authoritative native observation, so native coverage could falsely remain complete/zero while Code Mode had run and inner parents resolved only to an internal token.Corrective child PR #54 observes the real outer call at stock
execute.before, binds inner parents to that observed invocation, and makes dangling/wrong-identity parents invalid. #54 is green across CI, the stock native probe, and the 10-scenario provider-free acceptance gate.Review status: NOT READY until #54 is merged. No threat-model expansion is required.