Skip to content

[SUPERSEDED by #45] feat: preserve eval:live semantics while observing runtime tools - #41

Closed
bateau84 wants to merge 62 commits into
mainfrom
feat/trustworthy-execution-observer-export
Closed

bateau84 wants to merge 62 commits into
mainfrom
feat/trustworthy-execution-observer-export

Conversation

@bateau84

@bateau84 bateau84 commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Why

Loom evals must exercise the real OpenCode/Loom host rather than a replacement tool runner. This PR preserves:

bun run eval:live
  -> scripts/run-evals.py
  -> opencode-eval-runner invoke
  -> normal OpenCode/Loom execution

while adding runtime-owned diagnostic observation and an opt-in evidence-safety path underneath invoke.

A final Loom-side scrub cannot protect credentials the runner already printed, persisted, or clipped. The runner therefore accepts Loom's versioned private credential inventory, projects typed evidence before runner output, and requires a matching acknowledgement before host file/print sinks.

The latest addition also provides an explicit disposable OpenCode state profile for provider-free composition. It does not reuse installed auth/database state and does not fabricate database schema or migration journals.

Current source: 2940ae47b51c3210f1805fa5907089d0e4a87f85
Matching tested safety image: ghcr.io/bateau84/opencode-eval-runner@sha256:42381aeb8c44f81527db7593c79d0f9b62a47c645cc481f1bf0d01bfa49d47ab

No merge, default-image-pin change, or installed OpenCode update is authorized.

What changed

  • Native and Code Mode observation under normal invoke: native starts use the actual decoded input and actor/session context; native terminals follow Session truncation and canonical success/failure publication. Code Mode inner observations retain runtime invocation IDs, actual outer-parent binding, final caller-visible values/errors, and independent start/completion ordering.

  • Independent observation flags: Code Mode dispatch is now marked explicitly even when inner observation is disabled. OPENCODE_EVAL_OBSERVATIONS=0 can no longer misclassify inner Code Mode tools as native. The patched runtime is OpenCode 2.0.18-eval.5, and the safety image is rebased onto that verified runtime digest.

  • Safety before runner sinks: container/evidence_safety.py projects typed fields before clipping/output. runner/safe_invoke.py validates the result and acknowledgement before atomic mode-0600 file writes or --print-result. Missing/incompatible policy, timeout, and failure paths cannot fall back to raw output.

  • Explicit fidelity: fields are exact, redacted, or omitted. Short credentials such as 0, 1, text, and low remain protected without renaming JSON keys or protocol discriminators. Retained tool events require dispositions for selectors, input, status, and output/error availability; unexplained missing fields are rejected.

  • Disposable OpenCode state lifecycle: --opencode-state-profile disposable rejects implicit runner seed overrides and database seeds, does not forward ambient provider credentials merely because they exist, and lets pinned OpenCode bootstrap a fresh database in its disposable XDG tree. The runner then attests the runtime-created session_v2 / credential / migration schema, exactly 48 migration IDs, zero pre-inference Session rows, and zero credential rows.

  • Inventory selection matches actual execution: under the disposable profile, credential_seed must be not_selected; explicit auth/config/models/config-root selections must be declared as such. Missing/incompatible policy or explicit database selection fails before provider inference.

  • Reviewed image provenance: the safety image copies only the three reviewed container sources (__init__.py, evidence_safety.py, invoke.py), excluding generated/untracked files from the executed image context.

  • Signing authority remains separate: the default-branch signing workflow validates exact summary kind/version/image bindings, checks runtime source revision where applicable, authenticates Cosign to GHCR, and remains gated by the protected release-signing environment. It cannot activate while this PR remains unmerged.

Exact Loom integration

Loom keeps its existing entrypoint and adds the runner options internally:

/path/to/opencode-eval-runner/bin/opencode-eval-runner invoke \
  --image ghcr.io/bateau84/opencode-eval-runner@sha256:42381aeb8c44f81527db7593c79d0f9b62a47c645cc481f1bf0d01bfa49d47ab \
  --opencode-state-profile disposable \
  --require-evidence-safety \
  --evidence-policy-file /private/disposable/inventory.json \
  --config /path/to/synthetic-provider-config.json \
  --model <existing-model> \
  --workspace <existing-workspace> \
  --prompt-file <existing-prompt> \
  --output <host-result.json>

The disposable profile is opt-in. Ordinary invoke retains its existing seed/default behavior.

The safety schemas are:

Purpose Schema
Private policy loom-eval-credential-inventory/v1
Field availability loom-eval-evidence-safety/v1
Safety acknowledgement opencode-eval-runner/evidence-safety-ack/v1
Safe result opencode-eval-runner/safe-result/v1
Disposable runtime state opencode-eval-runner/runtime-state/v1

See runner evidence safety and disposable OpenCode state.

Why the database lifecycle is legitimate

The disposable profile begins with no database seed. Before any model/provider request, the pinned OpenCode runtime initializes its normal storage/session stack against a fresh XDG data directory.

OpenCode's own bootstrap creates the current generated schema and migration journal. The runner does not create tables or synthetic migration records.

The runner then verifies:

  • session_v2, credential, and migration exist;
  • migration count is exactly 48;
  • first migration is 20260127222353_familiar_lady_ursula;
  • last migration is 20260923013825_project_time_active;
  • Session rows before inference = 0;
  • credential rows before inference = 0.

This avoids the unsupported hand-built database shape that caused the first migration to replay against an already-created session table.

Verification

All five workflows are green on 2940ae47b51c3210f1805fa5907089d0e4a87f85:

Workflow Result
CI #194 Passed; 192 Python tests
Observer boundary #52 Passed
Protected runtime #48 Passed
Normal-invoke runtime #49 Passed
Evidence safety boundaries #25 Passed; 122/122 actual-image checks

The actual-image safety probe demonstrates:

  • ordinary safety projection still works;
  • short/escaped credentials are protected without structural JSON damage;
  • missing/incomplete/future policy variants remain fail-closed;
  • real failure/timeout behavior is preserved;
  • fresh OpenCode database bootstrap succeeds before the provider is contacted;
  • ambient host auth/database traps are not read;
  • explicit DB selection and missing policy fail before provider inference;
  • disposable runtime-state acknowledgement is exact and bound to the pinned migration boundaries;
  • actual Podman 4.9.3 preflight accepts Podman's bare 64-hex image ID, reports canonical image_config=sha256:<64>, and executes/re-inspects the exact content-addressed config ID without weakening RepoDigest, source-revision, initializer/module/invoke-hash, or safety-label checks.

Actual-image evidence, artifact ZIP SHA-256 ce45474e94f67af8fb5b85cd77e0a01193f0598cfe551e081c39fda9d3b434da.

Review status

Copilot independently identified both new and older findings across the safety, image-provenance, event-validation, and signing paths. The current head fixes:

  • safety-option diagnostic abbreviation handling;
  • create-time resource ownership/cleanup;
  • signer summary/image/source binding;
  • GHCR authentication before Cosign;
  • image build-context provenance;
  • pinned migration boundary validation;
  • complete per-event field dispositions.

Copilot's earlier high-severity documentation/integration mismatch (advertising an older safety image than the source-hash handshake accepted) is resolved by the current source/image pair above. The demonstrated Loom Podman preflight blocker is also fixed and covered by an actual Podman 4.9.3 CI proof. The latest high-severity Copilot finding—Code Mode calls being misclassified as native when inner observation was disabled—was fixed in ff61143 and carried into the current safety image. Its thread is resolved. Per explicit instruction, no further Copilot re-review was requested; historical review-summary bodies retain their original "Open" text because those submitted summaries are immutable snapshots.

Scope boundaries

This enables provider-free RSP composition using the real invoke path. It does not claim:

  • protected capture against arbitrary plugins sharing OpenCode process authority;
  • general normal-host run-wide completeness/redaction;
  • signing activation before the workflow/environment exist on the default branch;
  • full protected-capture acceptance.

No merge or default-pin change has been performed.

Import authenticated, invocation-correlated observer records separately from
script-controlled outer output. Reject ambiguous, forged, incomplete or unsafe
evidence and expose explicit coverage and field availability.

Add 35 provider-free exporter/CLI tests and document the proposed Loom producer
contract. Keep the runtime image unchanged; real Loom hook boundary, protected
signer integration and deterministic OpenCode acceptance remain draft blockers.

bateau84 commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Verification update for 406784b9ac584a40d614989c2265dbf051f1288b:

  • The five published file blobs match the locally tested files.
  • Local observer/export suite: 35 tests passed; Python compilation passed.
  • GitHub CI run 142: Python checks passed, including the full unittest discover -s tests -p 'test_*.py' step. Fake Copilot action/token smoke steps also passed.
  • At this check, the OpenCode image build is still running; subsequent image/plugin smoke steps are not yet complete. This is not an all-green CI claim.

CI builds test images using the unchanged image inputs; no replacement release image/digest has been published by this work. The PR remains draft for real Loom observer integration and end-to-end acceptance, independently of CI success.

bateau84 commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Acceptance clarification following handoff feedback

This PR is runner-side export infrastructure, not fulfillment of the feature handoff. Exporter tests do not establish usable end-to-end capture, protected producer feasibility, or independent code approval.

Keep the PR draft while resolving:

  1. Feasible producer/export contract with Loom's Worker. Establish where the trusted observer/signing boundary can actually live, how evaluated code is prevented from forging records or invoking a signer on arbitrary claims, and how unique inner-invocation identities survive concurrent starts/terminals. The current HMAC contract is a proposal, not an agreed prerequisite that Loom must somehow satisfy. Adapt the exporter to a feasible agreed contract rather than treating signing feasibility as a documentation obligation.
  2. Final-result boundary proof. Demonstrate that the observer captures the value/error Code Mode actually receives, not a value that a later hook or runtime transformation changes.
  3. Full CI and independent implementation review. CI run 142 has now completed successfully for 406784b9ac584a40d614989c2265dbf051f1288b. No PR reviews are currently recorded. The handoff feedback is explicitly not code approval.

Integration and deterministic testing can use the PR checkout without merging first.

Distinguish two future decisions:

  • Guarded infrastructure merge: potentially appropriate once the interface is agreed and feasible, the capture boundary is demonstrated, and the opt-in exporter is independently approved with green CI. Broader Loom integration may still continue separately. This is not current approval or authorization to merge.
  • Full handoff acceptance: requires actual deterministic-provider OpenCode native/Code Mode execution through the Loom observer and assertion consumer, satisfying CAP-AC-001 through CAP-AC-004. Runner fixture tests alone do not satisfy this decision.

No code, merge state, or feature-completion claim is changed by this clarification.

…rovider

Exercise overlapping identical calls, caught errors, discarded results, a later
result-mutating hook, and target-written forged capture using the pinned image.
Use real Docker/OpenCode and the PR host importer, not the fake OCI engine.

Unsigned hook/oracle records remain diagnostic only. No substitute producer or
signing requirement is added. Capture acceptance stays explicitly BLOCKED and
nonzero until a demonstrated Loom producer/export boundary is available.
Accept the pinned binary's actual version string and expose measured shared-ID,
caught-throw terminal, and later-hook mutation counterexamples in the summary.
Document fixture scope and keep full capture acceptance blocked. Signing remains
a proposal, and unsigned diagnostic logs are never imported as trusted results.

bateau84 commented Oct 2, 2026

Copy link
Copy Markdown
Owner Author

Runtime-owner handoff: blocked upstream, not an exporter fix

PR #41 cannot export execution facts the runtime does not provide. Keep this PR draft; neither exporter CI nor rejection of a forged file establishes usable producer capture. No merge, runtime modification, or handoff acceptance is authorized by this update.

Coordination status

The published @opencode/cli package points to anomalyco/opencode. Its contribution guide requires core-team design review for new core features. I prepared and attempted to submit a focused feature request there: “Supported Code Mode inner-dispatch observations for deterministic evals.” GitHub rejected creation with HTTP 403: Resource not accessible by integration. No upstream issue was created, no runtime owner has accepted this handoff, and no design agreement is claimed. The request still needs delivery through an account/integration with upstream issue permission.

Boundary requested from OpenCode

Use the existing Code Mode dispatcher, not a competing interpreter or wrapper implementation. Supply a supported observation point with:

  • A unique child invocation identity carried through the actual execution context, plus executable registration identity, actual dispatched input, executing agent/session/message/call, and exact enclosing execute invocation. Session ancestry is not execute nesting; absent parent data is not proof of native execution.
  • Owned snapshots of final returned values or genuine thrown errors after output-changing hooks and relevant Code Mode conversion/error mapping. Returned denial JSON, JSON-looking strings, and null remain distinct; unsupported error/value representations stay omitted and ineligible.
  • Reliable separate start/terminal ordering, explicit omissions/missing terminals, and a drained/fenced capture end. No FIFO, parent-ID, input-equality, object-reference assumptions, or script-output inference may manufacture provenance.
  • Observation failure that leaves tool inputs, results, errors, and permissions unchanged, with a demonstrated channel that evaluated code cannot use to forge eligible records.

Signing/HMAC and exact wire names remain proposals. Runtime owners establish the supported boundary; Loom owns its observer and assertion consumer; this runner adapts transport/import/export to that demonstrated boundary. The attached Architect draft is not already accepted by the v1 parser: message IDs, actual runtime call IDs, and parent observer-invocation IDs require explicit version/interface agreement, not silent relabeling.

Preserve the two empirical reports separately

The newly supplied Loom Worker addendum reports native/outer hooks but no inner-tool hooks on image 68ef7322…, with only inner names/inputs/status in outer metadata. PR #41's separate probe at 16394ac9672b5c3cff687c6360402c2fe6f7410e reported some inner callbacks, shared IDs, missing thrown-error terminals, and a later output mutation.

These are distinct experiments. Do not erase either result or declare their different callback coverage reconciled. Runtime-owner reproduction must compare exact producer/tool registrations, loaded configuration, runtime paths, and bytes. Neither report establishes the full supported boundary.

The supplied contract is draft revision 0.1 plus a later empirical addendum; its prototype files were uncommitted at inspection. Attachment SHA-256: 1a6e55024a64049e6aabde0716b38adcc2590e738232ea26b52a3c8af0b1dfab. Loom commit 6e255092388a57f141e609fee954cb7ae5977d4b must not be presented as pinning those prototype bytes.

Release and preserved-smoke gate

Before any positive capture claim:

  1. Obtain the runtime owner's supported boundary/version and pin the actual Loom producer, smoke fixtures, consumer, and runner checkout. Preserve scripts/fixtures/eval-tool-observer.ts, scripts/fixtures/eval-observer-smoke-tools.ts, and scripts/test_eval_observer_image.py at their exact bytes; do not replace Loom's smoke with the runner diagnostic.
  2. If bundled runtime bytes change, build and publish the resulting image and record its new immutable registry digest, actual OpenCode version, and runtime source revision. A repository commit, mutable tag, or local image ID is not the deployed digest.
  3. Rerun Loom's preserved provider-free smoke against that exact digest and PR checkout, through the actual producer → host CLI/exporter → Loom assertion consumer. Record loaded configuration and all revisions. Use disposable HOME/XDG/workspace state, no installation-wide database, and no real-provider inference.
  4. Include identical overlapping calls with reversed completion, genuine child actor/input/parent binding, discarded/transformed returns, caught throws, later output changes, malformed/forged/interrupted capture, and observation/I/O-failure noninterference. Missing evidence remains non-PASS. Independent code review remains a separate requirement.

No runtime/image bytes changed, no release image was published, and no new smoke execution is claimed in this coordination update. Existing artifacts and image pin remain unchanged.

Patch the existing OpenCode dispatcher locally instead of waiting for an
upstream PR. Carry unique child IDs, observe decoded executable inputs and
post-conversion terminal outcomes, and expose an opt-in read-only snapshot
hook for Loom's producer.

Add source and published-image diagnostics plus an experimental image build.
Keep default images unchanged and do not admit the new hook as authenticated
evidence. Protected export, native capture, preserved Loom smoke and
independent review remain separate integration requirements.
@bateau84 bateau84 changed the title feat: export fail-closed execution observer evidence on the host feat: add a local Code Mode observation seam and fail-closed export Oct 2, 2026
Keep untrusted tool code in a separate container with no capture mount or
runtime process access. Reuse execute.observed, redact before persistence,
and import a run-bound, fully accounted stream on the host without HMAC.

Add a provider-free observe command and actual-connection fault tests.
Native, delegated-session and in-process untrusted-plugin coverage remain
unsupported; this scoped path is not full Loom handoff acceptance.
…tity

Disable the script network extension only in the explicit protected profile,
with capture on or off. Carry the runtime invocation ID and dispatch ordinal
to the trusted remote adapter without changing tool input or parent call ID.

Build a distinct eval.2 downstream runtime; do not update default images.
Use the eval.2 runtime for the isolated direct-session profile. Keep the
collector mount and immutable adapter outside target tools, disable target
logs, bind image and input snapshots, and check actual Docker image IDs.

Carry runtime invocation metadata through the remote tool transport without
changing tool inputs. Exercise the public CLI and actual returned streams
with independent fixture observations, spoofing, replay, deletion, I/O
failure and redaction checks. Version the new launch-bound projection
explicitly; native, delegated-session and in-process plugin scope stays open.
Record the 33-check actual connection run and immutable experimental image
digests separately from full Loom acceptance. Explain the required remote
execution adapter, versioned consumer and preserved smoke without implying
that the current in-process Loom plugin is already isolated.
@bateau84 bateau84 changed the title feat: add a local Code Mode observation seam and fail-closed export feat: protect runtime observations with isolated tool execution Oct 3, 2026
Retain the concurrent isolated-launch implementation and its documented
33-check baseline. Require a separate receipt for received capture, reject
resealed-result forgery, validate execution snapshots, block readiness
redirects and redact script diagnostics before retention.

Document the v4 receipt profile, precise Loom integration work and prior
version-3 evidence separately. Signing/native/delegation remain open.
@bateau84 bateau84 changed the title feat: protect runtime observations with isolated tool execution feat: isolate runtime evidence collection and verify host receipts Oct 3, 2026
Run the pinned original-image diagnostic as an explicit negative control.
Require all 12 expected checks and blocked/ineligible capture before a zero
CI exit; missing checks, unexpected eligibility and wrong images still fail.
Keep the default diagnostic exit 4 and the real protected-channel tests.

Add ten regression tests for the exit policy and clarify that a passing
rejection test is not capture acceptance. No runtime or importer changes.

bateau84 commented Oct 3, 2026

Copy link
Copy Markdown
Owner Author

CI correction — 17877f6c458fd7e8d2537227100ca8186c7ae532

The current receipt-based implementation was already on this PR at a84260c. The failing check was Observer boundary integration, not the protected-channel implementation: run 37113830701 completed all 12 old-image diagnostic checks successfully, then exited 4 by design. I had left that historical counterexample wired as a permanently failing PR job.

The pushed fix makes CI explicitly run --expect-unsupported-baseline. That mode passes only when every required old-image check is present and true, the immutable baseline image matches, capture remains BLOCKED, and no capture is eligible. Unexpected acceptance or any failed/missing check still fails. Normal diagnostic invocation still exits 4. No continue-on-error, arbitrary exit-code suppression, change to importer eligibility, or historical result rewriting was added.

Added 10 exit-policy regression tests; all 112 Python tests passed locally with disposable HOME/XDG state. The four published file blobs match the local files. The real protected-channel tests remain separate and enabled. No runtime/collector/image source changed in this commit.

New head workflows have started; this comment does not yet claim green CI. PR remains draft and unmerged. Full Loom composition, native/delegated coverage, Cosign and independent review remain open.

### Background
Assessment found that invalid captures discarded their validated start and
terminal counts, reporting zero missing calls. Parent and child invocation
maps also allowed the same ID to be reused across the two namespaces.

### Changes
Preserve verified-prefix diagnostic counts without admitting invalid records,
leave unverified missing-call accounting unknown, and enforce capture-wide
parent/child ID uniqueness. Malformed terminal fields no longer settle a call.

### Verification
Five new parser regression cases failed on the prior implementation and pass
with the fixes. Full Python suite: 117 tests passed in disposable HOME/XDG
state. These are parser tests, not new runtime-origin or Loom acceptance proof.
### Background
The two new workflows executed PR-controlled build/test code in jobs with
package-write authority and later registry credentials. Build residue must
not share the publisher's process or credential boundary.

### Changes
Use read-only build and verification jobs, with an intermediate fresh
publisher that has no checkout and only loads/pushes Docker image data to
fixed experimental destinations. Transfer artifacts within the same run
and test the published immutable digests on a read-only runner.

### Verification
Publication-boundary regression checks reproduced the original permission
failure and pass after separation. Workflow YAML parses successfully; all
external actions are pinned. Actual split-job publication and image checks
remain to be run by CI; publication is not signer or safety approval.

bateau84 commented Oct 3, 2026

Copy link
Copy Markdown
Owner Author

Reviewer Assessment — recheck

Reviewed baseline 17877f6; fixes are now in d34fa82 and b6536ad.

Scoped conformance verdict: PASS. Open implementation findings from this Assessment: 0. This is a standalone role-guided review by the same assistant, not a separate independent reviewer or a native Loom gate/authority action.

Finding Mitigation Verification
R1 / MAJOR: build/test execution shared package-write/publishing authority Both new workflows now separate read-only build/test, a fresh no-checkout/no-execution publisher, and read-only verification of published digests. Declared-boundary regression tests plus successful actual build → publication → verification jobs.
R2 / MAJOR: invalid captures reported zero missing calls Validated-prefix counts remain diagnostic; unknown missing-call accounting stays null; invalid terminal fields do not settle calls. New counterexamples fail on the original importer and pass after the fix. Invalid records remain unadmitted.
R3 / MINOR: parent/child invocation-ID collision IDs must be unique across both parent and child maps. Both directions of reuse rejected; legitimate captures still pass.

Local full suite: 119 Python tests passed, using disposable HOME/XDG state. All five new parser regression cases also failed against preserved original bytes, confirming that they detect the defects rather than merely exercising code.

All four workflows completed successfully on b6536ad68e30f5423113967f1398db8f9d4927f9: CI, old-image negative control, protected connection, and runtime build/probe.

Methods read: Reviewer plus software-engineering, Python-testing and GitHub-workflow Assessment companions at Loom 9a13cc634bee3d2a969b1fb3f814692eb6c68ad1. No loom_assessment, OKF validation, separate-agent execution, or full Loom acceptance is claimed. Critic QA and Product Readiness follow separately. The PR remains draft and unmerged.

### Background
QA found that an image-declared VOLUME could create anonymous Docker storage
before the post-run mount check rejected it. Container cleanup did not ask
Docker to remove that storage, contrary to the disposable-state profile.

### Changes
Reject declared image volumes during digest inspection before launching any
container, retain the fixed policy rejection reason, and remove anonymous
volumes with this launch's own containers during cleanup.

### Verification
Three focused launch-policy tests pass, including the rejected-image path
and unchanged volume-free inputs. The prior preflight and diagnostic-code
counterexamples failed before the fix. Tests substitute only Docker I/O;
they are not new kernel/container-isolation proof.
### Background
QA found that the importer admitted values above the producer's documented
16 KiB bound. The Assessment accounting fix also changed a consumer-visible
field from zero to unknown without explicitly revising the projection.

### Changes
Measure sanitized values as compact UTF-8 JSON with the 16 KiB bound and
distinguish undersized framing from truncation. Version the projection as
v5, with explicit accounting_complete and nullable missing-call counts.
Update the real connection assertion and Loom migration instructions.

### Verification
Limit and version regressions failed on the prior implementation and now
pass. Full local suite: 127 Python tests passed in disposable HOME/XDG
state; compilation and diff whitespace checks passed. Actual image-based
verification is required separately; no new Loom or signing claim is made.

bateau84 commented Oct 3, 2026

Copy link
Copy Markdown
Owner Author

Critic QA — recheck complete

Head: 584795ff4787affc28f3f94b741e7c4f89301d1d.

PASS for the explicit restricted-profile implementation; confidence SOUND. Open findings from this QA pass: 0. This is a standalone role-guided review by the same assistant, not a separately independent reviewer or a native Loom gate.

Finding Fix and recheck
Q1 / MAJOR: image-declared volumes could create persistent state before the post-run check ce7bd9e rejects inherited volumes before launch, retains the safe policy reason, and removes anonymous volumes belonging to this launch's containers. Focused launch-policy tests exercise the real code with Docker I/O substituted; this is not an additional kernel-isolation proof.
Q2 / MINOR: importer admitted fields above the 16 KiB contract and confused short framing with clipping 584795f enforces compact UTF-8 JSON size, retains the exact boundary value, rejects overflow, and reports short framing without a false size-truncation claim.
Q3 / MINOR: corrected accounting changed the consumer contract without a new version Projection v5 explicitly adds truthful nullable accounting and accounting_complete; the actual connection test and Loom handoff now require v5. Runtime/wire/receipt profiles are unchanged. Historical v4 artifacts are untouched.

The new defect cases were observed failing before their fixes. Final local and CI Python suites: 127 tests passed. Python compilation and diff checks passed. Reviewer R1–R3 remain mitigated.

All four workflows passed on this head: CI #154, old-image rejection regression #12, protected channel #8, and runtime build/probe #9.

Downloaded protected-channel evidence: all 37 real-connection checks passed; legitimate public-CLI capture contains eight eligible v5 records; resealed forgery, replay and rehashed deletion admit no records. All 47 archived source files match the locally tested checkout bytes. ZIP SHA-256: 615a9fd39485097a62635402564bd0567f8fee3ddb1cb0f3285a12a1fed4a8c6.

New tested protected image:
ghcr.io/bateau84/opencode-eval-runner@sha256:c2511bd114e2ad83dbbed019e0f2a6ee46898c8dcd4086a5156259a5a07279bb

Adversarial fixture tool image (not actual Loom):
ghcr.io/bateau84/opencode-eval-runner@sha256:b0609b966f2ff83bc658b35af8f464a654d24008f4da3597ef97394ecbfe99f8

The publisher now runs separately from builds/tests, and published digests are tested on a fresh read-only job. No signing or approved release identity is implied. Runtime/collector source bytes were not modified by these review fixes; host admission, workflow boundaries, tests and migration documentation were.

Methods: Critic and software-engineering/GitHub-workflow QA companions at Loom 9a13cc634bee3d2a969b1fb3f814692eb6c68ad1. Full Product Readiness is assessed separately; native/delegated coverage, actual Loom composition and Cosign are not marked complete by this QA result. No squash or merge performed.

bateau84 commented Oct 3, 2026

Copy link
Copy Markdown
Owner Author

Product Readiness — HOLD

Assessed 584795ff4787affc28f3f94b741e7c4f89301d1d after the Reviewer and Critic QA rechecks. This is not a passing full-product gate. Confidence: SOUND in the HOLD verdict.

The six implementation findings from Assessment/QA are mitigated. The restricted runtime-to-host connection is supported by the actual public CLI, 37 connection checks and green CI. Those results do not establish the entire original handoff.

Open product-release obligations

ID Remaining blocker Required owner/action
P1 / MAJOR The real Loom producer/tool execution → runner → actual Loom predicates has not been demonstrated. Native/delegated coverage remains explicitly excluded; fixture-only redaction is not a general live-eval policy. Loom + runner: pin the preserved smoke and actual implementations, preserve permissions/identity, integrate the explicit v5 profile and run the real composition. Do not replace Loom behavior with fixture copies or relabel isolated calls as native ones.
P2 / MAJOR Requested Cosign image signing and expected approved-build signer verification are not implemented. image_signatures_verified remains false. Runner/release owner: implement the image signing/verification policy outside evaluated execution and exercise wrong-signer/changed-image rejection. The discussed input/result signing layouts remain proposals, not silently imposed or completed requirements.
P3 / MAJOR These role-guided passes used the requested Reviewer/Critic definitions, but the same assistant authored the fixes. They are not the separate independent review required for full acceptance. Independent implementation/security review of the pinned final diff and actual evidence.

Readiness attack

The plausible failure is a consumer accepting evidence_eligible: true while ignoring the restricted profile, full_handoff_eligible: false, unsigned-image state or missing native/delegated coverage. That could turn a correct narrow fixture result into an incorrect full-Loom PASS. The current actual Loom consumer is not present in this proof, so documentation and green runner tests cannot close that gap.

The valid runner work remains intact. HMAC is not necessary for the implemented protected path; adding a signing oracle or another interpreter would not close these obligations. Host/runtime/kernel and authorized workflow editors remain the explicit trusted base.

No existing artifacts were backfilled, no installation-wide database was accessed, and no real-provider inference was used. No native Loom critic-final/loom_complete action or independent approval is claimed.

Disposition: keep draft and unmerged. P1–P3 remain open, not waived or called mitigated. The updated Loom handoff preserves the exact integration and version requirements. No squash or merge performed.

@bateau84 bateau84 changed the title feat: isolate runtime evidence collection and verify host receipts feat: add protected Code Mode evidence collection Oct 3, 2026

bateau84 commented Oct 3, 2026

Copy link
Copy Markdown
Owner Author

Handoff: preserve actual Loom host semantics in protected capture

Loom’s observer/parser/smoke checkpoint is committed at f8439e4. Preserve these tests and historical results.

Current limitation

The codemode-inner/direct-session/v1 profile is intentionally restricted. It does not support normal multi-turn invoke, native observation, delegated-session capture, or the in-process Loom plugin.

Putting Loom executor closures behind /call does not automatically preserve:

  • Actual authenticated sessions and executing roles.
  • OpenCode and Loom permission enforcement.
  • Grants, child attachment, independent gates, and OQ continuation delivery.
  • Cancellation, lifecycle, and final caller-facing result conversion.

Do not substitute fabricated session responses, permissive permission callbacks, or successful continuation stubs.

Requested outcome

Investigate and propose the smallest supported protected-capture path that retains actual OpenCode/Loom execution semantics for the existing eval harness.

The realization is yours to determine. Please establish feasibility before requiring Loom to build a remote-host adapter.

The path must provide:

  1. Actual native and Code Mode input/final-result/error observation.
  2. Genuine actor/session/message/parent identity, including delegated sessions.
  3. Reliable invocation correlation and separate start/completion ordering.
  4. Protected collection that evaluated code cannot forge.
  5. Disposable state, without access to the installation-wide Loom database.
  6. Normal permission enforcement and explicit unsupported coverage.

The direct-session profile may remain a supplemental smoke path, but cannot replace these obligations.

Contract correction

At c1629415…, protected-channel.md describes wire v1/projection v2, while implementation uses wire v2/projection v3 and dispatch context. Reconcile the pinned documentation and implementation.

Division of responsibility

  • Runner/runtime: supported observation, protection, transport, execution profile, and immutable image.
  • Loom: launcher/consumer adaptation and tests against actual Loom implementation.
  • Joint: reviewed interface, loaded revision/image proof, and end-to-end composition tests.

Return a feasibility assessment, supported/unsupported matrix, and exact Loom-facing interface. If preserving these semantics requires a broader project, identify that boundary rather than implementing it implicitly.

Signing remains separate. No merge or acceptance is authorized by this handoff.

Next step: feasibility. Do not turn this into a general remote OpenCode platform project.

### Background
The prior protected profile proved an isolated direct-session subset by moving
evaluated tools behind a remote adapter. Loom's real eval harness must instead
keep bun run eval:live, its existing cases, normal invoke, sessions, permissions
and plugin behavior intact.

### Goal
Establish feasibility underneath the existing runner invoke entrypoint without
silently replacing Loom's host behavior.

### Reasoning
Native/direct tools and Code Mode inner calls already cross real runtime-owned
boundaries. Expose those semantics diagnostically on the normal patched image,
but keep positive protection blocked because arbitrary Loom plugin code shares
the OpenCode process and therefore shares collector authority.

### Changes
- build a semantics-preserving 2.0.18-eval.3 image with apply.py only;
- add native Tool-service observations and bind inner calls to the real outer
  execute observation on one runtime sequence;
- test the public opencode-eval-runner invoke path with a provider-free fixture;
- document the supported/unsupported matrix, fixed eval:live interface and the
  broader plugin-isolation boundary;
- correct the historical c162 wire/projection documentation.

### Verification
CI must rebuild/publish the experimental image, run source tests, the runtime
probe and the normal-invoke eval:live compatibility probe. No merge, Loom
acceptance, protected in-process producer or signing claim is made.
The eval.3 image intentionally enables diagnostic runtime observations underneath
normal invoke. Replace the stale opt-in-only assertion with a check that the
image-owned seam is present while product results remain unchanged.

Verification: rerun the published-image runtime and eval:live compatibility probes.
The compatibility image owns native host observation, while the probe still
toggles the older inner Code Mode seam for noninterference coverage. Assert
those controls independently instead of requiring both to follow one flag.

Verification: rerun the normal-invoke image and eval:live compatibility probes.

bateau84 commented Oct 4, 2026

Copy link
Copy Markdown
Owner Author

Latest Copilot disposition — current head 03339e5

The high-severity initializer-provenance finding was accepted and fixed:

  • container/__init__.py now has a build arg, OCI label, in-image SHA check, host-side digest comparison, and evidence_load provenance binding.

Copilot's subsequent medium workflow-test findings were also accepted and fixed:

  • signer escape regression now catches the original single-backslash syntax;
  • quoted workflow job IDs are parsed;
  • unsupported/duplicate top-level job keys fail closed;
  • nested mappings such as with: are not misclassified as jobs;
  • evidence-safety.yml is covered by package-authority checks;
  • evidence-safety.yml and observer-integration.yml are covered by action-pinning/no-ignored-error checks.

Current head: 03339e592433ef7edfc07736d57482efc51cdd17.

All five workflows are green:

GitHub currently reports 0 unresolved review threads. A final Copilot re-review of this exact head has been requested but has not posted yet.

@bateau84
bateau84 marked this pull request as draft October 4, 2026 22:43

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The documented integration pins an older image that cannot satisfy the current source-hash handshake.

Review effort: Balanced
Findings: 1 High severity

Open (1)
Resolved since last review (1)

Comment thread runner/safe_invoke.py
Podman may render Image.Id as bare 64-hex while Docker renders sha256:<64>.
Normalize both valid forms to canonical evidence_load.image_config while
executing the exact engine-returned config ID, never the mutable image tag.

Keep immutable RepoDigest, initializer/module/invoke hashes, source revision,
and safety labels unchanged. Reject malformed/uppercase/non-SHA256 IDs.

Add Docker/Podman unit regressions and an actual provider-free Podman preflight
against the published safety image. No defaults, inventory, ACK, or capture
acceptance are changed.
@bateau84
bateau84 requested a balanced review from Copilot October 5, 2026 09:23
@bateau84
bateau84 marked this pull request as ready for review October 5, 2026 09:23
@bateau84
bateau84 marked this pull request as draft October 5, 2026 09:25

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Code Mode calls are emitted as native observations when only the inner observation flag is disabled.

Review effort: Balanced
Findings: 1 High severity

Open (1)
Resolved since last review (1)

Comment thread runtime-patches/apply.py Outdated
Code Mode calls were classified as native whenever the inner observation callback
was disabled, because native detection overloaded observedInput absence as the
dispatch marker.

Carry an explicit codeModeDispatch boolean into the central Tool execution path.
Native observations are now created only for genuine native/direct calls,
independent of OPENCODE_EVAL_OBSERVATIONS. Add an actual-image regression proving
that disabling inner observation leaves only the real native tool and outer
execute call on the native stream.

Bump the patched runtime identity to OpenCode 2.0.18-eval.5 because this changes
runtime semantics. No merge or default-pin change.
The normal-invoke runtime now carries an explicit Code Mode dispatch marker so
disabling inner observation cannot misclassify Code Mode tools as native. Rebase
the opt-in evidence-safety image onto the tested eval.5 digest so Loom's RSP
composition uses the corrected runtime as well.

No default pin or merge change.

bateau84 commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Superseded by #45.

The threat model has been corrected: normal Loom evaluation uses an explicitly trusted checkout and requires authoritative runtime observation, evidence fidelity, completeness, and pre-persistence credential protection. It does not require resistance to a deliberately malicious plugin sharing the trusted runtime.

Useful implementation/research from this PR will be selectively reused in #45. This PR is being closed rather than merged so its stronger hostile-runtime architecture does not become the baseline.

The branch is intentionally retained as reference/provenance.

@bateau84 bateau84 changed the title feat: preserve eval:live semantics while observing runtime tools [SUPERSEDED by #45] feat: preserve eval:live semantics while observing runtime tools Oct 6, 2026
@bateau84 bateau84 closed this Oct 6, 2026
bateau84 added a commit that referenced this pull request Oct 7, 2026
## Integrated direction

This PR now integrates the completed child work from #46–#51 into one
trusted-checkout runtime-evidence implementation on stock OpenCode
**2.0.23**.

The normal Loom eval profile remains a trusted-checkout evaluation. This
PR does **not** add an OpenCode patch/fork, remote PluginHost, protected
channel, HMAC/signing boundary, capability broker, or hostile-plugin
isolation.

## Canonical public contract

The only authoritative public runtime-evidence object is:

`opencode-eval-runner/runtime-evidence/v1`

Raw observer records are internal adapter input only.

Existing `tools`, `actions`, `tool_result_evidence`, stdout/stderr,
Session/model text, and workspace files remain convenience/diagnostic
data and are not promoted into runtime evidence. Final cleanup #52 also
removes competing runtime `evidence_eligible` signals from the
diagnostic tool-result/safety projections.

There is one canonical builder/validator that owns:

- `complete | incomplete | unsupported | invalid`;
- capture closure and unknown coverage;
- missing terminals;
- observer/callback loss;
- timeout/interruption;
- duplicate/ambiguous invocation and sequence detection;
- identity-based concurrency;
- per-boundary support;
- assertion-scoped boundary/field eligibility.

Unknown coverage is never converted to zero.

## Boundary semantics

The v1 contract exposes:

- `native`
- `code_mode_execution`
- `code_mode_finality`

Overall `complete` means the supported capture boundaries are complete.
It does **not** mean every possible assertion is supported.

This allows:

```text
overall: complete / eligible
native: complete / eligible
code_mode_execution: complete / eligible
code_mode_finality: unsupported / ineligible
```

An assertion that only requires `native` evidence can therefore remain
eligible even though Code Mode finality is unsupported. Assertions that
require a redacted/omitted/unsupported exact field remain ineligible for
that field.

## Native observation

The integrated observer preserves #51's stock-2.0.23 boundary:

```text
decoded executable input -> transformed tool.execute wrapper

terminal success/error ->
  session.tool.success
  session.tool.failed
```

It records opaque invocation identity, actual tool, agent, Session,
message, real CallID, executable input, Session ancestry, terminal
result/error, and monotonic start/terminal ordering.

Correlation is identity-based. It does not use FIFO, input equality,
tool name, or completion order.

The raw former `native_tool_observations` projection is no longer a
public result.

## Code Mode

The #50 stock-runtime result is retained faithfully.

Supported facts:

- unique per-inner invocation identity;
- actual effective tool;
- decoded/executable input;
- Session/message/agent;
- actual outer `execute` CallID;
- binding to the outer invocation;
- start/handler-terminal ordering;
- identical concurrent calls and reverse completion.

The transformed handler's earlier value/error is **not** promoted as
caller-final evidence.

Exact final script-visible value/error is represented as:

```json
{
  "state": "unsupported",
  "reason": "stock_codemode_final_boundary_not_exposed"
}
```

> Stock OpenCode 2.0.23 does not expose a supported boundary that proves
the exact final value/error seen by a Code Mode script for each inner
call. That assertion is reported as unsupported.

## Evidence safety

The #47 safety behavior is applied to authoritative dynamic values
before their first observation persistence/output sink:

```text
raw value in observer memory
  -> sanitize/redact/omit
  -> size decision
  -> internal capture
  -> validate/account
  -> result serialization
  -> stdout/host-file persistence
```

Public field states are only:

- `available`
- `redacted`
- `omitted`
- `unsupported`

Product `exit_code`, timeout, and product success/failure remain
separate from evidence eligibility.

## Child PR disposition

| PR | Disposition |
| --- | --- |
| #46 fail-closed runtime evidence accounting | **Incorporated /
superseded as a separate implementation.** Its stronger accounting is
folded into the canonical v1 builder/validator. No second
public/accounting status engine remains. |
| #47 evidence safety | **Incorporated.** Pre-sink sanitization,
redaction/omission, size ordering, credential inventory, and output
guard behavior are retained. Internal `exact` terminology is not part of
the public runtime-evidence schema. |
| #48 provider-free acceptance gate | **Incorporated and tightened.** It
now requires the exact final v1 contract rather than rollout-compatible
alternate shapes. |
| #49 runtime evidence v1 contract | **Incorporated and evolved.** It
remains the public wire contract; overall vs boundary/assertion
eligibility was revised to preserve partial support. |
| #50 stock Code Mode experiment | **Incorporated as capability
proof/diagnostic probe.** Supported execution facts are used; exact
final caller value/error remains explicitly unsupported. |
| #51 stock native observer | **Incorporated / superseded as a public
projection.** Its stock observer boundary is retained as internal
adapter input to `runtime_evidence`; no public
`native_tool_observations` object remains. |

Child PRs #46–#51 have no remaining implementation authority after this
integration. Final cleanup #52 removes parallel-work duplication without
adding architecture. Explicit unsupported areas are: Code Mode
caller-final value/error on stock 2.0.23, OpenCode runtime observation
for the `github-copilot-cli` transport, and protection against a
deliberately hostile plugin sharing the trusted OpenCode process.

## Final integration cleanup

#52 is the final child PR targeting this branch. Its head
`a3f575d78f4b14d62b76d14dceea49b76496fd2d`:

- removes duplicate runtime-evidence constants and dead `truncated`
accounting;
- removes competing `evidence_eligible` signals from diagnostic
`tool_result_evidence` / safety projections;
- removes the unused legacy tool-result clipping helper;
- confirms PR #41 protected-runtime/signing/broker machinery is absent;
- documents Code Mode finality, Copilot runtime observation, and
hostile-plugin isolation as explicit unsupported areas;
- keeps stock OpenCode 2.0.23 and normal `invoke` behavior.

Validation on #52:
- CI #297 / run 37521327070: **PASS**, 93/93 unit tests;
- provider-free runtime evidence acceptance #11 / run 37521327281:
**PASS**, 6/6 helper tests and all 10 scenarios;
- stock native observer integration #23 / run 37521327205: **PASS**.

#52 has now been squash-merged into this branch as
`b10c52d50aaf0c038e77e80eaca3890347d5fa82`.

## Final-head validation

Head: `b10c52d50aaf0c038e77e80eaca3890347d5fa82`

- **CI #296 / run 37519197332 — PASS**
  - Python/unit suite: **92/92 PASS**
  - stock Code Mode probe: **PASS**
  - exact caller terminal capability reported unsupported: **PASS**
  - OpenCode plugin activation preflight: **PASS**
  - config-root plugin dependency preflight: **PASS**
  - pinned OpenCode remains **2.0.23**
- **Provider-free runtime evidence acceptance #10 / run 37519197280 —
PASS**
  - helper/unit checks: **5/5 PASS**
  - `native_success`: PASS
  - `native_error`: PASS
  - `code_success`: PASS
  - `code_caught_error`: PASS
  - `concurrent_reverse`: PASS
  - `delegation`: PASS
  - `timeout`: PASS
  - `interrupted`: PASS
  - `redaction`: PASS
  - `collector`: PASS
- **Stock native observer integration #22 / run 37519197390 — PASS**
  - stock 2.0.23: PASS
  - complete capture: PASS
  - provider-free probe: PASS

## Review state

The integration cleanup from #52 is merged. PR #45 now contains the
final integrated implementation and remains ready for final independent
review.

PR #45 remains open/draft and is **not merged**.


## Independent final review

Final review of head `b10c52d50aaf0c038e77e80eaca3890347d5fa82` found
one blocking correctness defect: stock Code Mode's synthetic
model-facing `execute` invocation was not emitted as an authoritative
native observation, so native coverage could falsely remain
complete/zero while Code Mode had run and inner parents resolved only to
an internal token.

Corrective child PR #54 observes the real outer call at stock
`execute.before`, binds inner parents to that observed invocation, and
makes dangling/wrong-identity parents invalid. #54 is green across CI,
the stock native probe, and the 10-scenario provider-free acceptance
gate.

**Review status: NOT READY until #54 is merged.** No threat-model
expansion is required.
@bateau84
bateau84 deleted the feat/trustworthy-execution-observer-export branch October 7, 2026 08:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants