Skip to content

[SUPERSEDED by #45] docs: plan TRUST-001 experiment on stock OpenCode - #43

Closed
bateau84 wants to merge 35 commits into
mainfrom
planning/trust-001-stock-opencode
Closed

bateau84 wants to merge 35 commits into
mainfrom
planning/trust-001-stock-opencode

Conversation

@bateau84

@bateau84 bateau84 commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Status

Planning-only documentation for the TRUST-001 bounded experiment.

This PR does not authorize or contain candidate implementation, candidate OCI execution, model inference, provider credential use, OpenCode patches/forks/custom binaries, default-pin changes, or production adoption.

Hard boundary

Wave-1 package

This PR persists:

  • stock OpenCode 2.0.23 capability assessment;
  • planning evidence provenance;
  • provisional TCB;
  • effective-authority manifest;
  • bounded capability manifest derived from Loom 149406d;
  • callback/request lifecycle;
  • scope/completeness contract;
  • first-sink inventory;
  • revised architecture recommendation and experiment plan;
  • independent Gate-1 architecture review.

Current gate

Gate 1: PASS

The complete planning package was reviewed as a whole and corrected before the verdict. The material corrections include:

  • pinned Loom legacy storage is generation-lifetime bounded product access, not setup-only;
  • capability and evidence channels have distinct authority and peer-integrity requirements;
  • stock OpenCode plugin loading is bounded by an immutable plugin-source closure manifest;
  • cancellation does not overclaim rollback of already-completed remote product side effects;
  • completeness requires trusted admission close plus a contiguous final collector sequence and generation seal;
  • candidate-controlled diagnostic/log/export paths are included in first-sink confidentiality;
  • stock ctx.event.subscribe() is source-supported for access, while ordering/drain remain unproven;
  • provenance classifications were normalized without promoting future experiment claims.

A Gate-1 PASS only makes a separate owner Authorization A eligible for consideration. It does not authorize construction or execution.

Known load-bearing open questions

  • Code Mode per-inner final caller result/error remains UNPROVEN under stock APIs.
  • evidence-channel resistance to stock shell subprocess authority remains UNPROVEN until provider-free construction/preflight.
  • capability-channel peer authenticity/integrity must be proven against shell/Loom child impersonation.
  • synchronous agent-transform replay fidelity remains bounded/UNPROVEN.
  • trusted root Session admission, public-event ordering/correlation, and generation seal/drain require provider-free proof.
  • remote cancellation side-effect fidelity remains UNPROVEN.
  • plugin-source closure and all candidate-controlled first sinks must be verified against the constructed checkpoint.

bateau84 commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Superseded by #45.

The threat model has been corrected: normal Loom evaluation uses an explicitly trusted checkout and requires authoritative runtime observation, evidence fidelity, completeness, and pre-persistence credential protection. It does not require resistance to a deliberately malicious plugin sharing the trusted runtime.

Useful implementation/research from this PR will be selectively reused in #45. This PR is being closed rather than merged so its stronger hostile-runtime architecture does not become the baseline.

The branch is intentionally retained as reference/provenance.

@bateau84 bateau84 changed the title docs: plan TRUST-001 experiment on stock OpenCode [SUPERSEDED by #45] docs: plan TRUST-001 experiment on stock OpenCode Oct 6, 2026
@bateau84 bateau84 closed this Oct 6, 2026
bateau84 added a commit that referenced this pull request Oct 6, 2026
## Scope

Focused corrective child of #45, targeting
`refactor/trusted-checkout-evidence`.

The base now includes merged #54 (`b10c8e9`), which closes the
authoritative outer Code Mode `execute` blocker. This PR is rebased
directly on that commit and contains only the remaining
authoritative-capture transport correction plus its regression coverage.

Stock OpenCode remains **2.0.23**. No threat-model expansion.

## Blocker 1 — outer Code Mode `execute`

**Closed in the base by #54 and revalidated here.**

The current integrated runtime evidence:

- observes the real synthetic outer `execute` at the supported stock
`execute.before` boundary;
- records Session, agent, message, real CallID, supported hook input,
and a unique runtime invocation ID;
- uses stock `session.tool.success` / `session.tool.failed` as terminal
authority;
- binds every inner Code Mode call to that exact observed outer
invocation;
- rejects dangling or identity-mismatched parents;
- keeps exact Code Mode caller-finality explicitly unsupported.

The Code Mode probe remains green after the capture-transport change.

## Blocker 2 — remove target-writable authoritative capture

The production observer no longer uses:

`/tmp/runtime/runtime-observer.jsonl`

Authoritative observer records now cross the process boundary through a
**runner-owned one-connection loopback stream**:

1. the runner binds an ephemeral listener on `127.0.0.1` immediately
before the real OpenCode invocation;
2. the trusted observer connects during plugin startup;
3. the listener closes after accepting that one connection;
4. the observer removes the endpoint from `process.env` before evaluated
tool subprocesses run;
5. sanitized observer records flow over the established stream;
6. the runner drains and validates those bytes through the existing
canonical builder/validator/accounting path.

A direct inherited memfd/FD was tried first, but stock
`@opencode/cli@2.0.23` crosses an internal process boundary that does
not preserve arbitrary extra descriptors. The one-connection stream
keeps stock OpenCode unchanged and avoids filesystem authority without
introducing security-platform machinery.

Protection against a deliberately malicious same-process plugin remains
explicitly out of scope.

## Tamper regression

Provider-free acceptance includes an evaluated tool that spawns
`/bin/sh` and:

- creates the old capture path;
- deletes it;
- recreates it;
- appends a forged observer record.

The scenario passes only if the shell tamper completes, authoritative
evidence remains complete/eligible, the real tamper tool is observed,
and the forged invocation/tool never appears in `runtime_evidence`.

## Preserved

- `opencode-eval-runner/runtime-evidence/v1`;
- canonical builder / validator / accounting;
- pre-sink credential sanitization;
- stock Session terminal authority;
- Code Mode exact caller-finality remains unsupported;
- normal `invoke` behavior;
- assertion-scoped eligibility;
- stock OpenCode 2.0.23 only.

No signing/HMAC, protected channel, remote PluginHost, capability
broker, hostile-plugin isolation, or OpenCode patch is introduced.

## Validation

Current clean head is based directly on merged #54 and is mergeable.

- **CI #329 / run 37532684965 — PASS**
  - **95/95 Python tests**
  - stock Code Mode probe: **PASS** (`diagnostics_passed: true`)
  - OpenCode **2.0.23**
- **Provider-free runtime evidence acceptance #43 / run 37532684985 —
PASS**
  - all **11/11** scenarios:
    - native_success
    - native_error
    - code_success
    - code_caught_error
    - concurrent_reverse
    - delegation
    - timeout
    - interrupted
    - redaction
    - collector
    - capture_tamper
- **Stock native observer integration #55 / run 37532684878 — PASS**
  - stock OpenCode 2.0.23
  - capture complete
  - provider-free native observation

## Merge scope

This PR is ready for review against
`refactor/trusted-checkout-evidence`.

Do **not** merge PR #45 as part of this change.
@bateau84
bateau84 deleted the planning/trust-001-stock-opencode branch October 7, 2026 08:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant