diff --git a/docs/experiments/trust-001/ARCHITECTURE-RECOMMENDATION.md b/docs/experiments/trust-001/ARCHITECTURE-RECOMMENDATION.md new file mode 100644 index 0000000..5023015 --- /dev/null +++ b/docs/experiments/trust-001/ARCHITECTURE-RECOMMENDATION.md @@ -0,0 +1,92 @@ +# Architect recommendation — stock OpenCode TRUST-001 candidate + +Status: **Gate 1 PASS; recommended experiment candidate, not selected architecture**. + +## Recommendation + +Following independent Gate-1 review, the owner may consider a separate Authorization A to investigate: + +> A runner-owned trusted bridge plugin loaded by **stock OpenCode v2.0.23**, proxying the bounded Loom plugin API to an isolated Loom execution domain, while host-side runner code owns evidence collection, safety, scope accounting, and persistence. + +No OpenCode modification is permitted. + +## Why this candidate + +It preserves the existing normal runner entrypoint and stock OpenCode Tool/Session semantics while moving evaluated Loom module execution out of the trusted OpenCode process. + +It reuses a natural supported extension point—the stock plugin API—without requiring a new workflow engine, state API, permission engine, or OpenCode fork. + +## Important revision from the earlier candidate + +The candidate is **not** simply "put Loom in another process." + +The trust design must separately address: + +- stock shell subprocesses executing under the OpenCode runtime's OS authority; +- project/config plugin-loading escape paths; +- capability/evidence channel authority; +- synchronous transform fidelity; +- Code Mode inner finality. + +Process topology alone is not evidence of TRUST-001. + +## Proposed boundary + +### Trusted + +- host runner launcher/verifier; +- host collector/scope/safety/writer; +- stock OpenCode v2.0.23 core; +- runner bridge plugin; +- reviewed protocol/correlation implementation; +- kernel/OCI enforcement assumptions. + +### Untrusted for evidence authority + +- Loom module/dependencies; +- Loom callbacks/handlers/policy logic; +- Loom children; +- model/product data; +- built-in shell subprocesses; +- writable workspace; +- arbitrary external plugins. + +## No OpenCode patch fallback + +If stock OpenCode's supported APIs cannot establish a Loom contract requirement, the only valid outcomes are: + +- supported by runner-owned wrapping; +- `UNPROVEN`; +- `UNSUPPORTED`. + +"Patch OpenCode" is not an allowed resolution. + +## Why v2.0.23 + +Compared with v2.0.18 it exposes more Session lifecycle API, including parent Session creation, removal, compaction, and metadata updates. + +However the plugin loader, plugin hooks, Tool execution path, Tool runtime, and plugin supervisor relevant to TRUST-001 remain unchanged. v2.0.23 does not itself isolate plugins. + +## Candidate strengths + +- pinned Loom needs only a bounded subset of PluginHost; +- Loom moves most normal state to its own SQLite after setup, reducing capability-broker state surface, while bounded stock plugin-storage `get/set/scan` remains available for pinned Loom's lazy legacy-compatibility path; +- stock live Session events provide strong native ancestry/called/terminal facts; +- runner proxy Loom tools can allocate trustworthy child correlation without changing tool inputs; +- Loom semantics remain executed by Loom. + +## Candidate risks + +1. **Code Mode inner finality** — stock public metadata does not expose each inner final script-visible result/error. +2. **Mixed OS trust** — shell subprocesses share the OpenCode runtime domain. +3. **Plugin loading** — evaluated paths must not cause new in-process plugin imports. +4. **Synchronous transforms** — stock transform callbacks are synchronous/replayable. +5. **Callback/cancellation fidelity** — RPC separation must preserve shared mutable event behavior and late-response rules. +6. **First-sink confidentiality** — new bridge/capability/collector paths must not persist secrets before projection. +7. **Scope sealing** — apparent quiescence is not complete evidence until trusted admission is closed and the collector has drained a generation-scoped final sequence. + +## Recommendation status + +**Gate 1: PASS.** The corrected planning package is coherent and bounded enough for the owner to consider a separate Authorization A for construction and explicitly named provider-free preflights. + +This review does not grant Authorization A, prove feasibility, or authorize semantic/adversarial/model-backed execution. diff --git a/docs/experiments/trust-001/AUTHORITY-MANIFEST.md b/docs/experiments/trust-001/AUTHORITY-MANIFEST.md new file mode 100644 index 0000000..6a38c32 --- /dev/null +++ b/docs/experiments/trust-001/AUTHORITY-MANIFEST.md @@ -0,0 +1,115 @@ +# Effective-authority manifest + +Status: **proposed Wave-1 boundary; not yet experimentally verified**. + +## Principle + +Different processes/containers do not automatically create TRUST-001. Effective write/control authority is the criterion. + +## Proposed filesystem and mount authority + +| Resource | Host collector | Stock OpenCode domain | Isolated Loom domain | Requirement | +|---|---|---|---|---| +| evaluated workspace | no direct execution | requested RO/RW | same requested RO/RW where Loom semantics require it | product state only | +| runner/bridge source | host-owned | read-only | none | never workspace-writable | +| stock OpenCode executable/root | no mutation | read-only image | none | immutable stock v2.0.23 | +| Loom source/dependencies | provenance only | **not imported/executed** | read-only | module boundary begins before evaluation | +| OpenCode disposable HOME/XDG | no | private tmpfs/state | none | invocation-local | +| Loom HOME/XDG/state | no | none | private | invocation-local; Loom owns its local runtime state | +| collector/evidence storage | read/write | **not mounted** | **not mounted** | evaluated code cannot modify evidence | +| safety policy/inventory | host-private | delivered only to trusted bridge/collector through reviewed private path | no raw policy | no credential-policy disclosure | +| container-engine socket | host runner only if required | none | none | never exposed | +| evidence-channel bootstrap endpoint | host-created private runtime dir | bridge bootstrap only | **not mounted/reachable** | never under workspace; removed after admitted connection | +| capability-channel bootstrap endpoint | host-created private runtime dir | bridge endpoint only | Loom endpoint only | separate from evidence channel; removed after admitted connection | + +## Channel authority and bootstrap + +The design requires **two distinct channels with different authority**: + +1. **evidence channel** — trusted runner bridge → host collector. This is the only channel allowed to introduce candidate runtime observations into evidence authority. +2. **capability channel** — runner bridge ↔ isolated Loom. This transports product callbacks, requests, mutations, and results only. It is untrusted for evidence authority. + +The channels MUST NOT share an accepted descriptor, listener path, authentication material, or message type that the collector could interpret as evidence. The capability channel has no direct evidence-write authority, but its transport admission/correlation must bind every accepted response to the admitted Loom generation. Loss of that authenticity/integrity makes affected evidence ineligible. + +The candidate should avoid a reusable credential in `process.env`, because stock shell inherits `process.env` when no Session-specific environment overrides it. + +### Evidence-channel bootstrap + +1. runner creates a private Unix-domain endpoint outside the workspace; +2. only the stock OpenCode domain receives bootstrap reachability; the isolated Loom domain never receives the collector endpoint; +3. the bridge connects during activation **before inference/tool execution**; +4. host admits only the expected stock OpenCode process hosting the reviewed bridge, using runner-owned process/container identity plus kernel peer identity where available; +5. the listener is closed/unlinked immediately after the expected connection is admitted; +6. the established descriptor is close-on-exec and must not be inherited or duplicated by stock shell subprocesses; +7. reconnect is unsupported in the bounded experiment. + +### Capability-channel bootstrap + +1. runner creates a separate private endpoint outside the workspace; +2. only the expected bridge side and isolated Loom generation may establish it; +3. the listener is closed/unlinked after the expected peers are admitted; +4. this channel has no collector credential, evidence framing authority, or evidence-write capability; +5. duplication, impersonation, or integrity loss makes affected product work failed/unresolved and affected evidence ineligible; an unadmitted process may never be silently accepted as the Loom peer. + +Both bootstraps remain **UNPROVEN** until provider-free construction/preflight. They are specified here so Gate 1 reviews concrete authority separation rather than process topology alone. + +## Process authority + +Required candidate controls: + +- read-only container roots; +- non-root execution; +- dropped Linux capabilities; +- no-new-privileges; +- default or stricter seccomp; +- no host PID namespace; +- no Docker/Podman control socket; +- no evidence volume in evaluated domains; +- no bridge secret in process environment; +- core dumps disabled or guaranteed non-persistent/unavailable to evaluated code. + +The provider-free preflight must specifically test whether evaluated shell/Loom children can: + +- ptrace or use `process_vm_readv/writev` against the trusted OpenCode process; +- use `pidfd_getfd` or `/proc//fd` to duplicate bridge descriptors; +- access `/proc//mem`; +- discover a still-connectable collector listener; +- signal/kill trusted work. + +Killing/suppressing trusted work is not by itself a false-evidence violation: it must become incomplete/non-evidence. Reading/duplicating the trusted channel or creating eligible records is a rejection. + +## Plugin-loading authority + +Stock OpenCode builds an effective plugin-source operation set from configuration documents, configuration roots, auto-discovered `plugin/` and `plugins/` entries, configured package/local targets, and watched source/config changes. Protecting only the final plugin file is therefore insufficient. + +Before activation the candidate MUST build a **plugin-source closure manifest** that: + +1. records the exact effective plugin add/remove operation set; +2. enumerates every configuration document/root, discovery directory, configured local source, package declaration, symlink target, and watched input that can change that operation set for the tested profile; +3. admits only runner-owned trusted bridge code for in-process execution; +4. makes every operation-set input runner-owned and read-only/unreachable to evaluated write authority for the generation; +5. preserves required non-plugin project inputs separately; +6. treats any post-activation operation-set change, newly discovered source, or changed admitted source identity as a checkpoint stop/rejection, not a hot reload. + +A practical runner-owned approach may use read-only submounts/overlays for the exact source/config closure while leaving the requested workspace mode unchanged elsewhere. Authorization A must name the concrete paths and operation set for its checkpoint. + +If stock OpenCode cannot keep this source closure immutable without changing required Loom behavior, the tested profile is **UNSUPPORTED**. + +## Network authority + +- Collector endpoint is not a general network service. +- Loom capability channel exposes only the reviewed typed capability protocol. +- Evaluated network access cannot provide an alternate route to collector, runner, container engine, or another trusted service. +- Any network mode needed by Loom/provider behavior is reported separately from capability/evidence-channel reachability. + +## Credentials + +- Provider credentials are product inputs, not evidence authority. +- Evidence-channel/capability-channel admission material, if ultimately required, is not carried in argv, process environment, workspace, project config, Loom state, or mounted readable files. +- A product credential can never authenticate evidence. + +## Current verdict + +Effective-authority separation: **UNPROVEN**. + +This is expected before candidate construction. The Gate 1 PASS means only that these proposed controls are coherent enough for the owner to consider a separately authorized provider-free prototype; it does not mark them experimentally proven. diff --git a/docs/experiments/trust-001/CALLBACK-LIFECYCLE.md b/docs/experiments/trust-001/CALLBACK-LIFECYCLE.md new file mode 100644 index 0000000..d4f51db --- /dev/null +++ b/docs/experiments/trust-001/CALLBACK-LIFECYCLE.md @@ -0,0 +1,207 @@ +# Callback and request lifecycle + +Status: **Wave-1 planning**. + +The remote boundary must preserve stock OpenCode callback semantics without making the transport an alternate product or evidence authority. + +## Request state machine + +Every core → Loom callback/tool request uses a trusted request identity and one lifecycle: + +```text +allocated + ↓ +outstanding + ├─→ responded + ├─→ failed + ├─→ cancelled + └─→ transport-lost +``` + +Rules: + +- exactly one response may be accepted for one outstanding request; +- response identity is allocated by the trusted bridge; +- unsolicited, unknown, duplicate, replayed, stale-generation, and post-cancellation responses are rejected; +- a response from an older plugin generation is never rebound to a newer generation; +- channel loss does not trigger product replay; +- reconnect is outside the bounded experiment; a lost channel makes affected work unresolved/incomplete. + +## Plugin generation + +Activation binds: + +- admitted Loom source revision/snapshot; +- bridge generation; +- isolated Loom process/container identity; +- capability manifest version; +- request sequence namespace. + +All registrations and outstanding requests belong to that generation. + +A restart creates a new generation and cannot inherit outstanding request authority. + +## Generation sealing and collector drain + +Request at-most-once rules do not by themselves prove completeness. The generation therefore needs a trusted two-phase close: + +1. stop admitting new case-required work at the reviewed runner/OpenCode boundary; +2. let all already-admitted callback/tool requests and stock events settle or receive an explicit non-success classification; +3. bridge assigns a monotonic collector sequence to every accepted evidence observation for the generation; +4. only after its accepted observation queue is drained, bridge sends a trusted `seal(generation, finalSequence)`; +5. collector may call the generation complete only after it has a contiguous sequence through `finalSequence` and no post-seal accepted observation/request exists. + +A late request/event after the seal is a protocol/completeness failure, not something silently ignored. A disposable OpenCode process exit may support the close boundary only after provider-free proof that relevant event and bridge queues are drained before the seal. + +## Registration lifecycle + +### Tool transform + +The isolated Loom process executes Loom's transform callback against a remote-side recording editor. + +It returns declarative: + +- namespace definitions; +- tool schemas/metadata; +- opaque handler IDs. + +The trusted bridge installs proxy tool handlers through stock `tool.transform`. + +No Loom function/object crosses the boundary for execution in core. + +### Agent transform + +Stock transforms are synchronous and replayable. Pinned Loom conditionally selects General as default. + +The bounded experiment must either: + +- use a reviewed generic declarative recording/replay form for the exact transform; or +- limit the claim to initial activation and mark dynamic agent-registry reload `UNPROVEN`. + +Trusted bridge code may not contain a Loom-specific hard-coded defaulting rule. + +## Hook semantics + +Stock `PluginHooks.trigger` invokes registered callbacks sequentially for one event. The bridge must preserve that observable behavior. + +For one Loom callback: + +1. core creates/owns the real event identity; +2. bridge allocates callback request ID; +3. event's allowed product fields are serialized; +4. isolated Loom callback executes; +5. returned mutation is structurally validated; +6. bridge applies only fields mutable in the stock API; +7. stock core continues normally. + +### Mutable fields by family + +- `tool.execute.before`: Loom may mutate `tool` / `input` according to stock semantics and may fail the call. +- `tool.execute.after`: Loom may mutate the completed `result` or error object as allowed by stock semantics; callback itself cannot fail the hook channel. +- `permission.evaluate`: Loom may mutate `effect` / `message`. +- `session.context`: Loom may mutate context/system/tool presentation according to stock API. +- `session.retry`: Loom may mutate retry decision. + +Bridge validation is structural and host-semantic only. It does not invent Loom authorization, retry, OQ, budget, cancellation, or workflow behavior. + +## Remote Loom tool execution + +When a stock Tool proxy is invoked: + +1. stock core supplies real `Tool.Context` with Session/agent/message and outer call ID; +2. bridge allocates a fresh remote-handler request ID; +3. bridge may also allocate a runner-owned **proxy child invocation ID** for trustworthy correlation where stock Code Mode reuses the outer call ID; +4. isolated Loom handler receives normal Loom input/context; +5. Loom returns/throws legitimate product data; +6. bridge maps the result into the registered stock Tool schema; +7. stock OpenCode continues through its normal hooks/normalization/session settlement. + +The proxy child ID is evidence correlation owned by the bridge. It is not injected into Loom tool input and does not replace stock Session/Tool identity. + +## Native finality + +Trusted evidence should prefer the stock live Session event surface for canonical native terminal settlement: + +- `session.tool.called`; +- `session.tool.success`; +- `session.tool.failed`. + +The stock Session runner publishes the terminal only after normal tool execution and ToolOutput truncation. The candidate must consume the **live event stream**, not replay Session storage later as a substitute. + +## Code Mode finality + +Stock v2.0.23 does not expose a demonstrated public event containing each inner call's final script-visible value/error. + +For remote Loom proxy tools, the bridge sees: + +- exact proxy invocation entry; +- exact input; +- product result/error returned by isolated Loom; +- stock outer context/parent. + +But stock Code Mode may still transform that Tool result into the JavaScript caller value after the proxy returns. + +Therefore: + +**CAP-004 inner finality remains UNPROVEN.** + +The experiment may prove a runner-owned proxy/wrapping construction sufficient for the selected surface. If it cannot, the result is `UNSUPPORTED` for that stock profile; no OpenCode patch is allowed. + +## Cancellation + +Cancellation ownership remains stock OpenCode/Loom product semantics. + +The bridge owns only request-channel authority: + +- once trusted request authority is cancelled, a late response cannot mutate trusted OpenCode/bridge state or create eligible evidence; +- rejecting that response does **not** prove that isolated Loom made no earlier side effect in its SQLite, workspace, subprocesses, or other product state; +- where stock semantics expose interruption/cancellation, the bridge must propagate it and prove the selected callback/tool behavior; +- cancelling transport request authority is not represented as Loom workflow cancellation; +- a runner timeout is not represented as Loom cancellation; +- a Loom cancellation decision remains produced by Loom code; +- if the timing or effect of a post-cancel remote side effect cannot be shown equivalent to stock behavior, the product result is unresolved and evidence completeness is false for that case. + +No cancellation path may retry product work implicitly. + +## Channel loss + +### Capability channel loss before product terminal + +- request becomes `transport-lost`; +- generation stops admitting new capability work and the isolated generation is fenced/terminated according to the reviewed runner path; +- product outcome is unresolved or follows the reviewed stock transport failure path; +- remote side effects completed before the fence are product state and may be indeterminate; +- evidence remains incomplete; +- no automatic retry. + +### Evidence channel loss + +- collector sequence can no longer be proven contiguous/sealed; +- product completion cannot be rewritten or replayed to repair evidence; +- evidence remains incomplete; +- no reconnect/replay under the bounded experiment. + +### After product completion but before evidence seal + +- completed product result must remain completed; +- collector/capability loss cannot replace it; +- evidence may remain incomplete; +- no replay/retry. + +## Required provider-free fidelity checks before Authorization B + +- transform registration order; +- hook ordering; +- `execute.before` mutation and failure; +- handler return and throw; +- `execute.after` result/error mutation; +- native final Session event; +- permission mutation; +- context/retry mutation; +- duplicate/replayed/stale/late response rejection; +- cancellation + late response, including remote side-effect timing; +- capability/evidence channel loss before/after product completion; +- generation seal/drain, contiguous final sequence, and post-seal late request/event rejection; +- same-parent and different-parent overlapping requests. + +Missing proof is `UNPROVEN`, not PASS or behavioral FAIL. diff --git a/docs/experiments/trust-001/CAPABILITY-MANIFEST.md b/docs/experiments/trust-001/CAPABILITY-MANIFEST.md new file mode 100644 index 0000000..59bcd9e --- /dev/null +++ b/docs/experiments/trust-001/CAPABILITY-MANIFEST.md @@ -0,0 +1,123 @@ +# Bounded capability manifest + +Status: **Wave-1 planning**. + +The capability surface is derived from pinned Loom `149406dfa0a01f94491d17054e50a1bc84bb97be`, not from the full stock OpenCode PluginHost. + +## Key reduction + +Loom setup initially receives OpenCode plugin storage as `legacyStorage`, then creates its own `execution-state.sqlite` and proxies `ctx.storage` to that Loom-local transactional store. + +Most steady-state Loom state therefore stays local to Loom. However pinned Loom retains `legacyStorage`: fresh-session cancellation admission can still perform a legacy `get`, and resumed pre-epoch sessions can invoke lazy migration that reads/scans legacy state and may write migration-refusal records. + +Therefore stock plugin storage is **not setup-only**. The bounded broker must keep Loom-plugin-namespace `get/set/scan` available for the generation, with structural/size bounds and no evidence authority. These values remain product state. + +## Required capabilities + +| Capability | Direction | Why pinned Loom needs it | Trusted identity owner | Authority / notes | Planning status | +|---|---|---|---|---|---| +| immutable `location` facts | core → Loom | runtime/project identity, workspace paths | core/runner | value only; Loom cannot redefine trusted Location | SUPPORTED-DESIGN | +| legacy storage `get/set/scan` | Loom → core | activation import + lazy runtime compatibility/migration | core Loom-plugin storage namespace | generation-lifetime bounded product access; fresh-session compatibility can `get`, resumed legacy migration can `get/set/scan`; never evidence authority | SUPPORTED-DESIGN | +| `rpc.register` | Loom → core registration; calls core → Loom | sidebar RPC | core owns effective registration | async proxy handler | SUPPORTED-DESIGN | +| `agent.transform` | Loom registration → core | set General as default when present | core owns active registry | synchronous stock transform is a fidelity challenge; bounded experiment may support initial generation only | **UNPROVEN** | +| `agent.list` | Loom → core | roster tool | core | normal request/response | SUPPORTED-DESIGN | +| `tool.transform` | Loom registration → core | register Loom namespace/tools | core owns effective registration | remote setup returns declarative registrations + opaque handler IDs | SUPPORTED-DESIGN | +| `tool.list` | Loom → core | attestation checks effective roster registration | core | remote view must preserve Loom-owned handler identity through token mapping | **UNPROVEN** | +| `permission.hook("evaluate")` | core → Loom → core | Loom authorization/grants/budgets | core owns event Session/agent/action; Loom owns product decision | mutable effect/message only; at-most-once callback | SUPPORTED-DESIGN | +| `session.hook("context")` | core → Loom → core | guidance + observed user message + upgrade notice | core | mutable system context | SUPPORTED-DESIGN | +| `session.hook("retry")` | core → Loom → core | Loom retry limit | core | mutable retry decision | SUPPORTED-DESIGN | +| `session.get` | Loom → core | legacy session/project checks, background parent validation | core | returned Session data remains product data | SUPPORTED-DESIGN | +| `session.context` | Loom → core | exact subagent background request discovery | core | request/response | SUPPORTED-DESIGN | +| `session.synthetic` | Loom → core | background/OQ/coordinator delivery | core owns Session identity/admission | Loom supplies product message/metadata | SUPPORTED-DESIGN | +| `tool.hook("execute.before")` | core → Loom → core | cancellation fences, locks, question/budget admission, evidence bookkeeping | core owns call context | mutable input / may fail | SUPPORTED-DESIGN | +| `tool.hook("execute.after")` | core → Loom → core | question decisions, mutation/evidence bookkeeping | core owns call context | mutable result/error; stock finality occurs later | SUPPORTED-DESIGN | +| remote Loom tool execute handler | core → Loom → core | all Loom native/Code Mode tools | bridge allocates proxy request; stock core owns outer context | Loom returns legitimate product result/error only | SUPPORTED-DESIGN | +| trusted event subscription (runner bridge only) | core → host collector | Session ancestry, native called/terminal records, scope accounting | stock core + host collector | not exposed to Loom as evidence authority | SUPPORTED-SOURCE | + +## Source-derived host-use closure + +Static inspection of pinned Loom `149406dfa0a01f94491d17054e50a1bc84bb97be` found direct host use limited to: + +- immutable `location` facts; +- storage `get/set/scan`; +- `rpc.register`; +- `agent.transform` and `agent.list`; +- `tool.transform`, `tool.list`, and tool hooks; +- permission hook; +- Session `get`, `context`, `synthetic`, and Session hooks. + +The runner bridge's trusted `event.subscribe()` use is an observation surface, not a Loom-requested capability and is not exposed to isolated Loom. + +Any newly discovered pinned-Loom host call or later Loom revision adds capability surface and creates a new candidate checkpoint with affected authority review. + +## Not required by pinned Loom selected surface + +The experiment MUST NOT implement the full PluginHost for completeness. Current pinned Loom does not directly require, among others: + +- provider/model transforms; +- MCP transforms; +- VCS API; +- websearch; +- worktree; +- skill transforms; +- shell hooks; +- Session create/remove/compact/move; +- general generation APIs. + +If later source inspection or an authorized scenario demonstrates a need, adding one is a **new candidate checkpoint** and requires affected authority review. + +## Transform fidelity + +### Tool transform + +Pinned Loom's tool transform primarily adds a namespace and tool definitions. The isolated setup can execute Loom's callback against a remote-side recording editor and send: + +- declarative namespace; +- declarative tool metadata/schema; +- opaque execute-handler ID. + +The bridge installs local proxy handlers; Loom code itself never crosses into trusted core. + +### Agent transform + +Pinned Loom performs: + +```text +if general exists -> make general default +``` + +Stock transform callbacks are synchronous and replayable. A remote async callback cannot simply replace that mechanism. + +For the bounded experiment, acceptable planning options are: + +1. prove a generic declarative recording/replay representation for this exact transform; or +2. mark dynamic agent-transform reload outside the tested surface and prove initial activation parity. + +Embedding a Loom-specific `default("general")` rule directly in trusted bridge code is **not acceptable**. + +## Tool identity and remote function identity + +Stock `tool.list` returns function-bearing registrations in-process. Serialized RPC cannot preserve JavaScript function identity directly. + +The broker must retain an opaque registration mapping so the isolated Loom view can recognize its own registered handler identity without trusting a caller-supplied provenance claim. + +This is needed by Loom's attestation tool and remains **UNPROVEN** until the provider-free callback/registration preflight. + +## Code Mode + +For a remote Loom tool invoked inside Code Mode, the trusted bridge can allocate a fresh proxy-child invocation identity when its proxy execute handler is actually entered, while retaining stock outer `Tool.Context.id` as the parent association. + +This can improve correlation without changing OpenCode. + +However stock public APIs still do not independently expose the final per-inner script-visible value/error after every Code Mode conversion. CAP-004 remains **UNPROVEN** pending the bounded experiment. + +## Rejection rule + +No capability may accept: + +- executable callback/module payload for execution in the trusted domain; +- caller-selected trusted invocation/session/parent identity; +- caller-selected sequence/scope/completeness/eligibility; +- generic arbitrary OpenCode method dispatch. + +A large typed surface is acceptable; authority escalation is not. diff --git a/docs/experiments/trust-001/EVIDENCE-SINKS.md b/docs/experiments/trust-001/EVIDENCE-SINKS.md new file mode 100644 index 0000000..dfa3fc6 --- /dev/null +++ b/docs/experiments/trust-001/EVIDENCE-SINKS.md @@ -0,0 +1,88 @@ +# Provisional first-sink inventory + +Status: **Wave-1 planning**. + +A path is in scope because it actually persists, clips, logs, or exports candidate evidence—not because of component ownership or filename. + +## Rule + +Synthetic credentials that enter candidate observation or diagnostic paths must be protected **before** the first candidate-controlled persistence, clipping, logging, or export boundary. + +Calling a stream "non-evidence" is not permission for candidate code to persist or export raw sensitive values. A later checksum/rejection can protect integrity/eligibility but cannot undo disclosure. + +Normal stock OpenCode and Loom product stores are separate product semantics. They are outside the runner evidence-safety claim unless the candidate copies, exports, or relies on their contents; such a copy/export becomes a new first sink. + +## Proposed candidate paths + +| Path | Kind | Raw evidence allowed? | First protection requirement | Status | +|---|---|---:|---|---| +| bridge callback/event objects in process memory | transient trusted memory | yes, bounded | no persistence/logging before handoff; bounded input validation | design | +| bridge → host collector evidence socket | transient trusted IPC | yes, bounded | distinct collector-only endpoint/protocol; no disk/log clipping; bounded frames; trusted peer/channel authority | design | +| bridge ↔ isolated Loom capability socket | transient product IPC | yes, bounded | separate endpoint/protocol; no raw transcript; no direct collector ingress; peer identity/correlation protected and integrity loss makes affected evidence ineligible | design | +| host collector in-memory correlation table | trusted memory | yes, bounded | safety projection before any persistence/preview | design | +| collector temporary/intermediate evidence file | evidence persistence | **no** | only projected safe representation may be written | required | +| runner final result file | evidence persistence | no | existing safe atomic writer/admission pattern | existing design to reuse | +| `--print-result` stdout | evidence export | no | print only validated projected result | existing design to reuse | +| broker/bridge diagnostic logs | diagnostic persistence/stream | no payload | fixed reason codes/IDs only; never raw callback/product payload | required | +| OCI engine logs for candidate containers | persistence | no raw candidate payload | disable/raw-log-driver path for the confidentiality profile or prove projection before persistence; never rely on raw logs as evidence | required | +| Loom isolated stdout/stderr | product/diagnostic stream | transient product data may exist | if runner/engine persists or exports it, that path becomes a first sink and must be disabled or projected first | required | +| OpenCode raw stdout/stderr | product/diagnostic stream | transient product data may exist | if runner/engine persists or exports it, that path becomes a first sink and must be disabled or projected first | required | +| crash/core dump | persistence | no | disable or ensure unavailable/non-persistent | required | +| source/image/component metadata | provenance | safe metadata only | no credential-bearing paths/values | design | + +## Explicit non-evidence product stores + +The candidate should **not** make these stores authoritative evidence sources: + +- Loom's `execution-state.sqlite`; +- stock OpenCode Session database; +- shell output files; +- arbitrary workspace files; +- dashboard state; +- model transcript prose. + +These stores may contain ordinary product data under stock/Loom semantics. The runner confidentiality claim does not retroactively sanitize them. If candidate code copies, clips, logs, exports, or promotes their contents, that new path enters this sink inventory. + +State-dependent Loom assertions should use a later trusted runtime tool/query whose invocation/result is itself observed, rather than reading these stores directly as runner evidence. + +### Stock Session events + +The trusted bridge may consume the **live stock event stream** as a runtime observation source. This is distinct from replaying the Session database as evidence after the fact. + +Normal OpenCode persistence of its own product Session history remains product behavior. If a future design reads that stored history as runner evidence, it becomes a new evidence path and must return to sink review. + +## Capability channel versus evidence channel + +The Loom capability channel transports product callbacks and results. It is **not** an evidence-authority channel. + +Requirements: + +- physically/logically distinct endpoint and accepted descriptor from the evidence channel; +- no collector authentication material or evidence framing on the capability channel; +- no raw channel transcript persisted; +- no debug payload logging; +- bounded frame size before allocation growth; +- oversize/malformed requests rejected with fixed diagnostics; +- an unadmitted process must not impersonate the Loom capability peer; replay/identity/integrity loss makes affected product work unresolved and affected evidence ineligible; +- trusted collector records only observations received through the admitted evidence channel and its own correlated/safe projection. + +## Clipping + +No raw evidence field may be clipped and then described as complete. + +Order: + +```text +raw trusted observation + -> safety projection/redaction/omission + -> size decision + -> persistence/export +``` + +Unsupported/oversized values are omitted before evidence persistence. + +## Candidate construction obligation + +Wave 2 must replace this provisional table with the **actual** sink inventory. + +Discovery of any unreviewed first sink is a checkpoint stop condition. The candidate must update the sink manifest and rerun affected confidentiality review before later evidence is usable. diff --git a/docs/experiments/trust-001/EXPERIMENT-PLAN.md b/docs/experiments/trust-001/EXPERIMENT-PLAN.md new file mode 100644 index 0000000..2a448bc --- /dev/null +++ b/docs/experiments/trust-001/EXPERIMENT-PLAN.md @@ -0,0 +1,228 @@ +# TRUST-001 bounded experiment plan — stock OpenCode revision + +Status: **planning only**. + +## Goal + +Test whether `opencode-eval-runner` can satisfy the TRUST-001 producer boundary for the selected Loom surface while: + +- using stock OpenCode v2.0.23 unchanged; +- preserving normal Loom semantics; +- changing only runner-owned implementation; +- keeping TRUST feasibility separate from full Loom consumer-contract delivery. + +## Authorization boundaries + +### Currently allowed + +- source/history inspection; +- existing artifact inspection; +- planning documents/manifests; +- pure/provider-free tests of already-existing planning/test surfaces. + +### Not authorized by this plan + +- candidate implementation; +- candidate OCI execution; +- model-backed `eval:live`; +- provider credential use; +- merges/default-pin changes; +- OpenCode source/binary changes. + +## Wave 0 — checkpoint/provenance + +Planning source checkpoint: + +- Loom `149406dfa0a01f94491d17054e50a1bc84bb97be`; +- runner research `002aba96441da8c69c5ce19ac77de298cfeb28d2`; +- stock OpenCode v2.0.23 `0fd7e2829449b052abf0078666669302923d77af`. + +Historical PR #41 images remain reported/reference profiles only. + +Baseline rerun is execution and requires explicit authorization; model-backed rerun additionally requires inference authorization. + +## Wave 1 — candidate definition + +Completed planning outputs in this directory: + +1. TCB; +2. effective-authority manifest; +3. bounded capability manifest; +4. callback/request lifecycle; +5. scope/completeness contract; +6. provisional first-sink inventory; +7. stock OpenCode 2.0.23 assessment; +8. evidence provenance ledger. + +### Gate 1 + +Independent review only. + +Verdicts: + +- PASS; +- FAIL; +- UNPROVEN; +- NOT RUN. + +Current state: **PASS**. + +## Authorization A — construction + +Only after Gate 1 PASS. + +Must identify: + +- exact candidate source branch/checkpoint; +- stock OpenCode image/version; +- exact TCB; +- reviewed capability surface, including generation-lifetime legacy storage bounds; +- plugin-source closure manifest: effective add/remove operation set plus every config/source input that can change it; +- distinct evidence-channel and capability-channel endpoints, peer admission, descriptor policy, and no-reconnect rule; +- generation close/seal and contiguous collector-sequence protocol; +- mount/process/network policy; +- candidate-controlled diagnostic/first-sink policy; +- candidate files allowed to change; +- named provider-free local preflights; +- reviewers. + +Authorization A does not authorize model inference. + +## Wave 2 — candidate construction + +If authorized: + +1. runner-owned bridge plugin using stock API only; +2. isolated Loom plugin runtime; +3. reviewed declarative transform/proxy machinery; +4. at-most-once request correlation; +5. trusted live event collector; +6. trusted scope accounting; +7. evidence safety before first candidate-controlled sink; +8. immutable plugin-source closure fences; +9. generation admission close + trusted final-sequence seal/drain; +10. distinct capability and evidence channels; +11. no OpenCode patch. + +### Gate 2 + +Independent source/authority review. + +### Gate 2C + +Provider-free confidentiality/transport preflight. + +Must prove at least: + +- no raw secret at any candidate-controlled first sink, including diagnostics/logging; +- no evidence-channel FD/listener inheritance, duplication, or impersonation by evaluated subprocesses; +- an unadmitted process cannot impersonate the capability peer; replay/identity/integrity loss makes affected evidence ineligible; +- no plugin-source operation-set change or loading escape; +- duplicate/replay/stale/late response rejection; +- generation close/seal, contiguous final sequence, and post-seal late request/event rejection; +- cancellation/late-response behavior does not overclaim remote side-effect rollback; +- capability/evidence channel loss before/after product completion; +- no automatic product retry. + +FAIL rejects the exact checkpoint. + +UNPROVEN/NOT RUN blocks Authorization B. + +## Authorization B — candidate execution + +Separate from construction. + +Names exact candidate checkpoint and scenarios. + +Does not imply model/provider inference authorization. + +## Wave 3 — semantic scenarios + +Report independently: + +- baseline parity; +- Loom conformance; +- evidence integrity. + +Allowed verdicts: PASS / FAIL / UNPROVEN / NOT RUN. + +Scenarios: + +- SCN-00 callback/transform fidelity; +- SCN-01 native Loom tool; +- SCN-02 Code Mode return/transform/discard/caught throw; +- SCN-03 permission denial + trusted state query; +- SCN-04 foreground delegation; +- SCN-05 background delegation; +- SCN-06 OQ continuation/reconciliation; +- SCN-07 identical concurrent calls with reverse completion under same and different parents. + +## Wave 4 — adversarial authority + +Provider-free where possible. + +Test: + +- collector-shaped product data; +- evidence-path modification; +- shared-workspace path aliases; +- plugin load/reload escape; +- process/proc/ptrace/fd attacks; +- container-engine/control access; +- broker/network bypass; +- malformed/code-bearing protocol payload; +- Loom/shell child attacker; +- duplicate/replay/stale/late responses; +- evidence deletion/reordering; +- fake terminal/completeness; +- crash/hang/channel loss; +- cancellation + late response and remote side-effect timing; +- post-seal late request/event and sequence-gap attempts. + +## Wave 5 — confidentiality + +Reconcile every actual first persistence/clipping path. + +For each test literal, short, escaped, deep-key, sensitive-key, boundary-spanning and failure-path synthetic credentials. + +Report confidentiality, integrity and eligibility separately. + +## Wave 6 — runner→Loom composition + +Use only normal consumer path. + +Runner fixture success is supporting evidence, not composition proof. + +Verify state-dependent assertions, descendants/background scope, concurrency and non-evidence vetoes. + +Any real model run requires explicit inference authorization. + +## Wave 7 — independent adjudication + +Final dimensions: + +- baseline parity; +- Loom conformance; +- TRUST-001 tested feasibility; +- confidentiality; +- scoped completeness. + +Success means feasibility only for the tested checkpoint/surface. + +## Rejection/recovery + +On a rejection criterion: + +1. stop exact candidate checkpoint; +2. preserve evidence; +3. mark checkpoint REJECTED; +4. record reproducer/reason; +5. corrections create a new checkpoint; +6. rerun rejecting and affected reviews; +7. authority/scope/sink/protocol changes return to earlier gates. + +No silent continuation under a rejected checkpoint. + +## Remaining program after TRUST + +Even a successful experiment leaves the rest of the Loom contract delivery plan, including any remaining CAP-005/006/007, EVAL-001/002/003, ART-001, and full assembled acceptance gaps. diff --git a/docs/experiments/trust-001/GATE-1-REVIEW.md b/docs/experiments/trust-001/GATE-1-REVIEW.md new file mode 100644 index 0000000..72b95c6 --- /dev/null +++ b/docs/experiments/trust-001/GATE-1-REVIEW.md @@ -0,0 +1,104 @@ +# Gate 1 independent architecture review + +**Current gate verdict: PASS** + +The complete Wave-1 planning package was independently reviewed as one architecture package before any correction was made. Source checks were performed against pinned Loom `149406d` and stock OpenCode v2.0.23. + +The review found correctable planning gaps in legacy-storage lifetime, capability/evidence channel separation, plugin-source closure, cancellation semantics, first-sink wording, and completeness sealing. Those corrections were traced across the affected artifacts and then the full package was reread. + +This PASS is a **design-readiness verdict only**. It does not authorize construction or execution. + +## Candidate under review + +Runner-only experiment using: + +- immutable stock OpenCode v2.0.23; +- runner-owned trusted bridge plugin loaded through the supported OpenCode plugin API; +- isolated Loom plugin execution domain; +- host-side trusted collector/evidence safety; +- bounded typed capability protocol. + +No OpenCode patch/fork/upstream PR is allowed. + +## Required reviewer inputs + +- [STOCK-OPENCODE-2.0.23.md](STOCK-OPENCODE-2.0.23.md) +- [TCB.md](TCB.md) +- [AUTHORITY-MANIFEST.md](AUTHORITY-MANIFEST.md) +- [CAPABILITY-MANIFEST.md](CAPABILITY-MANIFEST.md) +- [CALLBACK-LIFECYCLE.md](CALLBACK-LIFECYCLE.md) +- [SCOPE-COMPLETENESS.md](SCOPE-COMPLETENESS.md) +- [EVIDENCE-SINKS.md](EVIDENCE-SINKS.md) +- [PLANNING-EVIDENCE-PROVENANCE.md](PLANNING-EVIDENCE-PROVENANCE.md) + +## Load-bearing review questions + +### Effective authority + +1. Does the proposed channel/mount/process model prevent isolated Loom and stock shell subprocesses from manufacturing eligible evidence? +2. Can either evaluated domain duplicate/inherit/impersonate the evidence connection? +3. Is the capability channel distinct from the evidence channel, and does its transport bind responses to the admitted Loom generation so impersonation/integrity loss cannot remain eligible evidence? +4. Are evidence storage and runner control endpoints absent from evaluated authority? +5. Does suppression/crash remain incomplete rather than false-complete? + +### Plugin loading + +6. Can evaluated workspace/config mutation cause stock OpenCode to import another untrusted in-process plugin? +7. Does the plugin-source closure manifest include every config document/root, discovery directory, configured source, and watched input that can change the effective add/remove operation set? +8. Is that complete source closure immutable to evaluated authority for the generation without changing required Loom semantics? + +### Broker authority + +9. Is every remote capability tied to a concrete pinned Loom use, including runtime lazy legacy-storage compatibility rather than a false setup-only assumption? +10. Can any operation become arbitrary code execution or generic OpenCode dispatch in trusted core? +11. Are identity, parentage, sequence, scope, terminal, and completeness exclusively trusted-side facts? + +### Fidelity + +12. Is the synchronous `agent.transform` strategy sufficiently bounded for the experiment? +13. Can tool transform/function identity be represented without moving Loom code into trusted core? +14. Are hook mutation order, cancellation, late responses, remote side effects, and at-most-once behavior specified sufficiently for a prototype? +15. Does the design preserve Loom product logic outside the trusted adapter? + +### Observation + +16. Are stock live Session events sufficient for the claimed native terminal evidence? +17. Is public `ctx.event.subscribe()` only claimed as a source-supported access surface, with ordering/drain left for provider-free proof? +18. Is Code Mode inner finality correctly left `UNPROVEN` rather than assumed? +19. Does concurrency use trusted unique correlation or explicit ambiguity rejection? +20. Does completeness require a trusted admission close, contiguous final collector sequence, and generation seal rather than quiescence alone? + +### Confidentiality + +21. Are all proposed candidate-controlled first sinks enumerated, including diagnostics/logging/export paths? +22. Does any proposed candidate-controlled path persist/clip raw credential-bearing observation data before safety projection? +23. Are ordinary stock/Loom product stores kept distinct from runner evidence claims unless candidate code copies/exports them? + +## Known unresolved experimental questions + +These do not automatically prevent Gate 1 PASS for a bounded feasibility prototype, but they MUST remain explicit: + +- Code Mode per-inner final caller result/error: `UNPROVEN`; +- effective channel resistance to stock shell subprocess authority: `UNPROVEN`; +- remote function identity for Loom attestation: `UNPROVEN`; +- dynamic agent-transform reload parity: outside initial bounded claim unless independently solved; +- trusted root Session admission for scope: `UNPROVEN` until provider-free candidate preflight; +- stock public-event ordering/correlation and seal/drain behavior under concurrency: `UNPROVEN` until provider-free candidate preflight; +- remote cancellation side-effect fidelity: `UNPROVEN` until provider-free candidate preflight. + +## Gate 1 verdict vocabulary + +- `PASS` — sufficiently specified and bounded to consider Authorization A. +- `FAIL` — candidate design already violates a requirement or cannot preserve required semantics. +- `UNPROVEN` — review cannot establish enough to authorize the next stage. +- `NOT RUN` — independent review has not occurred. + +A Gate 1 PASS does **not** establish TRUST-001 feasibility and does not authorize candidate execution. It only permits the owner to consider the separate Authorization A for construction/named provider-free preflights. + +## Independent Gate 1 assessment + +**GATE 1: PASS.** + +The corrected planning package is sufficiently coherent, bounded, source-grounded, and falsifiable for the owner to consider a **separate Authorization A** for prototype construction and explicitly named provider-free preflights. + +This PASS does not establish TRUST-001 feasibility. No implementation, candidate OCI execution, model inference, provider credential use, merge, default-pin change, Loom modification, or OpenCode modification was performed by this Gate-1 review. diff --git a/docs/experiments/trust-001/PLANNING-EVIDENCE-PROVENANCE.md b/docs/experiments/trust-001/PLANNING-EVIDENCE-PROVENANCE.md new file mode 100644 index 0000000..8ddbef4 --- /dev/null +++ b/docs/experiments/trust-001/PLANNING-EVIDENCE-PROVENANCE.md @@ -0,0 +1,43 @@ +# Planning evidence provenance + +This ledger prevents reported, source-verified, and future experiment claims from being conflated. + +## Vocabulary + +- **VERIFIED-SOURCE** — independently read from the exact named source revision. +- **REPORTED** — supplied by prior workflow/report/consumer evidence but not rerun as part of this planning branch. +- **UNVERIFIED-IMAGE** — an image/digest is named, but its bytes and full source correspondence are not established by this planning work. +- **FUTURE-EXPERIMENT** — only candidate execution can establish the claim. + +| Claim | Provenance | Notes | +|---|---|---| +| Loom `149406d` pins runner source `002aba…` and safety image `8d7c…` | VERIFIED-SOURCE | Read from Loom source at the named revision. | +| `149406d` directly follows `f2f56c55…` | VERIFIED-SOURCE | Git parent relation. | +| OpenCode v2.0.23 tag resolves to `0fd7e2829449b052abf0078666669302923d77af` | VERIFIED-SOURCE | Git tag ref. | +| OpenCode v2.0.18 tag resolves to `cd9a14a6b688d4021bee381dfd39d2cef9c0f862` | VERIFIED-SOURCE | Git tag ref. | +| v2.0.23 still evaluates configured plugin modules in the OpenCode process | VERIFIED-SOURCE | `packages/core/src/plugin/module.ts` is unchanged from v2.0.18 and invokes `plugin.effect(...)` in-process. | +| v2.0.23 `tool.execute.after` still precedes later result normalization/return | VERIFIED-SOURCE | `packages/core/src/tool.ts` is unchanged from v2.0.18. | +| v2.0.23 durable Session events expose canonical native tool called/success/failed records | VERIFIED-SOURCE | `packages/schema/src/session-event.ts`. | +| v2.0.23 Code Mode public metadata records inner tool name/status/input but not per-inner returned value/error | VERIFIED-SOURCE | `packages/core/src/codemode/tool.ts`. | +| v2.0.23 PluginHost adds parent Session creation, remove, compact, and metadata update since v2.0.18 | VERIFIED-SOURCE | Exact upstream commits inspected. | +| OpenCode v2.0.23 does **not** provide plugin isolation from arbitrary in-process plugins | VERIFIED-SOURCE | Configured module/hook execution remains in-process. | +| Stock v2.0.23 public plugin Context exposes `event.subscribe()` over OpenCode events | VERIFIED-SOURCE | This proves an access surface, not ordering/drain/completeness. | +| Stock v2.0.23 plugin-source discovery derives operations from config documents/roots and discovered/configured sources and watches changes | VERIFIED-SOURCE | Plugin-loading closure must freeze every input that can change the effective operation set. | +| Pinned Loom `149406d` retains `legacyStorage` after proxying normal `ctx.storage` and can access it later for lazy compatibility/migration | VERIFIED-SOURCE | Stock plugin storage cannot be classified as setup-only. | +| Safety image `8d7c…` passed the reported Loom provider-free composition at `149406d` | REPORTED | Do not promote to independently rerun evidence in this branch. | +| Normal-observation image `df50…` demonstrated patched-runtime semantics | REPORTED | Research/reference only; patched OpenCode is out of scope. | +| Safety image `8d7c…` bytes exactly correspond to current planning sources | UNVERIFIED-IMAGE | Must be re-established if ever used for a future authorized baseline. | +| A stock-v2.0.23 runner bridge can satisfy TRUST-001 | FUTURE-EXPERIMENT | Not established by source inspection. | +| A stock-v2.0.23 runner bridge can satisfy Code Mode final caller evidence | FUTURE-EXPERIMENT | Status is currently `UNPROVEN`; no sufficient stock public final-inner-result surface has yet been demonstrated. | +| Proposed OS/channel controls prevent shell subprocess authority from forging evidence | FUTURE-EXPERIMENT | Must be verified against the constructed candidate. | + +## Source anchors + +Planning source inspection used: + +- `bateau84/loom@149406dfa0a01f94491d17054e50a1bc84bb97be` +- `bateau84/opencode-eval-runner@002aba96441da8c69c5ce19ac77de298cfeb28d2` +- `anomalyco/opencode@v2.0.23` +- `anomalyco/opencode@v2.0.18` + +No model inference or candidate execution was performed to produce this ledger. diff --git a/docs/experiments/trust-001/README.md b/docs/experiments/trust-001/README.md new file mode 100644 index 0000000..cde46d8 --- /dev/null +++ b/docs/experiments/trust-001/README.md @@ -0,0 +1,42 @@ +# TRUST-001 stock-OpenCode experiment planning + +Status: **planning only**. No prototype construction, candidate execution, inference, merge, default-pin change, or OpenCode modification is authorized by this directory. + +Gate 1 status: **PASS** after independent whole-package architecture review and planning corrections. This does not grant Authorization A or authorize candidate construction/execution. + +## Program goal + +The delivery goal remains full `opencode-eval-runner` compliance with Loom's consumer contract. This work addresses the producer-trust gap only and does not turn TRUST-001 feasibility into full consumer-contract acceptance. + +## Hard scope constraints + +- OpenCode is an immutable external dependency. +- Target stock OpenCode version for planning: **v2.0.23** (`0fd7e2829449b052abf0078666669302923d77af`). +- No OpenCode source patch, fork, custom binary, or upstream PR is part of this work. +- All candidate implementation, if later authorized, belongs to `opencode-eval-runner`. +- Loom production semantics remain unchanged. Loom changes only as required by its consumer contract. +- PR #41 patched-runtime work is research/reference evidence only, not an allowed runtime dependency. + +## Pinned planning inputs + +- Loom consumer/product checkpoint: `149406dfa0a01f94491d17054e50a1bc84bb97be` +- Runner research checkpoint: `002aba96441da8c69c5ce19ac77de298cfeb28d2` +- Stock OpenCode: `v2.0.23` / `0fd7e2829449b052abf0078666669302923d77af` +- Historical safety image reported by Loom: `ghcr.io/bateau84/opencode-eval-runner@sha256:8d7c90223f43bb3c9954918caf1e0a578745c5db2f18dd8f5b69f00a0c6d4723` +- Historical normal-observation image: `ghcr.io/bateau84/opencode-eval-runner@sha256:df50c03b1efc5d42eb94cf7ffcb7d4ee5c51b7c4970fa20585b8048713ed05d9` + +The two images above demonstrate different historical profiles. Neither is a permitted dependency for a stock-OpenCode candidate without renewed checkpoint proof. + +## Wave-1 outputs + +- [PLANNING-EVIDENCE-PROVENANCE.md](PLANNING-EVIDENCE-PROVENANCE.md) +- [STOCK-OPENCODE-2.0.23.md](STOCK-OPENCODE-2.0.23.md) +- [TCB.md](TCB.md) +- [AUTHORITY-MANIFEST.md](AUTHORITY-MANIFEST.md) +- [CAPABILITY-MANIFEST.md](CAPABILITY-MANIFEST.md) +- [CALLBACK-LIFECYCLE.md](CALLBACK-LIFECYCLE.md) +- [SCOPE-COMPLETENESS.md](SCOPE-COMPLETENESS.md) +- [EVIDENCE-SINKS.md](EVIDENCE-SINKS.md) +- [GATE-1-REVIEW.md](GATE-1-REVIEW.md) + +Gate 1 is `PASS`. The next possible step is a **separate owner Authorization A** naming the exact construction checkpoint and provider-free preflights; this planning PR does not grant it. diff --git a/docs/experiments/trust-001/SCOPE-COMPLETENESS.md b/docs/experiments/trust-001/SCOPE-COMPLETENESS.md new file mode 100644 index 0000000..f9b6266 --- /dev/null +++ b/docs/experiments/trust-001/SCOPE-COMPLETENESS.md @@ -0,0 +1,152 @@ +# Scope and completeness contract + +Status: **Wave-1 planning**. + +Coverage is relative to the case-required observation scope. Loom defines what must be observed; the runner resolves membership and completeness from trusted runtime facts. + +## Bounded experiment scope descriptor + +For the initial experiment, use a deliberately strong but finite scope: + +```text +generation = one admitted bridge/Loom generation for this runner invocation +root = actual target Session for this runner invocation +include = root + + all actual descendant Sessions created from root during the case + + all required tool/proxy invocations in those Sessions +interval = target turn start until trusted closure predicate +``` + +A future contract may support narrower explicit selectors, but the experiment does not need an unlimited global monitor. + +## Root admission + +The host launch record must bind the experiment to the actual target invocation. + +The collector must not select "first Session seen" or infer root identity from model prose. + +The candidate must define how the normal runner invocation admits its target root Session ID before Gate 2. Acceptable proof must come from stock runtime/runner identity, not Loom claims. + +Until provider-free candidate construction demonstrates this binding, root admission is **UNPROVEN**. + +## Session membership + +Trusted stock Session events provide: + +- `session.created.sessionID`; +- `session.created.parentID`; +- Session step/message identities. + +A Session joins scope only when trusted ancestry reaches the admitted root under the declared inclusion rule. + +Loom-returned child IDs may be used as product data/corroboration but cannot create authoritative membership. + +## Tool membership + +Native Tool identity is reconstructed from trusted stock Session events: + +- assistant message / Session; +- input start/name; +- called input; +- call ID; +- terminal success/failure. + +For runner proxy Loom tools, the bridge additionally owns the remote proxy request/child correlation. + +A plugin cannot add or remove an authoritative tool member by emitting collector-shaped payloads. + +## Background descendants + +A root/foreground turn ending does not close scope while a required admitted background descendant remains unresolved. + +For the bounded experiment, every descendant Session admitted by the inclusion rule must reach an accounted terminal/idle state or an explicit non-success classification. + +## Trusted member states + +Each required member is represented as one of: + +- started; +- terminal-success; +- terminal-error; +- interrupted; +- timed-out; +- unresolved; +- unsupported. + +No unknown state is silently counted as zero/missing-free. + +## Trusted closure seal + +Completeness is a two-phase property: **quiescence, then trusted seal**. Observing that all currently known members look terminal is not enough because a late admitted callback, descendant, or live event could otherwise arrive after the collector declares success. + +For the bounded experiment: + +1. a reviewed runner/OpenCode boundary closes admission for new case-required work for the generation; +2. every already-admitted Session/tool/proxy/capability member settles or receives an explicit non-success classification; +3. the trusted bridge assigns a monotonic collector sequence to every accepted observation for that generation; +4. after all accepted observations are emitted, the bridge sends `seal(generation, finalSequence)`; +5. the collector accepts completeness only if it has a contiguous sequence through `finalSequence`, all required members satisfy the closure predicate, and no accepted post-seal work/event exists. + +A late required request/event after seal is a protocol/completeness failure. It is never silently ignored. + +A disposable OpenCode process exit may be used as part of the trusted close boundary only after provider-free proof that the relevant stock-event and bridge queues are drained before the seal. Process exit by itself is not completeness. + +## Closure predicate + +The bounded experiment scope can close as complete only when all are true: + +1. root identity is admitted; +2. no required Session membership is ambiguous; +3. trusted generation admission is closed for new case-required work; +4. root reaches the required lifecycle boundary; +5. every in-scope descendant required by the scope has reached its required lifecycle boundary; +6. every in-scope required tool/proxy invocation has one trusted terminal; +7. there are no unresolved outstanding capability requests; +8. evidence channel did not lose an unaccounted interval; +9. collector has the trusted generation seal and a contiguous sequence through its `finalSequence`; +10. no required request/event was accepted after that seal; +11. no required operation is classified unsupported; +12. no evidence-safety loss required by the case prevents the deterministic assertion. + +If any condition is false, completeness is false/unknown as appropriate. + +## Forbidden absence + +A forbidden-absence assertion is eligible only after the trusted generation seal and only if every in-scope path on which the forbidden behavior could occur is completely covered. + +Examples: + +- If forbidden behavior may occur in a background child, root-only coverage is insufficient. +- If Code Mode inner capture is unsupported for the selected operation, absence inside Code Mode is not proved. +- A missing terminal cannot be interpreted as "did not happen." + +## Concurrency + +Calls are not correlated by FIFO or completion order. + +The experiment must cover: + +- identical calls, same parent, reverse completion; +- identical calls, different parents, overlapping execution. + +For remote Loom proxy tools, bridge-owned proxy request IDs may provide unique child correlation while retaining the stock outer call as parent. + +For stock inner tools that cannot be uniquely correlated through public APIs, report ambiguity/unsupported rather than guess. + +## Scope authority + +Evaluated Loom/plugin/product data cannot authoritatively provide: + +- member list; +- parentage; +- sequence; +- terminal state; +- close signal; +- missing count; +- completeness. + +These come only from the trusted collector/bridge and stock runtime surfaces admitted by this contract. + +## Experiment limitation + +A successful bounded scope proves only the selected scope rule. It does not establish unlimited descendant coverage or every future Loom case scope. diff --git a/docs/experiments/trust-001/STOCK-OPENCODE-2.0.23.md b/docs/experiments/trust-001/STOCK-OPENCODE-2.0.23.md new file mode 100644 index 0000000..730a809 --- /dev/null +++ b/docs/experiments/trust-001/STOCK-OPENCODE-2.0.23.md @@ -0,0 +1,89 @@ +# Stock OpenCode 2.0.23 capability assessment + +## Decision boundary + +OpenCode v2.0.23 is treated as **stock and immutable**. A missing public capability becomes `UNPROVEN` or `UNSUPPORTED`; it is not permission to patch OpenCode. + +## Changes since v2.0.18 that help + +Stock v2.0.23 adds useful supported plugin/session API surface: + +- parent Session creation through `session.create({parentID})`; +- `session.remove`; +- `session.compact`; +- Session metadata forwarding through `session.update`; +- explicit provider request headers carrying Session and parent Session IDs; +- improved parallel permission-rejection handling. + +These improve lifecycle fidelity and future compatibility, but pinned Loom `149406d` does not depend on most of these additions in its normal control plane. + +## Critical components unchanged + +Exact Git blobs are unchanged between v2.0.18 and v2.0.23 for: + +- `packages/core/src/plugin/module.ts`; +- `packages/core/src/plugin/hooks.ts`; +- `packages/core/src/tool.ts`; +- `packages/core/src/tool/runtime.ts`; +- `packages/core/src/plugin/supervisor.ts`. + +Consequences: + +1. configured plugin code is still loaded/evaluated in the OpenCode process; +2. plugin hooks still execute in-process; +3. the plugin supervisor is lifecycle management, not security isolation; +4. `tool.execute.after` still occurs before later core normalization and return. + +## Useful stock evidence surfaces + +| Requirement | Stock surface | Planning status | +|---|---|---| +| Public live event access | plugin Context `ctx.event.subscribe()` returns OpenCode event stream | SUPPORTED-SOURCE for access; ordering/drain/completeness remain UNPROVEN | +| Session creation / ancestry | durable `session.created` with `parentID`; Session API | SUPPORTED-SOURCE | +| Agent for a step | durable `session.step.started` binds assistant message to agent | SUPPORTED-SOURCE | +| Native tool call ID/input | durable Session tool input/called events | SUPPORTED-SOURCE | +| Native terminal success/failure | durable `session.tool.success` / `session.tool.failed` | SUPPORTED-SOURCE | +| Permission mutation | `permission.hook("evaluate")` | SUPPORTED-SOURCE | +| Tool input mutation/failure | `tool.hook("execute.before")` | SUPPORTED-SOURCE | +| Tool handler result/error mutation | `tool.hook("execute.after")` | SUPPORTED-SOURCE, but not final caller boundary | +| Session context/retry mutation | Session hooks | SUPPORTED-SOURCE | +| Background/foreground child mechanics | stock Subagent/Session APIs/events | SUPPORTED-SOURCE; composition still required | +| Code Mode inner name/input/status | Code Mode metadata + hooks | SUPPORTED-SOURCE | +| Code Mode per-inner final caller value/error | no demonstrated public final boundary | **UNPROVEN** | +| Trusted run-wide ordering | public event delivery + runner sequence could support it | **UNPROVEN until provider-free ordering test** | +| Plugin isolation | none supplied by stock OpenCode | **UNSUPPORTED by OpenCode itself** | + +## Code Mode finality + +Stock v2.0.23 Code Mode records inner calls as roughly: + +`{ tool, status, input }` + +and returns them as outer `execute` metadata. Per-inner returned value/error is not retained there. + +The inner tool execution path returns a `Tool.Result`, then Code Mode derives the script-visible value from structured `output` or textual content. Public `execute.after` runs before later core processing. Therefore CAP-004 cannot be marked supported merely from the public hook shape. + +The experiment may investigate whether a runner-owned proxy tool can establish the tested inner finality without inference or OpenCode modification. Until demonstrated, the row stays `UNPROVEN`. + +## Built-in shell authority + +Stock OpenCode's built-in shell tool spawns real subprocesses from the OpenCode runtime domain. Unless a Session-specific environment exists, shell invocation begins from `process.env`. + +This is load-bearing for TRUST-001: + +- bridge/collector secrets MUST NOT be placed in the OpenCode process environment if shell can inherit them; +- shell subprocesses are evaluated workload authority for the threat model; +- evidence storage and collector paths must remain outside their writable/reachable authority; +- the candidate must prevent shell subprocesses from stealing/injecting the evidence channel or impersonating the admitted Loom capability peer, while allowing suppression to result only in incomplete/ineligible evidence. + +## Plugin loading escape + +Stock OpenCode derives plugin operations from configuration documents/directories, auto-discovers entries under `plugin/` and `plugins/`, resolves configured package/local targets, watches relevant config/source inputs, and evaluates admitted plugin modules in-process. + +The candidate must therefore freeze the **effective plugin-source operation set and every input that can change it**, not merely make one bridge file read-only. The runner must enumerate those config roots/documents, discovery directories, configured targets, symlink/source identities, and watched inputs for the exact profile before activation and keep them outside evaluated write authority for the generation. + +Any post-activation change in the effective operation set or admitted source identity is a checkpoint stop/rejection. If this cannot be guaranteed using stock behavior without changing required Loom semantics, TRUST-001 is unsupported for that profile. + +## Conclusion + +v2.0.23 is a better stock target than v2.0.18, mainly due to improved Session APIs. It does not provide the missing trust boundary and does not solve Code Mode finality. The runner-only plan remains viable for investigation; the Gate 1 review treats these gaps explicitly and leaves their runtime proof to later gates. diff --git a/docs/experiments/trust-001/TCB.md b/docs/experiments/trust-001/TCB.md new file mode 100644 index 0000000..c7dc297 --- /dev/null +++ b/docs/experiments/trust-001/TCB.md @@ -0,0 +1,110 @@ +# Provisional trusted computing boundary + +Status: **Wave-1 planning artifact. Gate 1 PASS; runtime properties remain unproven.** + +This TCB is the candidate boundary accepted for Gate-1 design readiness. It is not a claim that the boundary is already effective. + +## Trusted components + +For the bounded experiment, the proposed trusted domain is: + +1. host `opencode-eval-runner` launcher; +2. host collector, scope accountant, evidence-safety projector, and final evidence writer; +3. host kernel and selected OCI engine only to the extent required to enforce the declared container/process boundary; +4. **stock OpenCode v2.0.23** core runtime; +5. a runner-owned bridge plugin loaded by stock OpenCode through its supported plugin API; +6. runner-owned bridge protocol framing/correlation code. + +OpenCode itself is immutable. No patched binary or source fork belongs to the TCB. + +## Untrusted/evaluated authority + +Treat as untrusted for evidence authority: + +- the complete Loom plugin module at `149406d`; +- Loom dependencies; +- Loom tool handlers, hooks, transforms, workflow/authorization/OQ logic; +- the Loom plugin's direct filesystem, Git, dashboard, and subprocess activity; +- processes spawned by Loom; +- model-produced data; +- arbitrary tool/product return values; +- **commands spawned by stock OpenCode's built-in shell tool**; +- the writable evaluated workspace; +- any external plugin not explicitly admitted as trusted or isolated equivalently. + +The last two items are load-bearing. Isolating Loom alone does not establish TRUST-001 if an evaluated shell command or another plugin can control the bridge/core/collector. + +## Proposed topology + +```text +host opencode-eval-runner + ├─ trusted collector / scope accounting / safety / writer + │ + ├─ stock OpenCode v2.0.23 OCI domain + │ ├─ OpenCode core + │ ├─ runner-owned trusted bridge plugin + │ └─ untrusted built-in tool subprocesses (e.g. shell) + │ + └─ isolated Loom OCI domain + ├─ Loom module + dependencies + ├─ Loom callbacks/tool handlers + ├─ Loom local runtime state + └─ Loom-spawned children +``` + +The stock OpenCode domain is therefore a **mixed-authority OS domain**. The Gate 1 review requires later proof that the proposed channel and mount controls prevent its untrusted subprocesses from escalating into evidence authority. + +Two channels have different TCB meaning: + +- **evidence channel:** trusted bridge → host collector only; this is the sole candidate observation ingress with evidence authority; +- **capability channel:** trusted bridge ↔ isolated Loom; payloads are untrusted product data and have no direct evidence-write authority, while trusted transport admission/correlation must bind them to the admitted Loom generation. + +They must not share accepted descriptors, listener paths, authentication material, or collector-readable evidence framing. + +## Evidence authority + +Only the trusted side may establish: + +- effective admitted registration identity; +- trusted collector sequence; +- authoritative runtime Session/message/call facts obtained from stock OpenCode; +- runner-owned proxy invocation identity when stock OpenCode lacks a unique inner identity; +- scope membership, generation admission close, trusted seal, and final collector sequence/drain; +- evidence completeness/eligibility; +- final evidence persistence. + +The isolated Loom side may legitimately supply: + +- tool output; +- thrown errors; +- permission decisions; +- hook mutations; +- workflow/OQ state; +- other product data. + +Those values are not provenance merely because they contain provenance-shaped fields. + +## TCB exclusions + +The following are explicitly not trusted evidence authorities: + +- Loom's own evidence records; +- product prose; +- model claims; +- workspace files; +- stock OpenCode Session storage when read after the fact as a substitute for a reviewed live capture boundary; +- raw stdout/stderr; +- OCI logs; +- the historical PR #41 patched observer; +- signatures/HMACs created by evaluated code. + +## Open questions carried beyond Gate 1 + +1. Can the runner establish the evidence channel so shell subprocesses cannot inherit, duplicate, or impersonate it, while keeping the capability channel non-authoritative? +2. Can the exact plugin-source closure manifest be kept immutable to evaluated authority for the tested stock profile? +3. Can the runner prove generation admission close, queue drain, contiguous final sequence, and trusted seal under concurrency? +4. Can the runner prove the required Code Mode final-inner-result boundary using stock APIs/proxy wrapping? +5. Does the tested callback surface preserve stock ordering, cancellation, and remote side-effect semantics? +6. Can every candidate-controlled first sink be protected before persistence/clipping/logging/export? + +These are provider-free/source-review obligations for later gates, not permission to alter OpenCode and not reasons to claim the properties are already proven.