Skip to content

refactor: scope eval trust to reviewed runtime observations - #45

Merged
bateau84 merged 21 commits into
mainfrom
refactor/trusted-checkout-evidence
Oct 7, 2026
Merged

bateau84 merged 21 commits into
mainfrom
refactor/trusted-checkout-evidence

Conversation

@bateau84

@bateau84 bateau84 commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Integrated direction

This PR now integrates the completed child work from #46–#51 into one trusted-checkout runtime-evidence implementation on stock OpenCode 2.0.23.

The normal Loom eval profile remains a trusted-checkout evaluation. This PR does not add an OpenCode patch/fork, remote PluginHost, protected channel, HMAC/signing boundary, capability broker, or hostile-plugin isolation.

Canonical public contract

The only authoritative public runtime-evidence object is:

opencode-eval-runner/runtime-evidence/v1

Raw observer records are internal adapter input only.

Existing tools, actions, tool_result_evidence, stdout/stderr, Session/model text, and workspace files remain convenience/diagnostic data and are not promoted into runtime evidence. Final cleanup #52 also removes competing runtime evidence_eligible signals from the diagnostic tool-result/safety projections.

There is one canonical builder/validator that owns:

  • complete | incomplete | unsupported | invalid;
  • capture closure and unknown coverage;
  • missing terminals;
  • observer/callback loss;
  • timeout/interruption;
  • duplicate/ambiguous invocation and sequence detection;
  • identity-based concurrency;
  • per-boundary support;
  • assertion-scoped boundary/field eligibility.

Unknown coverage is never converted to zero.

Boundary semantics

The v1 contract exposes:

  • native
  • code_mode_execution
  • code_mode_finality

Overall complete means the supported capture boundaries are complete. It does not mean every possible assertion is supported.

This allows:

overall: complete / eligible
native: complete / eligible
code_mode_execution: complete / eligible
code_mode_finality: unsupported / ineligible

An assertion that only requires native evidence can therefore remain eligible even though Code Mode finality is unsupported. Assertions that require a redacted/omitted/unsupported exact field remain ineligible for that field.

Native observation

The integrated observer preserves #51's stock-2.0.23 boundary:

decoded executable input -> transformed tool.execute wrapper

terminal success/error ->
  session.tool.success
  session.tool.failed

It records opaque invocation identity, actual tool, agent, Session, message, real CallID, executable input, Session ancestry, terminal result/error, and monotonic start/terminal ordering.

Correlation is identity-based. It does not use FIFO, input equality, tool name, or completion order.

The raw former native_tool_observations projection is no longer a public result.

Code Mode

The #50 stock-runtime result is retained faithfully.

Supported facts:

  • unique per-inner invocation identity;
  • actual effective tool;
  • decoded/executable input;
  • Session/message/agent;
  • actual outer execute CallID;
  • binding to the outer invocation;
  • start/handler-terminal ordering;
  • identical concurrent calls and reverse completion.

The transformed handler's earlier value/error is not promoted as caller-final evidence.

Exact final script-visible value/error is represented as:

{
  "state": "unsupported",
  "reason": "stock_codemode_final_boundary_not_exposed"
}

Stock OpenCode 2.0.23 does not expose a supported boundary that proves the exact final value/error seen by a Code Mode script for each inner call. That assertion is reported as unsupported.

Evidence safety

The #47 safety behavior is applied to authoritative dynamic values before their first observation persistence/output sink:

raw value in observer memory
  -> sanitize/redact/omit
  -> size decision
  -> internal capture
  -> validate/account
  -> result serialization
  -> stdout/host-file persistence

Public field states are only:

  • available
  • redacted
  • omitted
  • unsupported

Product exit_code, timeout, and product success/failure remain separate from evidence eligibility.

Child PR disposition

PR Disposition
#46 fail-closed runtime evidence accounting Incorporated / superseded as a separate implementation. Its stronger accounting is folded into the canonical v1 builder/validator. No second public/accounting status engine remains.
#47 evidence safety Incorporated. Pre-sink sanitization, redaction/omission, size ordering, credential inventory, and output guard behavior are retained. Internal exact terminology is not part of the public runtime-evidence schema.
#48 provider-free acceptance gate Incorporated and tightened. It now requires the exact final v1 contract rather than rollout-compatible alternate shapes.
#49 runtime evidence v1 contract Incorporated and evolved. It remains the public wire contract; overall vs boundary/assertion eligibility was revised to preserve partial support.
#50 stock Code Mode experiment Incorporated as capability proof/diagnostic probe. Supported execution facts are used; exact final caller value/error remains explicitly unsupported.
#51 stock native observer Incorporated / superseded as a public projection. Its stock observer boundary is retained as internal adapter input to runtime_evidence; no public native_tool_observations object remains.

Child PRs #46–#51 have no remaining implementation authority after this integration. Final cleanup #52 removes parallel-work duplication without adding architecture. Explicit unsupported areas are: Code Mode caller-final value/error on stock 2.0.23, OpenCode runtime observation for the github-copilot-cli transport, and protection against a deliberately hostile plugin sharing the trusted OpenCode process.

Final integration cleanup

#52 is the final child PR targeting this branch. Its head a3f575d78f4b14d62b76d14dceea49b76496fd2d:

  • removes duplicate runtime-evidence constants and dead truncated accounting;
  • removes competing evidence_eligible signals from diagnostic tool_result_evidence / safety projections;
  • removes the unused legacy tool-result clipping helper;
  • confirms PR [SUPERSEDED by #45] feat: preserve eval:live semantics while observing runtime tools #41 protected-runtime/signing/broker machinery is absent;
  • documents Code Mode finality, Copilot runtime observation, and hostile-plugin isolation as explicit unsupported areas;
  • keeps stock OpenCode 2.0.23 and normal invoke behavior.

Validation on #52:

#52 has now been squash-merged into this branch as b10c52d50aaf0c038e77e80eaca3890347d5fa82.

Final-head validation

Head: b10c52d50aaf0c038e77e80eaca3890347d5fa82

  • CI #296 / run 37519197332 — PASS
    • Python/unit suite: 92/92 PASS
    • stock Code Mode probe: PASS
    • exact caller terminal capability reported unsupported: PASS
    • OpenCode plugin activation preflight: PASS
    • config-root plugin dependency preflight: PASS
    • pinned OpenCode remains 2.0.23
  • Provider-free runtime evidence acceptance fix: seed host OpenCode model catalog by default #10 / run 37519197280 — PASS
    • helper/unit checks: 5/5 PASS
    • native_success: PASS
    • native_error: PASS
    • code_success: PASS
    • code_caught_error: PASS
    • concurrent_reverse: PASS
    • delegation: PASS
    • timeout: PASS
    • interrupted: PASS
    • redaction: PASS
    • collector: PASS
  • Stock native observer integration fix: invoke OpenCode API with V2 CLI syntax #22 / run 37519197390 — PASS
    • stock 2.0.23: PASS
    • complete capture: PASS
    • provider-free probe: PASS

Review state

The integration cleanup from #52 is merged. PR #45 now contains the final integrated implementation and remains ready for final independent review.

PR #45 remains open/draft and is not merged.

Independent final review

Final review of head b10c52d50aaf0c038e77e80eaca3890347d5fa82 found one blocking correctness defect: stock Code Mode's synthetic model-facing execute invocation was not emitted as an authoritative native observation, so native coverage could falsely remain complete/zero while Code Mode had run and inner parents resolved only to an internal token.

Corrective child PR #54 observes the real outer call at stock execute.before, binds inner parents to that observed invocation, and makes dangling/wrong-identity parents invalid. #54 is green across CI, the stock native probe, and the 10-scenario provider-free acceptance gate.

Review status: NOT READY until #54 is merged. No threat-model expansion is required.

bateau84 commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Supersession checkpoint

CI #265 is green on 618b2e9.

Verified on this replacement head:

  • Python/unit checks
  • stock OpenCode 2.0.23 image build
  • OpenCode plugin activation preflight
  • config-root plugin workspace dependency resolution
  • pinned tool/version checks
  • reasoning-control checks

The old #41 implementation is intentionally not merged wholesale. Evidence-safety, disposable-state, observation semantics, delegation/concurrency probes, and other useful pieces will be ported only where they fit the new stock-runtime trust model. Patched-runtime, protected-channel, remote PluginHost/isolation, HMAC trust-boundary, and signing-as-TRUST-prerequisite work stay out.

@bateau84 bateau84 left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Final independent review — NOT READY

BLOCKING — outer Code Mode execute invocation is missing from authoritative evidence

container/native_observer.ts observes registered direct tools through tool.transform, but stock OpenCode creates the synthetic Code Mode execute tool later inside Tool.snapshot; it is therefore not transform-wrapped.

The observer's execute.before hook only creates/updates an internal activeOuter correlation token. It does not emit a public/native start record for the actual outer execute invocation, and Session terminal handling ignores the outer terminal because there is no corresponding nativeActive entry.

Consequences:

  • a Code Mode run can have an actual model-facing execute tool invocation while the public native boundary reports zero such calls and remains complete;
  • an absence assertion over native/direct tools can therefore be false-positive for execute;
  • Code Mode parent points to an internal opaque token that has no corresponding outer observation in runtime_evidence;
  • the acceptance gate currently only checks that this parent ID is non-empty/shared; it does not require a matching observed outer invocation.

This conflicts with the v1 contract's description of native as direct/native execution and its claim of parent binding to the actual outer invocation.

Required correction

Observe the model-facing outer execute call through a supported stock-2.0.23 runtime surface and include it in authoritative evidence, or explicitly add a separate unsupported/partial outer-execute boundary so absence cannot be claimed.

Do not infer it from product stdout/model data and do not patch OpenCode. Stock Session tool input/called/terminal surfaces or another reviewed supported boundary should be investigated.

Add provider-free regression coverage that proves:

  1. one Code Mode run contains an authoritative outer execute observation;
  2. inner records bind to that exact observed outer invocation;
  3. a no-native/absence assertion cannot report zero when execute actually ran;
  4. concurrent inner calls retain the same observed outer parent.

Other review results

The scoped threat model is otherwise preserved: I found no reintroduction of protected-channel, signing, broker, OpenCode patching, or hostile-plugin isolation. The native Session terminal boundary, fail-closed accounting, explicit Code Mode finality limitation, pre-sink evidence projection, host re-validation, and the final-head CI/acceptance runs are coherent.

Verdict: NOT READY until the outer-execute coverage hole is closed or explicitly represented as unsupported.

Final integration cleanup for PR #45: keep runtime_evidence/v1 as the sole evidence authority, remove duplicate/dead accounting and competing eligibility signals, retain stock OpenCode 2.0.23 behavior, and document unsupported areas.
## Purpose

Correct the blocking outer-`execute` evidence gap found during the
independent final review of #45.

This branch is based on the post-#52 target head
`b10c52d50aaf0c038e77e80eaca3890347d5fa82`.

## Defect

Stock OpenCode 2.0.23 creates the model-facing Code Mode `execute` tool
inside `Tool.snapshot`, after registration transforms have run. The
integrated observer therefore did not emit a native observation for that
real outer invocation. Inner records were linked only to an internal
opaque token.

That allowed the public native boundary to report complete/zero
outer-`execute` observations while Code Mode had actually run.

## Correction

- Observe the real model-facing outer `execute` invocation at stock
`execute.before`.
- Keep `session.tool.success` / `session.tool.failed` as terminal
authority.
- Bind Code Mode inner records to that observed outer invocation ID.
- Reject dangling, wrong-tool, or identity-mismatched parents in
canonical accounting/validation.
- Strengthen provider-free acceptance checks for outer observation,
exact parent resolution, identity agreement, non-zero native coverage,
concurrency, and reverse completion.
- Document that synthetic outer `execute` input is exact at
`execute.before` but pre-`CodeMode.Input` decode.

## Scope

No threat-model expansion. No OpenCode patch/fork, protected channel,
signing, broker, process isolation, or hostile-plugin machinery.

Targets `refactor/trusted-checkout-evidence`.

## Validation

Head: `221af580640360c16e263cc05136d3be0d090a0f`

- CI #307 / run 37525849151: **PASS**, 94/94 Python tests; stock Code
Mode diagnostic probe PASS; OpenCode **2.0.23**.
- Provider-free runtime evidence acceptance #21 / run 37525849081:
**PASS**, 6/6 helper tests and all 10 scenarios.
- Stock native observer integration #33 / run 37525849107: **PASS**.
- PR is mergeable against the current #45 branch.

#53 is closed as superseded; it became dirty only because #52 merged
while this review was in progress.

@bateau84 bateau84 left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Final re-review — NOT READY

#54 is merged and closes the outer Code Mode execute coverage defect. The integrated head is green across CI, provider-free acceptance, and the stock native observer gate.

BLOCKING — authoritative capture is still target-writable

The observer still writes authoritative records to:

/tmp/runtime/runtime-observer.jsonl

runner/cli.py mounts all of /tmp as a writable tmpfs for the evaluated OpenCode process. Normal model-driven tool/shell subprocesses execute under that same runtime user and can therefore potentially read, truncate, replace, or append to this file.

This conflicts with the trusted-checkout contract already documented in this PR: target-writable evidence files do not independently establish that an event occurred.

This is not hostile-plugin resistance and does not require the old TRUST-001 security architecture. It is ordinary eval correctness: model/tool execution must not be able to manufacture or alter the evidence used to score itself.

Required correction

Move authoritative observer transport off the evaluated process's writable filesystem. A small runner-owned pipe/socketpair/inherited-FD style transport is sufficient if normal tool subprocesses cannot write to it. No signing, HMAC, broker, remote PluginHost, or OpenCode patch is required.

Add one provider-free tamper regression proving an evaluated tool/subprocess cannot create, replace, truncate, or append authoritative records.

Status

  • outer execute authoritative observation: FIXED
  • inner-to-outer binding: FIXED
  • Code Mode finality remains correctly unsupported
  • stock OpenCode 2.0.23: PASS
  • CI / acceptance / native observer gates: PASS
  • authoritative transport integrity against normal evaluated tool writes: BLOCKING

Verdict: NOT READY until the target-writable capture file is removed from the authoritative path.

bateau84 commented Oct 6, 2026

Copy link
Copy Markdown
Owner Author

Final blocker follow-up is now green in #55.

  • fix: observe outer Code Mode execute evidence #54 is merged into refactor/trusted-checkout-evidence and closes the authoritative outer Code Mode execute blocker.
  • fix: close final runtime evidence blockers #55 is rebased directly on that base and removes the target-writable authoritative capture file in favor of a runner-owned one-connection loopback stream.
  • The evaluated-shell tamper regression creates/deletes/recreates/appends the old /tmp/runtime/runtime-observer.jsonl path and passes only when the forged record cannot enter authoritative runtime_evidence.

Current #55 validation:

Both final review blockers are closed by the combined base + #55 state. PR #45 remains open/unmerged pending integration/review of #55.

## Scope

Focused corrective child of #45, targeting
`refactor/trusted-checkout-evidence`.

The base now includes merged #54 (`b10c8e9`), which closes the
authoritative outer Code Mode `execute` blocker. This PR is rebased
directly on that commit and contains only the remaining
authoritative-capture transport correction plus its regression coverage.

Stock OpenCode remains **2.0.23**. No threat-model expansion.

## Blocker 1 — outer Code Mode `execute`

**Closed in the base by #54 and revalidated here.**

The current integrated runtime evidence:

- observes the real synthetic outer `execute` at the supported stock
`execute.before` boundary;
- records Session, agent, message, real CallID, supported hook input,
and a unique runtime invocation ID;
- uses stock `session.tool.success` / `session.tool.failed` as terminal
authority;
- binds every inner Code Mode call to that exact observed outer
invocation;
- rejects dangling or identity-mismatched parents;
- keeps exact Code Mode caller-finality explicitly unsupported.

The Code Mode probe remains green after the capture-transport change.

## Blocker 2 — remove target-writable authoritative capture

The production observer no longer uses:

`/tmp/runtime/runtime-observer.jsonl`

Authoritative observer records now cross the process boundary through a
**runner-owned one-connection loopback stream**:

1. the runner binds an ephemeral listener on `127.0.0.1` immediately
before the real OpenCode invocation;
2. the trusted observer connects during plugin startup;
3. the listener closes after accepting that one connection;
4. the observer removes the endpoint from `process.env` before evaluated
tool subprocesses run;
5. sanitized observer records flow over the established stream;
6. the runner drains and validates those bytes through the existing
canonical builder/validator/accounting path.

A direct inherited memfd/FD was tried first, but stock
`@opencode/cli@2.0.23` crosses an internal process boundary that does
not preserve arbitrary extra descriptors. The one-connection stream
keeps stock OpenCode unchanged and avoids filesystem authority without
introducing security-platform machinery.

Protection against a deliberately malicious same-process plugin remains
explicitly out of scope.

## Tamper regression

Provider-free acceptance includes an evaluated tool that spawns
`/bin/sh` and:

- creates the old capture path;
- deletes it;
- recreates it;
- appends a forged observer record.

The scenario passes only if the shell tamper completes, authoritative
evidence remains complete/eligible, the real tamper tool is observed,
and the forged invocation/tool never appears in `runtime_evidence`.

## Preserved

- `opencode-eval-runner/runtime-evidence/v1`;
- canonical builder / validator / accounting;
- pre-sink credential sanitization;
- stock Session terminal authority;
- Code Mode exact caller-finality remains unsupported;
- normal `invoke` behavior;
- assertion-scoped eligibility;
- stock OpenCode 2.0.23 only.

No signing/HMAC, protected channel, remote PluginHost, capability
broker, hostile-plugin isolation, or OpenCode patch is introduced.

## Validation

Current clean head is based directly on merged #54 and is mergeable.

- **CI #329 / run 37532684965 — PASS**
  - **95/95 Python tests**
  - stock Code Mode probe: **PASS** (`diagnostics_passed: true`)
  - OpenCode **2.0.23**
- **Provider-free runtime evidence acceptance #43 / run 37532684985 —
PASS**
  - all **11/11** scenarios:
    - native_success
    - native_error
    - code_success
    - code_caught_error
    - concurrent_reverse
    - delegation
    - timeout
    - interrupted
    - redaction
    - collector
    - capture_tamper
- **Stock native observer integration #55 / run 37532684878 — PASS**
  - stock OpenCode 2.0.23
  - capture complete
  - provider-free native observation

## Merge scope

This PR is ready for review against
`refactor/trusted-checkout-evidence`.

Do **not** merge PR #45 as part of this change.

@bateau84 bateau84 left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Final independent review — READY TO MERGE

Reviewed integrated head:

e3eb017ff070f3956119472a7d0fb74faa247102

Both previously blocking correctness defects are closed:

  1. Outer Code Mode execute authority — the real model-facing synthetic execute invocation is observed at stock execute.before, settled by Session terminal events, and inner Code Mode calls bind to that exact observed outer invocation.
  2. Target-writable capture authority — production no longer consumes /tmp/runtime/runtime-observer.jsonl; authoritative records cross a runner-owned one-connection loopback stream, the listener closes after the trusted observer connects, and evaluated subprocess writes to the old path cannot affect evidence.

Final-head validation:

  • CI #330: PASS
  • Provider-free runtime evidence acceptance #44: PASS, including capture_tamper
  • Stock native observer integration #56: PASS
  • stock OpenCode 2.0.23 retained

The canonical runtime_evidence/v1 contract remains the sole eligibility authority. Missing/lost/malformed capture fails closed. Product outcome remains separate from evidence outcome. Exact Code Mode caller-final value/error remains explicitly unsupported. No OpenCode patch, signing/HMAC, protected channel, broker, remote PluginHost, or hostile-plugin isolation architecture has been reintroduced.

No BLOCKING, IMPORTANT, or MINOR correctness findings requiring changes remain.

Verdict: READY TO MERGE #45.

## Product Readiness finding

PR #45 makes `runtime_evidence` mandatory and the host CLI rejects a
result that omits or violates
`opencode-eval-runner/runtime-evidence/v1`, but the main README's Result
contract example still showed the older shape without that object. The
compatibility consequence for legacy/custom images was also not stated.

That leaves the primary consumer-facing contract contradictory even
though the runtime implementation is correct.

## Changes

- make `runtime_evidence` explicit in the README Result contract
example;
- state that every official transport result carries it, with Copilot
reporting `unsupported`;
- document that host runner + overridden/custom image are a
compatibility pair and must be upgraded together;
- add the assertion-level consumer rule: required boundaries must be
`complete` and required exact fields `available`; diagnostic convenience
fields cannot fill gaps;
- document the current aggregate and per-field capture bounds and their
fail-closed meaning.

## Scope

Documentation only. No runtime behavior, OpenCode patching, evidence
channel redesign, signing, broker, PluginHost, or hostile-plugin
isolation changes.

Targets `refactor/trusted-checkout-evidence` and is based directly on
reviewed PR #45 head `e3eb017ff070f3956119472a7d0fb74faa247102`.

@bateau84 bateau84 left a comment

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Focused Product Readiness re-pass — PRODUCT READY

Reviewed integrated head:

a67c931696d1ae2b2d16f49beab6c01c912aa02b

PR #56 is merged and closes the only Product Readiness blocker identified in the previous gate.

Evidence-readiness decision procedure

The public contract now gives consumers an explicit assertion-level procedure:

  1. validate runtime_evidence/v1;
  2. overall incomplete or invalid makes every runtime assertion ineligible for PASS;
  3. declare the boundary/boundaries required by the assertion and require each to be complete;
  4. if an assertion needs an exact field, require that field to be available; redacted/omitted are incomplete for that assertion and unsupported is unsupported;
  5. never backfill authoritative facts from tools, actions, tool_result_evidence, stdout/stderr, model text, or workspace files.

This is sufficient for Loom to make evidence-readiness decisions without moving judging semantics into the runner.

Product Readiness findings

BLOCKING: none.

IMPORTANT non-blocking follow-ups:

  • Loom must migrate its verdict path to consume runtime_evidence before using this capability to authorize PASS.
  • Aggregate capture-limit failures currently diagnose generically; dedicated capture_size_limit / capture_record_limit reason codes would improve operations but cannot create false PASS.
  • Any OpenCode upgrade beyond reviewed stock 2.0.23 should re-run the observer capability probes as an upgrade gate.

Accepted release limitations

  • exact Code Mode script-visible final value/error remains unsupported;
  • github-copilot-cli runtime observation remains unsupported;
  • deliberately malicious same-process evaluated plugins remain outside the trusted-checkout profile;
  • stock OpenCode 2.0.23 is the reviewed runtime.

Compatibility / migration

The README now states that runtime_evidence is mandatory, host and image form a compatibility pair, and legacy/custom images without valid v1 evidence are rejected. The primary Result contract therefore matches actual host behavior.

Exact-head validation

  • CI #332: PASS — 95/95 tests
  • Provider-free acceptance #45: PASS — all 11 scenarios including capture_tamper
  • Stock native observer #57: PASS
  • stock OpenCode 2.0.23 confirmed

No additional Product Readiness pass is required for PR #45 unless the runtime or public contract changes again. Loom should perform its own integration/acceptance pass when its scoring path is migrated.

PRODUCT READY — PR #45 may merge

## Purpose

Close the remaining user-facing documentation gaps before PR #45 merges.

The runner is an isolated **invocation execution boundary**, not an
eval-suite engine. This PR makes that contract explicit and documents
every supported public invocation surface.

## Changes

- add `docs/invocation-usage.md` as the authoritative usage/interface
reference;
- explicitly state that there is no input-JSON request API today;
- document the invocation input contract: CLI/Action arguments plus
prompt/system/config seed files;
- document all **24/24 CLI options**, including defaults and transport
applicability;
- document all runner-specific environment overrides and recognized
provider credentials;
- document all **21/21 GitHub Action inputs** and defaults;
- document Action precedence: repository `command` mode, direct
invocation mode, and setup-only mode;
- document which CLI capabilities are not exposed as direct Action
inputs;
- add a complete "ways to run" matrix;
- clarify that eval cases/assertions/judging remain owned by the calling
harness;
- correct the stale README description of the current zero-inference
plugin activation preflight;
- link the complete reference prominently from README local and Action
usage sections.

## Validation

Cross-checked the docs against the current implementation:

- CLI flags documented: **24/24**
- Action inputs documented: **21/21**
- runner-specific environment overrides documented: **8/8**

Documentation only; no runtime or contract behavior changes.

Targets `refactor/trusted-checkout-evidence` / PR #45.
@bateau84
bateau84 marked this pull request as ready for review October 7, 2026 07:59
@bateau84
bateau84 merged commit fd9da10 into main Oct 7, 2026
3 checks passed
@bateau84
bateau84 deleted the refactor/trusted-checkout-evidence branch October 7, 2026 07:59
bateau84 added a commit that referenced this pull request Oct 7, 2026
## Scope

Worker B for #58 (Task 1/9): post-PR-#45 runner contract/invariant
inventory.

This PR is intentionally documentation-only. It does **not** implement
the eval engine, change `invoke`, migrate Loom, add a public `eval` CLI,
or define the final Task 1 architecture.

## What it captures

- one `invoke` = exactly one isolated invocation;
- `opencode-eval-runner/v1` result boundary;
- mandatory/authoritative `runtime_evidence/v1`;
- assertion-scoped evidence readiness and field states;
- product vs evidence vs infrastructure failure separation;
- inner vs outer timeout behavior;
- retry ownership above `invoke`, cross-checked against Loom's current
narrow retry policy;
- OpenCode vs Copilot transport capability differences;
- host output validation and host/image compatibility;
- CLI and GitHub Action interface limits;
- explicit unsupported areas and trusted-checkout scope;
- current Loom behaviors that must not accidentally become runner
contracts.

## Validation

Cross-checked against the landed post-#45 contract:

- `main` / PR #45 merge commit
`fd9da10cbe2a8182fc8910ec199a221250deb3ca`
- final reviewed PR #45 head `25478106773969870931af3fb34e7b2586746606`
  - `runner/cli.py`
  - `container/invoke.py`
  - `docs/invocation-usage.md`
  - `docs/runtime-evidence-contract.md`
  - `action.yml`
- Loom `functionality-anchor-requirements-coherance`
- `scripts/run-evals.py` blob `30f2ae87764be180173a9e5b4b609ee79b2f063d`

`eval-engine/01-architecture` now matches `main` at the PR #45 merge
commit. The worker branch has been rebased onto that exact state,
leaving a one-file documentation diff.

Part of #58.
bateau84 added a commit that referenced this pull request Oct 7, 2026
## Summary

Part of #58 (Task 1/9), parallel worker A.

Adds a focused inventory of the current Loom eval harness from:
- bateau84/loom branch functionality-anchor-requirements-coherance
- scripts/run-evals.py at 30f2ae87764be180173a9e5b4b609ee79b2f063d

The inventory traces and classifies:
- case loading/normalization;
- target/workspace setup;
- invocation and orchestration retries;
- runtime evidence and legacy observers;
- judge construction/parsing;
- deterministic checks;
- classification;
- artifacts;
- iterations/concurrency;
- skill ablation;
- reporting.

It separates each responsibility into generic orchestration,
Loom/project semantics, low-level invoke behavior, or
legacy/compatibility behavior that should not migrate.

## Post-#45 validation

Cross-checked against the merged runner contracts on the integration
base:
- one invoke remains one isolated invocation;
- retries remain outside invoke;
- runtime_evidence/v1 is the only authoritative runtime-evidence object;
- tools/actions/tool_result_evidence/stdout/stderr/model text are
diagnostic only;
- project code owns cases, assertions, judging, thresholds, and
behavioral meaning.

The document explicitly marks Loom's older direct-container fallback,
observer/tool-result reconstruction, runner-safety adapter, and
Loom-side evidence_safety eligibility machinery as non-migration paths.

## Scope

Documentation only.

No runtime changes, no Loom migration, no final engine API/module
design, no universal assertion DSL, and no public eval CLI.

PR target: eval-engine/01-architecture.
bateau84 added a commit that referenced this pull request Oct 7, 2026
## Task

Closes #58 — **Task 1 of 9: Define generic orchestration boundary and
extraction contract**.

This PR is architecture/documentation only. It does not implement the
eval engine.

## Inputs integrated

Parallel worker results were merged into the integration branch first:

- #65 — runner contract/invariant inventory
- #66 — Loom eval-runner responsibility inventory

Both are based on the post-#45 merged `main` contract.

## Architecture decisions

The final synthesis defines:

- one `invoke` = exactly one isolated invocation;
- retry ownership exclusively in orchestration;
- `runtime_evidence/v1` as the only authoritative runtime-evidence
source;
- explicit separation of product, evidence, and infrastructure outcomes;
- generic ownership for planning, iterations, concurrency, attempts,
evidence-readiness mechanics, judge lifecycle, tri-state classification,
artifacts, and summaries;
- project/profile ownership for case semantics, fixtures, prompts,
evidence requirements, deterministic assertions, judge meaning,
thresholds, and ablation policy;
- explicit non-migration of Loom's old direct-container fallback, legacy
observer/evidence reconstruction, and superseded runner-safety paths;
- concrete internal module decomposition for Tasks 2–5;
- concrete internal contracts for normalized cases/jobs, invocation
specs, attempt records, retry policy, evidence requirements/readiness,
check outcomes, semantic decisions, and durable artifacts;
- initial standard/runtime concurrency-lane behavior;
- versioned internal run/artifact schemas;
- explicit Task 2–5 handoff boundaries.

No universal assertion DSL, public `eval` CLI, input JSON API, Loom
migration, or skill-ablation implementation is introduced.

## Files

- `docs/eval-engine-architecture.md` — final Task 1 synthesis
- `docs/eval-engine-runner-contracts.md` — worker B inventory
- `docs/loom-eval-runner-responsibility-inventory.md` — worker A
inventory

## Branch topology

- integration branch: `eval-engine/01-architecture`
- base: `main`
- worker branches were merged into the integration branch before
synthesis.

## Acceptance

The architecture is intended to let Tasks #59–#62 proceed independently
without redesigning the ownership boundary.

Documentation only; no runtime behavior changes.
bateau84 added a commit that referenced this pull request Oct 7, 2026
## Summary

Consolidates PR CI after merged #45 without changing runtime/evidence
semantics.

### Workflow cleanup

- keeps `.github/workflows/ci.yml` and `.github/workflows/publish.yml`;
- removes `.github/workflows/native-observer-integration.yml`;
- removes `.github/workflows/runtime-evidence-acceptance.yml`;
- keeps stable PR job names `test` and `runtime-integration`.

### `test`

Runs on every PR and keeps fast/general checks only:

- Python compile checks;
- full unit-test discovery;
- fake Copilot image + GitHub Action token/sanitization smoke test.

It no longer builds the real OpenCode or Copilot runtime images.

### `runtime-integration`

The job exists on every PR. It decides applicability internally, so
documentation-only/unrelated PRs still receive a successful
`runtime-integration` check instead of a missing required check.

For runtime-relevant changes it:

1. builds the stock OpenCode image once;
2. asserts OpenCode remains 2.0.23 and checks the pinned model-variant
interface;
3. runs the provider-free native observer probe;
4. runs the stock Code Mode capability probe;
5. runs the full provider-free runtime-evidence acceptance suite;
6. runs the OpenCode plugin activation preflight;
7. runs the config-root plugin dependency-resolution preflight;
8. builds/checks the real Copilot image only when shared/Copilot runtime
inputs changed;
9. uploads the consolidated integration reports.

This removes the prior three-way rebuild of the same stock OpenCode
image on runtime PRs.

### Dead code cleanup

Confirmed `loaded_skills_from_export` and `assistant_from_export` have
no production/runtime/runner/documented caller on current `main`; their
only remaining references were their own definitions and the two
obsolete Session-export unit tests.

Removed:

- `loaded_skills_from_export`;
- `assistant_from_export`;
- `test_failed_exported_skill_call_is_not_reported_as_loaded`;
- `test_session_export_preserves_tool_inputs_as_actions`.

Observable timing compatibility fields were left untouched because they
are output data and were not proven dead.

### Required-check recommendation

After this PR is merged, the repository ruleset should require exactly:

- `test`
- `runtime-integration`

Keep strict/up-to-date required status checks enabled.

`Publish images` must not be required.

The ruleset is **not changed in this PR**: the available repository
tools do not expose ruleset mutation, and enabling the new required
check before the workflow lands on `main` could block other PRs that
cannot yet produce it.

Follow-up in GitHub UI after merge:

`Settings → Rules → Rulesets → <active main ruleset> → Require status
checks to pass`

Keep `test`, add `runtime-integration`, keep strict/up-to-date enabled,
and do not add `Publish images`.

### Validation

Final head: `a5bb16e9bbe55efde1ab7a7f5eb9de8b07336495`

CI run **#341** / **37593353163**: **PASS**

- `test`: **PASS**
  - Python compile checks: PASS
  - unit tests: **93/93 PASS**
  - fake Copilot image + direct Action smoke: PASS
  - repository-command token sanitization smoke: PASS
- `runtime-integration`: **PASS**
  - stock OpenCode image built **once**: PASS
  - pinned OpenCode: **opencode v2.0.23**
  - native observer probe: PASS
- Code Mode capability probe: PASS, including explicit unsupported
caller-final value/error behavior
  - provider-free runtime-evidence acceptance: **11/11 scenarios PASS**
    - native success
    - native error
    - Code Mode success
    - caught Code Mode error
    - concurrent reverse completion
    - delegated Session ancestry
    - timeout
    - interrupted execution
    - credential redaction
    - forged collector-shaped payload rejection
    - capture tamper regression
  - OpenCode plugin activation preflight: PASS
  - config-root dependency-resolution preflight: PASS
  - real Copilot image build/version/interface check: PASS

The internal path decision excludes documentation-only changes such as
`README.md` / `docs/**`; for those PRs the `runtime-integration` job
still starts and succeeds through its explicit not-applicable path, so a
future required check will not hang.

No `runtime_evidence/v1`, trusted-checkout threat-model, or `invoke`
semantic changes were made.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant