diff --git a/README.md b/README.md index 50d001d..f93709b 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ # opencode-eval-runner -Reusable OCI isolation for behavioral evals that invoke OpenCode or GitHub Copilot CLI. +Reusable OCI execution boundary for isolated eval invocations using OpenCode or GitHub Copilot CLI. The runner intentionally does **not** own an eval corpus, grading semantics, or agent policy. Those stay in the repository being evaluated. This project owns the execution boundary: @@ -28,6 +28,25 @@ eval harness The caller may use the same model for both, but they do not share OpenCode session state or filesystem state. +## What this repository runs + +This project runs **one isolated invocation at a time**. It is not an eval-suite engine: the calling repository owns cases, iterations, assertions, target/judge orchestration, grading, and final PASS/FAIL policy. + +The current public input interface is CLI or GitHub Action arguments plus prompt/system files; there is **no input JSON request API**. Each invocation produces the documented JSON Result contract. + +Supported usage modes are: + +| Mode | Interface | +| --- | --- | +| Local OpenCode invocation | `opencode-eval-runner invoke --transport opencode ...` | +| Local Copilot invocation | `opencode-eval-runner invoke --transport github-copilot-cli ...` | +| Local eval suite | Repository harness repeatedly calls the CLI | +| GitHub Action direct invocation | Action inputs with `model` set | +| GitHub Action repository harness | Action `command` mode | +| GitHub Action setup only | Leave `model` and `command` empty, then call the CLI in a later step | + +See **[Invocation usage and interface reference](docs/invocation-usage.md)** for the complete input contract, all CLI options, environment variables, all GitHub Action inputs, direct-mode limitations, and execution examples. + ## Transports ### `opencode` @@ -95,6 +114,8 @@ This transport reuses the trust-boundary pattern already proven in `nrkno/mats-o ## Local usage +For the full option/default/reference table, see [Invocation usage and interface reference](docs/invocation-usage.md). + Build the transport you need: ```bash @@ -173,7 +194,7 @@ opencode-eval-runner invoke \ ... ``` -Stock OpenCode 2.0.23 does not expose the old singular `debug agent ` command that returned a resolved tool map. The runner therefore performs the strongest supported zero-inference preflight: it requires the expected plugin entrypoint to be materialized in the isolated OpenCode config, runs `opencode debug agents` to prove the configured location starts successfully with plugins active, and requires the selected agent to resolve. Missing plugin materialization, plugin/startup failure, or missing agent is infrastructure/non-evidence, never a behavioral FAIL. Actual tool use remains a repository-owned behavioral assertion in the eval corpus. +The expected-plugin preflight is zero-inference. The runner requires the plugin entrypoint to be materialized in the isolated config, starts a private stock OpenCode 2.0.23 server, creates a non-resuming Session prompt so plugin activation reaches its barrier without model inference, and verifies that the named plugin is present and active in the runtime plugin inventory. Agent resolution remains part of the real `opencode run --agent` invocation. Missing materialization, activation failure, or later agent resolution failure is infrastructure/non-evidence, never a behavioral FAIL. Actual tool use remains a repository-owned behavioral assertion in the eval corpus. ### Evaluating a skill @@ -216,6 +237,8 @@ The token value is not placed on the container command line. ## GitHub Actions +For all 21 Action inputs, defaults, execution-mode precedence, and CLI-only capabilities, see [Invocation usage and interface reference](docs/invocation-usage.md#github-action-interface). + The repository is a composite GitHub Action. It supports either a single direct invocation or setup plus a repository-owned eval harness. For a direct Copilot invocation, the action uses the workflow's built-in `GITHUB_TOKEN`; no separate Copilot secret is needed. The calling workflow must grant `copilot-requests: write`: diff --git a/docs/invocation-usage.md b/docs/invocation-usage.md new file mode 100644 index 0000000..5e82aa9 --- /dev/null +++ b/docs/invocation-usage.md @@ -0,0 +1,287 @@ +# Invocation usage and interface reference + +This repository provides an isolated **invocation execution boundary** for eval harnesses. It does not own eval cases, suites, assertions, judging semantics, thresholds, or final PASS/FAIL policy. + +An eval harness uses this runner to execute one target or judge invocation at a time: + +~~~text +eval harness + -> invocation inputs + -> opencode-eval-runner invoke + -> isolated OpenCode or GitHub Copilot CLI process + -> result/v1 + runtime-evidence/v1 + -> eval harness assertions / judging / verdict +~~~ + +## Input contract + +There is currently **no input JSON request API**. + +The supported invocation input contracts are: + +1. the host CLI: opencode-eval-runner invoke ...; +2. the GitHub Action inputs in action.yml; +3. a repository-owned harness that calls the host CLI one or more times. + +Prompt and system content are supplied as UTF-8 files rather than as a JSON request body. + +The absence of an input JSON schema is intentional in the current interface. Do not assume stdin JSON or an invocation/v1 request envelope is supported. + +### Input files + +| Input | Required | Applies to | Meaning | +| --- | --- | --- | --- | +| --prompt-file PATH | Yes | Both transports | UTF-8 user/task prompt copied into the isolated invocation. | +| --system-file PATH | No | GitHub Copilot CLI | UTF-8 system/agent instructions. The OpenCode transport currently does not consume this file. | +| --config PATH | No | OpenCode | Explicit OpenCode JSON config seed. The host OpenCode config is not inherited automatically. | +| --auth PATH | No | OpenCode | Explicit legacy OpenCode auth JSON seed. | +| --database PATH | No | OpenCode | OpenCode V2 database source. The runner copies only credential rows and migration journals into a sanitized temporary database. | +| --models-catalog PATH | No | OpenCode | Explicit OpenCode model-catalog seed. | +| --config-root PATH | No | OpenCode | Explicit OpenCode config-root source used by the runtime plugin compatibility bridge. The current implementation materializes the supported local Loom plugin tree from plugins/loom when present. | + +## Ways to run + +| Mode | Interface | Typical use | +| --- | --- | --- | +| Local single OpenCode invocation | Host CLI | Debug or run one agent/tool-aware eval invocation with authoritative OpenCode runtime evidence. | +| Local single Copilot invocation | Host CLI | Run a model/role/judge invocation that does not require OpenCode runtime evidence. | +| Local eval suite | Repository-owned harness calling the CLI repeatedly | Compose cases, target/judge invocations, assertions, iterations, and verdicts outside the runner. | +| GitHub Action direct invocation | Action inputs with model set and command empty | Run one isolated invocation directly from a workflow. | +| GitHub Action repository harness | Action command input | Prepare the runner/images, then execute the repository's own eval harness. | +| GitHub Action setup only | Leave both model and command empty | Pull/build selected images and put the runner on PATH for later workflow steps. | + +Directly invoking the transport container entrypoint is an implementation detail, not a supported public interface. Use the host CLI or GitHub Action so workspace mounts, credential seeding, result validation, and runtime-evidence handling stay consistent. + +## Host CLI reference + +The only current subcommand is: + +~~~text +opencode-eval-runner invoke [options] +~~~ + +### Execution and transport + +| Option | Required / default | Applies to | Description | +| --- | --- | --- | --- | +| --transport {opencode,github-copilot-cli} | Default: opencode | Both | Selects the invocation transport. | +| --engine {auto,podman,docker} | Default: auto | Both | OCI engine. auto prefers Podman, then Docker. | +| --network MODE | Optional | Both | Explicit OCI network mode/name such as host, bridge, slirp4netns, or a custom network. | +| --image IMAGE | Optional | Both | Override the selected transport image. Host runner and custom image must be contract-compatible. | +| --workspace PATH | Default: . | Both | Workspace mounted at /workspace. | +| --workspace-mode {ro,rw} | Default: ro | Both | Workspace bind-mount mode. | +| --mount SOURCE:TARGET[:ro|rw] | Repeatable | Both | Add an explicit bind mount. Default mode is read-only. | + +### Model and invocation + +| Option | Required / default | Applies to | Description | +| --- | --- | --- | --- | +| --model MODEL | **Required** | Both | Model identifier passed to the selected transport. | +| --reasoning LEVEL | Optional | Both | OpenCode maps this to the model #variant; Copilot maps it to --effort. | +| --agent NAME | Optional | OpenCode | Select the OpenCode agent. Copilot uses its fixed isolated eval-runner profile. | +| --skill ID | Optional | OpenCode only | Records the skill under test. It does not force the skill to load. | +| --expected-plugin NAME | Optional | OpenCode only | Fail closed before inference unless the named plugin is materialized and active in the isolated OpenCode runtime. | +| --prompt-file PATH | **Required** | Both | UTF-8 invocation prompt. | +| --system-file PATH | Optional | Copilot | UTF-8 system instructions consumed by the Copilot transport. | +| --env NAME | Repeatable | Both | Explicitly forward an additional host environment variable when it is set. | + +### OpenCode seeds + +| Option | Required / default | Applies to | Description | +| --- | --- | --- | --- | +| --auth PATH | Optional; otherwise env/default discovery | OpenCode | Legacy auth JSON seed. | +| --database PATH | Optional; otherwise env/default discovery | OpenCode | V2 database source; sanitized before container use. | +| --models-catalog PATH | Optional; otherwise env/default discovery | OpenCode | Model-catalog seed. | +| --config PATH | Optional | OpenCode | Explicit OpenCode config JSON. | +| --config-root PATH | Optional | OpenCode | Explicit config-root source for supported local plugin materialization. | + +### Timeouts and result handling + +| Option | Required / default | Applies to | Description | +| --- | --- | --- | --- | +| --timeout-seconds N | Default: 240 | Both | Timeout for the model/OpenCode/Copilot invocation inside the container. | +| --container-timeout N | Default: 300 | Both | Outer host timeout for the OCI process. | +| --output PATH | **Required** | Both | Host path where the validated result JSON is written. | +| --print-result | Default: off | Both | Also print the validated result JSON to stdout. | + +### CLI examples + +OpenCode: + +~~~bash +opencode-eval-runner invoke \ + --transport opencode \ + --workspace /path/to/project \ + --model openai/gpt-5.5 \ + --agent general \ + --prompt-file /tmp/prompt.txt \ + --output /tmp/result.json +~~~ + +Copilot: + +~~~bash +opencode-eval-runner invoke \ + --transport github-copilot-cli \ + --workspace /path/to/project \ + --model gpt-5.4 \ + --prompt-file /tmp/prompt.txt \ + --system-file /tmp/system.txt \ + --output /tmp/judgment.json +~~~ + +A local eval harness simply invokes the CLI repeatedly for its target and judge calls. The harness, not this runner, owns case IDs, iterations, expected behavior, assertions, grading, and final verdicts. + +## Environment-variable reference + +### Runner configuration + +| Variable | Meaning | +| --- | --- | +| OPENCODE_EVAL_RUNNER_OPENCODE_IMAGE | Default OpenCode image override when --image is not supplied. | +| OPENCODE_EVAL_RUNNER_COPILOT_IMAGE | Default Copilot image override when --image is not supplied. | +| OPENCODE_EVAL_RUNNER_IMAGE | Generic fallback image override used after the transport-specific override. | +| OPENCODE_EVAL_RUNNER_AUTH | OpenCode auth JSON seed path. | +| OPENCODE_EVAL_RUNNER_DB | OpenCode V2 database source path. | +| OPENCODE_EVAL_RUNNER_MODELS | OpenCode model-catalog seed path. | +| OPENCODE_EVAL_RUNNER_CONFIG | Explicit OpenCode config JSON path. | +| OPENCODE_EVAL_RUNNER_CONFIG_ROOT | Explicit OpenCode config-root source path. | + +Explicit CLI paths take precedence over their environment equivalents. + +### Automatically recognized provider credentials + +For OpenCode, these host variables are forwarded automatically when set: + +- OPENAI_API_KEY +- ANTHROPIC_API_KEY +- OPENROUTER_API_KEY + +Use --env NAME for any additional variable. + +For GitHub Copilot CLI, authentication precedence is: + +1. COPILOT_GITHUB_TOKEN +2. GH_TOKEN +3. GITHUB_TOKEN +4. authenticated host gh auth token fallback + +The token value is passed through the child-process environment, not placed on the OCI command line. + +## GitHub Action interface + +The Action has three behaviors: + +1. command non-empty: setup images/runner, then run the repository-owned command. This takes precedence over direct invocation inputs. +2. command empty and model non-empty: run one direct isolated invocation. +3. both empty: setup only; no invocation is run. + +### Action inputs + +| Input | Default | Meaning | +| --- | --- | --- | +| opencode-image | ghcr.io/bateau84/opencode-eval-runner:opencode-edge | OpenCode transport image. | +| copilot-image | ghcr.io/bateau84/opencode-eval-runner:copilot-edge | Copilot transport image. | +| prepare-opencode | true | Pull/build the OpenCode image during setup. | +| prepare-copilot | true | Pull/build the Copilot image during setup. | +| engine | docker | Container engine used by action invocations/builds. | +| build | false | Build transport images from the action checkout instead of pulling published images. | +| transport | opencode | Direct-invocation transport. Ignored when command is supplied. | +| model | empty | Direct-invocation model. A non-empty value triggers direct mode when command is empty. | +| reasoning | empty | Optional reasoning level/effort. | +| agent | empty | Optional OpenCode agent. | +| skill | empty | Optional OpenCode skill ID under test. | +| workspace | . | Direct-invocation workspace. | +| workspace-mode | ro | Direct workspace mount mode. | +| prompt-file | empty | Direct prompt file. Required when model is set. | +| system-file | empty | Optional system prompt file. Consumed by the Copilot transport. | +| output | .opencode-evals/result.json | Direct-invocation result path. | +| mounts | empty | Newline-separated SOURCE:TARGET[:ro|rw] mounts. | +| timeout-seconds | 240 | Inner model invocation timeout. | +| container-timeout | 300 | Outer container timeout. | +| command | empty | Repository-owned eval command; takes precedence over direct invocation. | +| artifact-dir | .opencode-evals | Host artifact directory created during Action setup. | + +### Action direct-mode coverage + +The Action direct mode intentionally exposes a smaller interface than the host CLI. + +The following CLI capabilities are **not direct Action inputs**: + +- --network +- --expected-plugin +- --auth +- --database +- --models-catalog +- --config +- --config-root +- repeatable arbitrary --env +- --print-result + +For OpenCode seed paths, workflows can use the documented OPENCODE_EVAL_RUNNER_* environment variables where applicable. If a workflow needs the full CLI option surface, use setup-only or command mode and invoke opencode-eval-runner explicitly. + +### GitHub Action direct invocation + +~~~yaml +- uses: bateau84/opencode-eval-runner@main + with: + transport: opencode + model: openai/gpt-5.5 + agent: general + workspace: . + prompt-file: .github/evals/prompt.txt + output: .opencode-evals/result.json +~~~ + +### GitHub Action repository-owned harness + +~~~yaml +- uses: bateau84/opencode-eval-runner@main + with: + command: | + python3 scripts/run-evals.py --cases WORK-01,REVIEW-01 +~~~ + +### GitHub Action setup only + +~~~yaml +- uses: bateau84/opencode-eval-runner@main + with: + prepare-opencode: true + prepare-copilot: false + +- run: | + opencode-eval-runner invoke \ + --transport opencode \ + --model openai/gpt-5.5 \ + --prompt-file .github/evals/prompt.txt \ + --output .opencode-evals/result.json +~~~ + +## Output contract + +Each invocation produces one validated opencode-eval-runner/v1 JSON result. The complete representative result is documented in the repository README. + +Every official result includes runtime_evidence: + +- OpenCode results can contain authoritative runtime-evidence/v1 observations. +- Copilot results report runtime evidence as explicitly unsupported. +- the host validates runtime_evidence before persisting the result artifact. + +For runtime assertions, follow the evidence-readiness decision procedure in [runtime-evidence-contract.md](runtime-evidence-contract.md). Diagnostic fields such as tools, actions, tool_result_evidence, stdout/stderr, and model text cannot substitute for missing authoritative evidence. + +## What belongs outside this runner + +An eval harness or repository remains responsible for: + +- case definitions and IDs; +- iterations/repetitions; +- target-versus-judge orchestration; +- behavioral assertions; +- rubrics and judge prompts; +- expected results; +- thresholds; +- aggregation; +- PASS/FAIL/non-evidence decisions. + +This separation is deliberate: the runner provides faithful isolated invocations and runtime evidence; the consumer decides what those observations mean for the eval.