Skip to content

Bridge decision-capable registered models to DecisionTarget and surface the answer in run detail - #20

Open
avalyset wants to merge 1 commit into
SimulaMet:mainfrom
avalyset:feat/decision-target-registry
Open

avalyset wants to merge 1 commit into
SimulaMet:mainfrom
avalyset:feat/decision-target-registry

Conversation

@avalyset

@avalyset avalyset commented Oct 9, 2026 •

Copy link
Copy Markdown

Covers items 2 and 4 from the gap analysis in #17 (comment 6081082149); items 1 and 3 are SushantGautam's. Builds on c468a56 and f4e91d2.

What

A registered model that /connections/ probed as a decision model (RegisteredModel.is_decision) can now be used as an audit target, and the structured answer it returns — choice, per-option probabilities, calibrated confidence — is shown in run detail.

Item 2 — DecisionTarget bridge (infra/engine.py, infra/trace_target.py)

  • The bridge. build_model_auditor() builds a DecisionTarget and calls auditor.set_target(...) when the target snapshot carries capabilities.decision. The snapshot already copies RegisteredModel.capabilities (audits.services._endpoint_snapshot), so target setup never re-probes: a run uses what was true when it was frozen.
  • No provider allow-list. The gate is that flag alone. Since f4e91d2 probes every non-OpenRouter server at the same {base_url}/v1/systemone surface, target setup needs no per-provider endpoint logic either — a provider you add later works here unchanged. The .ollama() factory builds that URL and applies the 64 KiB body cap; a connection's own /v1 suffix comes off first, the same normalisation probe_systemone_decision does. OpenRouter is the one exception: a fixed remote URL whose key is resolved from the connection's secret_reference, like every other role, so a raw key never enters a snapshot.
  • max_turns forced to 1. Done in auditor_kwargs rather than at either construction site, so both run paths agree — the repetition runner derives its model entry and max_retries_per_rep from the same kwargs.
  • install_trace_context_target and DecisionTarget — checked. It does not cover one, and must not be applied to a decision run: it wraps auditor.target_client in a TraceContextModelTarget, a ModelTarget over an OpenAI-compatible chat client, which would replace the DecisionTarget outright. Nothing is lost by skipping it — DecisionTarget.send merges TargetContext.trace_headers into its own request headers, so per-turn traceparent still propagates. test_a_decision_target_forwards_the_per_turn_traceparent pins that, so it fails if the engine ever stops doing it.
  • The unused chat client is not built. SimpleAudit's own Auditor facade does the same (_skip_target_client), and building it would demand the provider's any_llm extra and an API key for a client that is never called: a local Ollama decision run otherwise fails on any-llm-sdk[ollama] not being installed.
  • Both run paths. The single-repetition path, and the repetition runner, which builds its own auditors and so needs the choice made on the class (decision_auditor_class).

Item 4 — picker and run detail (infra/ui.py, templates/)

  • New Experiment picker: decision-capable models carry the same decision badge as on /connections/, with the constraint spelled out per role — a single-turn run needing a scenario with a decision block for the target, "answers options, not prose" for judge and auditor.
  • Run detail: _rep_view carries the answer per reply (reply["decision"], which ModelAuditor copies from TargetResponse.decision) and per repetition, a later turn winning. The result panel renders a "Decision answer" card: the chosen option, confidence N% when the model reported one, and a probability bar per option, highest first. A chat reply has no such block and renders unchanged.

Not in this PR

  • Items 1 and 3 (choice_match judge surfacing; the scenario decision block). Nothing here touches ScenarioRevision, scenario_dict() or the scenario form — the end-to-end test builds its decision scenario as a dict in the test, so this is mergeable before or after that work.
  • Dependency on item 1: restricting the picker by which scenarios a set contains needs the decision field from item 3, so the badge is advisory for now. And a decision run is judge-less in practice — _validate_secrets demanding a judge secret_reference is item 1's call, which is why the tests here construct the auditor with the code-only choice_match judge directly rather than through a run's judge snapshot.

Tested

Tested end-to-end against clef-flash (9.1B, Q8_0) on local Ollama 0.35.1; the recorded response is the CI fixture (infra/fixtures/decision_clef_flash.json). CI runs Python 3.12.

The fixture is replayed through an httpx.MockTransport, so everything but the HTTP hop is the real chain: frozen snapshot → DecisionTarget → engine run with the code-only choice_match judge → _rep_view → rendered panel. test_live_clef_flash_answers_and_renders runs the same path against the real server and skips when no local Ollama serves the model.

54 new tests. Ruff clean. Five pre-existing failures (chat/tests/test_sync.py signal counts, infra/tests/test_db_resilience.py SQLite locking) and four errors (infra/tests/test_minimal_config.py embedded Hatchet lifecycle) reproduce identically on a clean f4e91d2 checkout, so they are not from this branch.

Refs #17

AI assistance. Implemented with Claude Code (Opus 5); planning and review with Claude Fable 5.1. The author ran the test suite, the live clef-flash run and the diff review locally before submission.

A registered model that /connections/ probed as a decision model
(RegisteredModel.is_decision) can now be used as an audit target, and the
structured answer it gives back is shown in run detail. Items 2 and 4 of
the gap analysis in issue SimulaMet#17.

Target setup (infra/engine.py). The capability is frozen into the run's
target snapshot, which already carries capabilities, so target setup never
re-probes: a run uses what was true when it was frozen. The gate is that
flag alone, never a list of known providers — detection probes every
non-OpenRouter server at the same {base_url}/v1/systemone surface, so a
provider added upstream later needs no change here. OpenRouter is the one
exception: a fixed remote URL whose key is resolved from the connection's
secret_reference, like every other role.

A decision target replaces the auditor's target outright rather than being
wrapped by the trace-context adapter, which wraps an OpenAI-compatible chat
client a decision target does not use. Nothing is lost by that:
DecisionTarget merges TargetContext.trace_headers into its own request
headers, so per-turn traceparent still propagates (pinned by a test).

The unused AnyLLM chat client is not built at all. SimpleAudit's own
Auditor facade does the same, and building it would demand the provider's
any_llm extra and an API key for a client that is never called — a local
Ollama decision run otherwise fails on any-llm-sdk[ollama] not being
installed. max_turns is forced to 1, since DecisionTarget answers once and
a follow-up turn would have nothing to send; forced in auditor_kwargs so
both run paths agree, the repetition runner deriving its model entry and
max_retries_per_rep from the same kwargs.

Run detail renders the answer per reply: the chosen option, the calibrated
confidence, and a bar per option sorted by probability. The New Experiment
picker badges decision-capable models, with the constraint spelled out per
role.

- model_registry/decision: is_decision_snapshot, build_decision_target,
  decision_target_for_model/_snapshot, decision_answer
- infra/engine: decision_target_for, install_target,
  skipping_target_client, max_turns forced for decision snapshots
- infra/trace_target: decision_auditor_class for the repetition runner
- infra/ui: _rep_view carries the answer per message and per repetition
- templates: "Decision answer" card, decision badge in the model picker
- 54 tests, including an end-to-end replay of a reply recorded from
  clef-flash (infra/fixtures/decision_clef_flash.json) and a live test
  that skips without a local Ollama serving it

Refs SimulaMet#17, SimulaMet/SimpleAudit#106

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants