Repository navigation
Conversation
A registered model that /connections/ probed as a decision model (RegisteredModel.is_decision) can now be used as an audit target, and the structured answer it gives back is shown in run detail. Items 2 and 4 of the gap analysis in issue SimulaMet#17. Target setup (infra/engine.py). The capability is frozen into the run's target snapshot, which already carries capabilities, so target setup never re-probes: a run uses what was true when it was frozen. The gate is that flag alone, never a list of known providers — detection probes every non-OpenRouter server at the same {base_url}/v1/systemone surface, so a provider added upstream later needs no change here. OpenRouter is the one exception: a fixed remote URL whose key is resolved from the connection's secret_reference, like every other role. A decision target replaces the auditor's target outright rather than being wrapped by the trace-context adapter, which wraps an OpenAI-compatible chat client a decision target does not use. Nothing is lost by that: DecisionTarget merges TargetContext.trace_headers into its own request headers, so per-turn traceparent still propagates (pinned by a test). The unused AnyLLM chat client is not built at all. SimpleAudit's own Auditor facade does the same, and building it would demand the provider's any_llm extra and an API key for a client that is never called — a local Ollama decision run otherwise fails on any-llm-sdk[ollama] not being installed. max_turns is forced to 1, since DecisionTarget answers once and a follow-up turn would have nothing to send; forced in auditor_kwargs so both run paths agree, the repetition runner deriving its model entry and max_retries_per_rep from the same kwargs. Run detail renders the answer per reply: the chosen option, the calibrated confidence, and a bar per option sorted by probability. The New Experiment picker badges decision-capable models, with the constraint spelled out per role. - model_registry/decision: is_decision_snapshot, build_decision_target, decision_target_for_model/_snapshot, decision_answer - infra/engine: decision_target_for, install_target, skipping_target_client, max_turns forced for decision snapshots - infra/trace_target: decision_auditor_class for the repetition runner - infra/ui: _rep_view carries the answer per message and per repetition - templates: "Decision answer" card, decision badge in the model picker - 54 tests, including an end-to-end replay of a reply recorded from clef-flash (infra/fixtures/decision_clef_flash.json) and a live test that skips without a local Ollama serving it Refs SimulaMet#17, SimulaMet/SimpleAudit#106
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Covers items 2 and 4 from the gap analysis in #17 (comment 6081082149); items 1 and 3 are SushantGautam's. Builds on c468a56 and f4e91d2.
What
A registered model that
/connections/probed as a decision model (RegisteredModel.is_decision) can now be used as an audit target, and the structured answer it returns — choice, per-option probabilities, calibrated confidence — is shown in run detail.Item 2 —
DecisionTargetbridge (infra/engine.py,infra/trace_target.py)build_model_auditor()builds aDecisionTargetand callsauditor.set_target(...)when the target snapshot carriescapabilities.decision. The snapshot already copiesRegisteredModel.capabilities(audits.services._endpoint_snapshot), so target setup never re-probes: a run uses what was true when it was frozen.{base_url}/v1/systemonesurface, target setup needs no per-provider endpoint logic either — a provider you add later works here unchanged. The.ollama()factory builds that URL and applies the 64 KiB body cap; a connection's own/v1suffix comes off first, the same normalisationprobe_systemone_decisiondoes. OpenRouter is the one exception: a fixed remote URL whose key is resolved from the connection'ssecret_reference, like every other role, so a raw key never enters a snapshot.max_turnsforced to 1. Done inauditor_kwargsrather than at either construction site, so both run paths agree — the repetition runner derives its model entry andmax_retries_per_repfrom the same kwargs.install_trace_context_targetandDecisionTarget— checked. It does not cover one, and must not be applied to a decision run: it wrapsauditor.target_clientin aTraceContextModelTarget, a ModelTarget over an OpenAI-compatible chat client, which would replace theDecisionTargetoutright. Nothing is lost by skipping it —DecisionTarget.sendmergesTargetContext.trace_headersinto its own request headers, so per-turn traceparent still propagates.test_a_decision_target_forwards_the_per_turn_traceparentpins that, so it fails if the engine ever stops doing it.Auditorfacade does the same (_skip_target_client), and building it would demand the provider's any_llm extra and an API key for a client that is never called: a local Ollama decision run otherwise fails onany-llm-sdk[ollama]not being installed.decision_auditor_class).Item 4 — picker and run detail (
infra/ui.py,templates/)decisionbadge as on /connections/, with the constraint spelled out per role — a single-turn run needing a scenario with a decision block for the target, "answers options, not prose" for judge and auditor._rep_viewcarries the answer per reply (reply["decision"], whichModelAuditorcopies fromTargetResponse.decision) and per repetition, a later turn winning. The result panel renders a "Decision answer" card: the chosen option,confidence N%when the model reported one, and a probability bar per option, highest first. A chat reply has no such block and renders unchanged.Not in this PR
choice_matchjudge surfacing; the scenariodecisionblock). Nothing here touchesScenarioRevision,scenario_dict()or the scenario form — the end-to-end test builds its decision scenario as a dict in the test, so this is mergeable before or after that work.decisionfield from item 3, so the badge is advisory for now. And a decision run is judge-less in practice —_validate_secretsdemanding a judgesecret_referenceis item 1's call, which is why the tests here construct the auditor with the code-onlychoice_matchjudge directly rather than through a run's judge snapshot.Tested
Tested end-to-end against clef-flash (9.1B, Q8_0) on local Ollama 0.35.1; the recorded response is the CI fixture (
infra/fixtures/decision_clef_flash.json). CI runs Python 3.12.The fixture is replayed through an
httpx.MockTransport, so everything but the HTTP hop is the real chain: frozen snapshot →DecisionTarget→ engine run with the code-onlychoice_matchjudge →_rep_view→ rendered panel.test_live_clef_flash_answers_and_rendersruns the same path against the real server and skips when no local Ollama serves the model.54 new tests. Ruff clean. Five pre-existing failures (
chat/tests/test_sync.pysignal counts,infra/tests/test_db_resilience.pySQLite locking) and four errors (infra/tests/test_minimal_config.pyembedded Hatchet lifecycle) reproduce identically on a cleanf4e91d2checkout, so they are not from this branch.Refs #17
AI assistance. Implemented with Claude Code (Opus 5); planning and review with Claude Fable 5.1. The author ran the test suite, the live clef-flash run and the diff review locally before submission.