Repository navigation
Jev (TypeSafe System One model): pre-answer classifier for scenario routing #17
Description
Activity
Update — the engine side of this landed.
SimulaMet/SimpleAudit#104 (structured choice answers + exact-match grading) was merged as SimpleAudit#106 and shipped in v0.4.0 today, and Studio's pin is now widened to allow it (
simpleaudit>=0.4.0,<0.5.0, commitcc01335).What the engine now provides that this issue depends on:
DecisionTarget— a System One decision model (OllamaPOST /v1/systemone, OpenRouter Decisions API) as a first-class audit target, withDecisionTarget.ollama(...)/DecisionTarget.openrouter(...)factory methods.- A scenario-level
decisionblock: fixed-option question + criteria; targets see it as data (no answer key), the full structured answer (choice, probabilities, confidence) is stored in the transcript. choice_matchjudge — code-only grading of the chosen option againstaccepted, no judge model, no judge API key. This is the "cheap, fast judge/quality-gate option" from the issue: calibrated probabilities are available in the result for inspecting/driving routing.
What's still left for Studio (this issue):
- Surface a "decision model" participant option in the target setup UI (Jev via OpenRouter, Clef via local Ollama), wired to
DecisionTarget+secret_referencefor the API key. - Expose the structured category + calibrated confidence in the audit result / run detail.
- The pre-answer routing use case (classify → branch in a scenario) — the engine covers grading, not branching, so scenario-routing support may need a design pass.
Suggested next step: a Studio PR that lets users register a
DecisionTargetendpoint through the model registry and seedecisionanswers in run detail.Thanks for the ping, and for landing the engine side. I'll take the Studio part: registering a DecisionTarget endpoint through the model registry (Clef via local Ollama first, Jev via OpenRouter behind secret_reference) and surfacing the structured answer — choice, probabilities, confidence — in run detail. PR to follow on this issue. I'm leaving the pre-answer routing (your item 3) out of that PR since it needs the design pass you mention.
Reacted by Sushant Gautam@avalyset — progress on the Studio side, in parallel with your PR: the model-registry half of item 1 is now in place (commit
c468a56), which should shrink the surface your PR needs to cover.What
c468a56does on /connections/:- Auto-detects decision models as they're registered, and persists the capability on
RegisteredModel.capabilities(with aRegisteredModel.is_decisionproperty + a badge in the model list and discover dialog).- Ollama — probed per model via
POST {api-root}/v1/systemonewith a 2-option question:200⇒ decision model, the"does not support decision"400 ⇒ not one (no inference spent on chat models). - OpenRouter — model-id prefix heuristic (
typesafe/,jev) since the Decisions catalog isn't publicly listed.
- Ollama — probed per model via
- So "register a
DecisionTargetendpoint through the model registry (Clef via local Ollama first, Jev via OpenRouter)" is largely already marked for you — the registry now knows which registered models are decision-capable.
What I'd still expect to land in your PR (not done here):
- Wiring the decision-capable registered model into the engine's
DecisionTargetfactory (DecisionTarget.ollama(...)/.openrouter(...)) in the target-setup flow +secret_referencefor the key. - Surfacing the structured answer — choice, probabilities, confidence — in run detail (item 2).
Happy to split or hand over whatever overlaps — let me know how you'd like to divide the registry-wiring vs. run-detail pieces.
- Auto-detects decision models as they're registered, and persists the capability on
Thanks — that fits well. I'll take both remaining pieces in the one PR: wiring decision-capable registered models (your
RegisteredModel.is_decision) intoDecisionTarget.ollama(...)/.openrouter(...)in the target-setup flow withsecret_referencefor the OpenRouter key, and surfacing choice, probabilities and confidence in run detail. Rebasing onto c468a56 now so the PR builds on your registry detection instead of duplicating it.One more refinement landed on this side —
f4e91d2(onmain):Detection now uses a single probe surface,
POST /v1/systemone, for every non-OpenRouter connection. Turns out that endpoint is the standard System One surface across providers (Ollama and vLLM both expose it at the API root; vLLM returns501when the model has no supported read strategy). So:- Any OpenAI-compatible server — Ollama, vLLM, anything — is probed once per model with a 2-option question.
200⇒ decision model;400 "does not support decision"(Ollama) or501(vLLM) ⇒ not one. No inference spent on non-decision models. - OpenRouter stays on the
typesafe//jevname heuristic: its/v1/systemoneis a paid proxy with no public model listing, so probing all ~469 listed models would bill real inference.
Practical implication for the
DecisionTargetwiring: when a model carriescapabilities.decision, the target endpoint is simply{base_url}/v1/systemonefor Ollama/vLLM/OpenAI-compatible servers — no per-provider endpoint logic needed in the target setup, only the auth handling (OpenRouter's key viasecret_reference).- Any OpenAI-compatible server — Ollama, vLLM, anything — is probed once per model with a 2-option question.
Full gap analysis for end-to-end decision-model support, for anyone picking pieces (splitting welcome):
Already on
main:- Engine 0.4.0 pin (
cc01335) - Auto-detection +
capabilities.decision+ badges on /connections/ (c468a56) - Unified
/v1/systemoneprobe for all non-OpenRouter providers, incl. vLLM 501 (f4e91d2)
Remaining, in suggested build order:
-
choice_matchjudge surfacing (small — verify + maybe fix)- Already in the engine's
JUDGE_CONFIGS, sojudges/services.bases()should mirror it automatically. To verify:ensure_starter_judgessurfaces it, and a run can go judge-less (code-only judge ⇒ no judge endpoint /secret_referencerequired —_validate_secretsininfra/engine.pymay currently demand one).
- Already in the engine's
-
DecisionTargetbridge wiring (core, worker-side)build_model_auditor()(infra/engine.py) always builds the stock OpenAI-compatibleModelTarget. When the target snapshot hascapabilities.decision, it needs to build aDecisionTargetinstead (endpoint{base_url}/v1/systemone, key via the snapshot'ssecret_reference, Ollama 64 KiB body cap via the.ollama()factory) and callauditor.set_target(...).max_turnsmust be forced to 1 (DecisionTarget.max_turns = 1) —auditor_kwargscurrently hardcodes it from the generation config.- Check whether
install_trace_context_target(infra/trace_target.py) coversDecisionTargettoo, or per-turn traceparent propagation is lost for decision runs.
-
Scenario
decisionblock (biggest piece)- New JSON field on
ScenarioRevision+scenario_dict()passthrough + form UI for a closed question (fixed options,acceptedanswer key that never reaches the target, optionalstate). - Validate 2–26 options at save time (the engine raises on broken limits).
- A scenario with a
decisionblock should require a decision-capable target.
- New JSON field on
-
Target picker + run detail (UI polish)
- New Experiment: show the decision badge next to registered models; flag/restrict decision models when the scenario set contains
decisionscenarios. - Run detail: surface
TargetResponse.decision— chosen option, per-option probabilities, confidence — not just thecontentstring (issue item 2).
- New Experiment: show the decision badge next to registered models; flag/restrict decision models when the scenario set contains
Overlap note: @avalyset — items 2 (target wiring) and 4 (run detail) look like the scope of your announced PR.
- Engine 0.4.0 pin (
Yes, let's split along those lines. I take 2 (DecisionTarget bridge in build_model_auditor via set_target, max_turns forced to 1, and checking install_trace_context_target covers DecisionTarget) and 4 (decision badge/restriction in the New Experiment picker, and choice/probabilities/confidence in run detail). You take 1 and 3. My PR won't touch ScenarioRevision or the scenario form: the end-to-end test builds its decision scenario as a dict in the test, so it stays mergeable before or after yours.
Upstream note for the bridge-wiring item: SimulaMet/SimpleAudit#108 (open) points DecisionTarget.openrouter() at /api/v1/systemone and drops the legacy /api/alpha/decisions URL. So once that lands in an engine release (0.4.1/0.5.0), the Studio bridge can use one canonical endpoint for every provider: {base_url}/v1/systemone for Ollama/vLLM/self-hosted, and the DecisionTarget.openrouter() factory for OpenRouter-hosted Jev. No alpha URL handling needed anywhere.
Studio PR: #20
Jev (TypeSafe AI's "System One Model") is a frontier model that makes fast, structured decisions: it outputs type-safe structured values with calibrated confidence scores in ~70–500ms at ~$0.042/MTok, and doesn't generate free text (so no hallucinated values).
Use case
A good fit for a pre-answer classification step in the Studio UI / audit flow. Example (medical assistant): before providing an answer, classify the patient's question into "health information" vs "patient rights" (etc.), with a confidence score, and route accordingly.
What this would look like in Studio
Engine-side exploration lives in SimulaMet/SimpleAudit#100 (API integration as target/auditor/judge + confidence-vs-judge comparison).