Skip to content

Jev (TypeSafe System One model): pre-answer classifier for scenario routing #17

Description

@SushantGautam

Jev (TypeSafe AI's "System One Model") is a frontier model that makes fast, structured decisions: it outputs type-safe structured values with calibrated confidence scores in ~70–500ms at ~$0.042/MTok, and doesn't generate free text (so no hallucinated values).

Use case

A good fit for a pre-answer classification step in the Studio UI / audit flow. Example (medical assistant): before providing an answer, classify the patient's question into "health information" vs "patient rights" (etc.), with a confidence score, and route accordingly.

What this would look like in Studio

  • New participant option: "Jev (TypeSafe)" alongside LLM targets/judges, configured with an API key from console.typesafe.ai (early access).
  • Expose the structured category + calibrated confidence in the audit result / run detail, so it's inspectable and can drive branching in a scenario.
  • Surface it as a cheap, fast judge/quality-gate option (calibrated probabilities are useful for scoring without a full LLM judge call).

Engine-side exploration lives in SimulaMet/SimpleAudit#100 (API integration as target/auditor/judge + confidence-vs-judge comparison).

Activity

  1. SushantGautam commented on Oct 7, 2026

    @SushantGautam
    CollaboratorAuthor
  2. SushantGautam commented on Oct 9, 2026

    @SushantGautam
    CollaboratorAuthor

    Update — the engine side of this landed.

    SimulaMet/SimpleAudit#104 (structured choice answers + exact-match grading) was merged as SimpleAudit#106 and shipped in v0.4.0 today, and Studio's pin is now widened to allow it (simpleaudit>=0.4.0,<0.5.0, commit cc01335).

    What the engine now provides that this issue depends on:

    • DecisionTarget — a System One decision model (Ollama POST /v1/systemone, OpenRouter Decisions API) as a first-class audit target, with DecisionTarget.ollama(...) / DecisionTarget.openrouter(...) factory methods.
    • A scenario-level decision block: fixed-option question + criteria; targets see it as data (no answer key), the full structured answer (choice, probabilities, confidence) is stored in the transcript.
    • choice_match judge — code-only grading of the chosen option against accepted, no judge model, no judge API key. This is the "cheap, fast judge/quality-gate option" from the issue: calibrated probabilities are available in the result for inspecting/driving routing.

    What's still left for Studio (this issue):

    1. Surface a "decision model" participant option in the target setup UI (Jev via OpenRouter, Clef via local Ollama), wired to DecisionTarget + secret_reference for the API key.
    2. Expose the structured category + calibrated confidence in the audit result / run detail.
    3. The pre-answer routing use case (classify → branch in a scenario) — the engine covers grading, not branching, so scenario-routing support may need a design pass.

    Suggested next step: a Studio PR that lets users register a DecisionTarget endpoint through the model registry and see decision answers in run detail.

  3. avalyset commented on Oct 9, 2026

    @avalyset

    Thanks for the ping, and for landing the engine side. I'll take the Studio part: registering a DecisionTarget endpoint through the model registry (Clef via local Ollama first, Jev via OpenRouter behind secret_reference) and surfacing the structured answer — choice, probabilities, confidence — in run detail. PR to follow on this issue. I'm leaving the pre-answer routing (your item 3) out of that PR since it needs the design pass you mention.

  4. SushantGautam commented on Oct 9, 2026

    @SushantGautam
    CollaboratorAuthor

    @avalyset — progress on the Studio side, in parallel with your PR: the model-registry half of item 1 is now in place (commit c468a56), which should shrink the surface your PR needs to cover.

    What c468a56 does on /connections/:

    • Auto-detects decision models as they're registered, and persists the capability on RegisteredModel.capabilities (with a RegisteredModel.is_decision property + a badge in the model list and discover dialog).
      • Ollama — probed per model via POST {api-root}/v1/systemone with a 2-option question: 200 ⇒ decision model, the "does not support decision" 400 ⇒ not one (no inference spent on chat models).
      • OpenRouter — model-id prefix heuristic (typesafe/, jev) since the Decisions catalog isn't publicly listed.
    • So "register a DecisionTarget endpoint through the model registry (Clef via local Ollama first, Jev via OpenRouter)" is largely already marked for you — the registry now knows which registered models are decision-capable.

    What I'd still expect to land in your PR (not done here):

    • Wiring the decision-capable registered model into the engine's DecisionTarget factory (DecisionTarget.ollama(...) / .openrouter(...)) in the target-setup flow + secret_reference for the key.
    • Surfacing the structured answer — choice, probabilities, confidence — in run detail (item 2).

    Happy to split or hand over whatever overlaps — let me know how you'd like to divide the registry-wiring vs. run-detail pieces.

  5. avalyset commented on Oct 9, 2026

    @avalyset

    Thanks — that fits well. I'll take both remaining pieces in the one PR: wiring decision-capable registered models (your RegisteredModel.is_decision) into DecisionTarget.ollama(...) / .openrouter(...) in the target-setup flow with secret_reference for the OpenRouter key, and surfacing choice, probabilities and confidence in run detail. Rebasing onto c468a56 now so the PR builds on your registry detection instead of duplicating it.

  6. SushantGautam commented on Oct 9, 2026

    @SushantGautam
    CollaboratorAuthor

    One more refinement landed on this side — f4e91d2 (on main):

    Detection now uses a single probe surface, POST /v1/systemone, for every non-OpenRouter connection. Turns out that endpoint is the standard System One surface across providers (Ollama and vLLM both expose it at the API root; vLLM returns 501 when the model has no supported read strategy). So:

    • Any OpenAI-compatible server — Ollama, vLLM, anything — is probed once per model with a 2-option question. 200 ⇒ decision model; 400 "does not support decision" (Ollama) or 501 (vLLM) ⇒ not one. No inference spent on non-decision models.
    • OpenRouter stays on the typesafe//jev name heuristic: its /v1/systemone is a paid proxy with no public model listing, so probing all ~469 listed models would bill real inference.

    Practical implication for the DecisionTarget wiring: when a model carries capabilities.decision, the target endpoint is simply {base_url}/v1/systemone for Ollama/vLLM/OpenAI-compatible servers — no per-provider endpoint logic needed in the target setup, only the auth handling (OpenRouter's key via secret_reference).

  7. SushantGautam commented on Oct 9, 2026

    @SushantGautam
    CollaboratorAuthor

    Full gap analysis for end-to-end decision-model support, for anyone picking pieces (splitting welcome):

    Already on main:

    • Engine 0.4.0 pin (cc01335)
    • Auto-detection + capabilities.decision + badges on /connections/ (c468a56)
    • Unified /v1/systemone probe for all non-OpenRouter providers, incl. vLLM 501 (f4e91d2)

    Remaining, in suggested build order:

    1. choice_match judge surfacing (small — verify + maybe fix)

      • Already in the engine's JUDGE_CONFIGS, so judges/services.bases() should mirror it automatically. To verify: ensure_starter_judges surfaces it, and a run can go judge-less (code-only judge ⇒ no judge endpoint / secret_reference required — _validate_secrets in infra/engine.py may currently demand one).
    2. DecisionTarget bridge wiring (core, worker-side)

      • build_model_auditor() (infra/engine.py) always builds the stock OpenAI-compatible ModelTarget. When the target snapshot has capabilities.decision, it needs to build a DecisionTarget instead (endpoint {base_url}/v1/systemone, key via the snapshot's secret_reference, Ollama 64 KiB body cap via the .ollama() factory) and call auditor.set_target(...).
      • max_turns must be forced to 1 (DecisionTarget.max_turns = 1) — auditor_kwargs currently hardcodes it from the generation config.
      • Check whether install_trace_context_target (infra/trace_target.py) covers DecisionTarget too, or per-turn traceparent propagation is lost for decision runs.
    3. Scenario decision block (biggest piece)

      • New JSON field on ScenarioRevision + scenario_dict() passthrough + form UI for a closed question (fixed options, accepted answer key that never reaches the target, optional state).
      • Validate 2–26 options at save time (the engine raises on broken limits).
      • A scenario with a decision block should require a decision-capable target.
    4. Target picker + run detail (UI polish)

      • New Experiment: show the decision badge next to registered models; flag/restrict decision models when the scenario set contains decision scenarios.
      • Run detail: surface TargetResponse.decision — chosen option, per-option probabilities, confidence — not just the content string (issue item 2).

    Overlap note: @avalyset — items 2 (target wiring) and 4 (run detail) look like the scope of your announced PR.

  8. avalyset commented on Oct 9, 2026

    @avalyset

    Yes, let's split along those lines. I take 2 (DecisionTarget bridge in build_model_auditor via set_target, max_turns forced to 1, and checking install_trace_context_target covers DecisionTarget) and 4 (decision badge/restriction in the New Experiment picker, and choice/probabilities/confidence in run detail). You take 1 and 3. My PR won't touch ScenarioRevision or the scenario form: the end-to-end test builds its decision scenario as a dict in the test, so it stays mergeable before or after yours.

  9. SushantGautam commented on Oct 9, 2026

    @SushantGautam
    CollaboratorAuthor

    Upstream note for the bridge-wiring item: SimulaMet/SimpleAudit#108 (open) points DecisionTarget.openrouter() at /api/v1/systemone and drops the legacy /api/alpha/decisions URL. So once that lands in an engine release (0.4.1/0.5.0), the Studio bridge can use one canonical endpoint for every provider: {base_url}/v1/systemone for Ollama/vLLM/self-hosted, and the DecisionTarget.openrouter() factory for OpenRouter-hosted Jev. No alpha URL handling needed anywhere.

  10. avalyset commented on Oct 9, 2026

    @avalyset

    Studio PR: #20

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions