Skip to content

Support decision models as audit targets (structured choice answers, exact-match grading) #104

Description

@KushtrimVisoka

Summary

SimpleAudit audits chat models. There is a growing class of decision models that work differently: they read a document (a "state") and a question with a fixed set of options, and return the chosen option with a probability for every option, without writing prose. Examples:

I'd like to propose first-class support for them, so that one scenario pack can audit both decision models and chat models on the same questions, and compare them in the same results.

Motivation: I maintain Iustitia Marigonae, a benchmark of closed questions about Kosovo court judgments (Albanian), which I already publish as a SimpleAudit scenario pack for chat models. I want to run the decision models on it through SimpleAudit too.

Why CallableTarget is not enough today

I first bridged a decision model with CallableTarget. It works, but:

  1. The target never sees the question as data. It receives only the prompt text and the history, so the adapter has to recognise the prompt and look the question and its options up elsewhere.
  2. Grading a fixed-option answer with an LLM judge costs tokens and adds noise where an exact comparison is possible. Small local judges (4–8B) often failed to return a usable checklist for these answers.

Proposal

Four small changes. Each is self-contained with tests, and nothing changes for existing scenarios, targets or judges.

1. An optional decision field in scenarios

"decision": {
    "id": "verdict",
    "instructions": "Did the court find the defendant guilty?",
    "criteria": {"yes": "Found guilty", "no": "Not found guilty"},
    "accepted": ["yes"],                 # never sent to the target
    "state": {"jurisdiction": "Kosovo"}, # optional, decision models only
}
  • It follows the same two rules as document marks: unknown keys raise, and the answer key (accepted) never reaches the target.
  • Targets receive the question without accepted in TargetContext.extra["decision"], so the send() signature does not change.
  • Judges: post-processing code gets the full block in scenario_meta["decision"].
  • Chat models: a scenario without a test_prompt asks the question as text, with the options listed and the chosen key requested on the first line.
  • Checks: run_async validates every block before any request, and check_scenario_pack.py reports invalid blocks. The guidelines get a "Decision Field" section (version 1.2).

2. DecisionTarget

  • Protocol: sends the question to a System One endpoint as {model, state, questions}. DecisionTarget.ollama(...) and DecisionTarget.openrouter(...) set the URL, key and limits.
  • State: decision.state plus the text of the scenario's documents; document marks are never sent.
  • Limits: checked before sending (2–26 options; 64 KiB per request for Ollama). A scenario over a limit is recorded as an error; documents are never shortened.
  • Answer: the content is the chosen option ("yes: Found guilty"), so judges and the summary work unchanged. The full answer (choice, probabilities, confidence) goes in a new optional TargetResponse.decision and is stored beside the reply in the transcript.
  • Turns: max_turns = 1; the auditor caps scenarios at a target's max_turns and warns once.

3. A choice_match judge, graded in code

  • Grading: compares the chosen option with accepted. The choice comes from a decision model's answer, or from the option key on the first line of a chat reply.
    • accepted → pass;
    • wrong or unrecognised → the scenario's designed severity;
    • no accepted → ungraded.
  • No judge model: judge configs may declare a grade function. The auditor then calls it instead of a judge model and creates no judge client, so no judge API key is needed.
  • Refusals: the re-grading paths (rejudge, PromptVariant.from_judge) refuse such configs, because stored transcripts lack the decision block. customize_judge refuses new criteria for them.

4. Documentation and an example

A README section "Decision Models", a choice_match row in the judge table, and examples/decision_models_ollama.py.

Compatibility

  • Existing scenarios, targets and judges behave exactly as before, and all existing tests pass.
  • No new dependencies.
  • 86 new tests. The new modules have full line coverage.

Tested

Against Clef on Ollama 0.35.1, with DecisionTarget, and a chat model through Ollama's OpenAI-compatible API, both graded by choice_match with no judge API key. Output of the example on three synthetic scenarios:

Scenario               Accepted   clef-64k                 gemma4:e2b
------------------------------------------------------------------------------
Verdict - Guilty       yes        yes (0.95) pass          yes pass
Verdict - Acquitted    no         no (0.95) pass           no pass
Sentence - Suspended   yes        yes (0.88) pass          yes pass

The implementation is on my fork: KushtrimVisoka/SimpleAudit@main...decision-docs (four commits, one per change above).

Questions

  1. Does the decision field fit the scenario schema as proposed, or would you prefer it somewhere else (e.g. under metadata)?
  2. Is a grade hook on judge configs (code-only grading, no judge model) acceptable, since it touches how results are judged?
  3. Would you like this as four stacked pull requests, in the order above, or as one?

Happy to adjust the design before opening pull requests.

Activity

  1. SushantGautam commented on Oct 9, 2026

    @SushantGautam
    Collaborator

    Merged — thank you @KushtrimVisoka! 🎉

    Your four changes are in as PR #106, taken straight from your decision-docs branch and rebased onto main. All three CI jobs (3.11 / 3.12 / 3.13) are green and the full suite passes (1108 passed).

    To answer your open questions, as we resolved them for the merge:

    1. decision field placement — kept as a top-level decision block on the scenario (not under metadata), following the same rules as document marks.
    2. grade hook on judge configs — accepted; code-only grading with no judge model or judge API key is exactly what makes these scenarios cheap and deterministic.
    3. PR structure — merged as a single PR (the four commits, in order), since they build on each other.

    On availability: it's now on main, and will ship with the next release — so pip install --upgrade simpleaudit will bring it in as soon as that version is published.

    Nice work on Iustitia Marigonae — happy to have decision models first-class here.

  2. SushantGautam commented on Oct 9, 2026

    @SushantGautam
    Collaborator

    Update: v0.4.0 is already published on PyPI — so no waiting.

    pip install --upgrade simpleaudit
    

    gives you the decision-models feature right now. (Release notes)

  3. SushantGautam commented on Oct 9, 2026

    @SushantGautam
    Collaborator

    Related follow-up: #108 points DecisionTarget.openrouter() at the promoted /api/v1/systemone route (dropping the legacy /api/alpha/decisions URL). Same System One wire format; v1 is the superset (bare jev-1.13 ids auto-map to typesafe/, provider routing, usage.cost) and matches the exact route Ollama and vLLM expose.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions