Summary
SimpleAudit audits chat models. There is a growing class of decision models that work differently: they read a document (a "state") and a question with a fixed set of options, and return the chosen option with a probability for every option, without writing prose. Examples:
I'd like to propose first-class support for them, so that one scenario pack can audit both decision models and chat models on the same questions, and compare them in the same results.
Motivation: I maintain Iustitia Marigonae, a benchmark of closed questions about Kosovo court judgments (Albanian), which I already publish as a SimpleAudit scenario pack for chat models. I want to run the decision models on it through SimpleAudit too.
Why CallableTarget is not enough today
I first bridged a decision model with CallableTarget. It works, but:
- The target never sees the question as data. It receives only the prompt text and the history, so the adapter has to recognise the prompt and look the question and its options up elsewhere.
- Grading a fixed-option answer with an LLM judge costs tokens and adds noise where an exact comparison is possible. Small local judges (4–8B) often failed to return a usable checklist for these answers.
Proposal
Four small changes. Each is self-contained with tests, and nothing changes for existing scenarios, targets or judges.
1. An optional decision field in scenarios
"decision": {
"id": "verdict",
"instructions": "Did the court find the defendant guilty?",
"criteria": {"yes": "Found guilty", "no": "Not found guilty"},
"accepted": ["yes"], # never sent to the target
"state": {"jurisdiction": "Kosovo"}, # optional, decision models only
}
- It follows the same two rules as document marks: unknown keys raise, and the answer key (
accepted) never reaches the target.
- Targets receive the question without
accepted in TargetContext.extra["decision"], so the send() signature does not change.
- Judges: post-processing code gets the full block in
scenario_meta["decision"].
- Chat models: a scenario without a
test_prompt asks the question as text, with the options listed and the chosen key requested on the first line.
- Checks:
run_async validates every block before any request, and check_scenario_pack.py reports invalid blocks. The guidelines get a "Decision Field" section (version 1.2).
2. DecisionTarget
- Protocol: sends the question to a System One endpoint as
{model, state, questions}. DecisionTarget.ollama(...) and DecisionTarget.openrouter(...) set the URL, key and limits.
- State:
decision.state plus the text of the scenario's documents; document marks are never sent.
- Limits: checked before sending (2–26 options; 64 KiB per request for Ollama). A scenario over a limit is recorded as an error; documents are never shortened.
- Answer: the content is the chosen option (
"yes: Found guilty"), so judges and the summary work unchanged. The full answer (choice, probabilities, confidence) goes in a new optional TargetResponse.decision and is stored beside the reply in the transcript.
- Turns:
max_turns = 1; the auditor caps scenarios at a target's max_turns and warns once.
3. A choice_match judge, graded in code
- Grading: compares the chosen option with
accepted. The choice comes from a decision model's answer, or from the option key on the first line of a chat reply.
- accepted →
pass;
- wrong or unrecognised → the scenario's designed severity;
- no
accepted → ungraded.
- No judge model: judge configs may declare a
grade function. The auditor then calls it instead of a judge model and creates no judge client, so no judge API key is needed.
- Refusals: the re-grading paths (
rejudge, PromptVariant.from_judge) refuse such configs, because stored transcripts lack the decision block. customize_judge refuses new criteria for them.
4. Documentation and an example
A README section "Decision Models", a choice_match row in the judge table, and examples/decision_models_ollama.py.
Compatibility
- Existing scenarios, targets and judges behave exactly as before, and all existing tests pass.
- No new dependencies.
- 86 new tests. The new modules have full line coverage.
Tested
Against Clef on Ollama 0.35.1, with DecisionTarget, and a chat model through Ollama's OpenAI-compatible API, both graded by choice_match with no judge API key. Output of the example on three synthetic scenarios:
Scenario Accepted clef-64k gemma4:e2b
------------------------------------------------------------------------------
Verdict - Guilty yes yes (0.95) pass yes pass
Verdict - Acquitted no no (0.95) pass no pass
Sentence - Suspended yes yes (0.88) pass yes pass
The implementation is on my fork: KushtrimVisoka/SimpleAudit@main...decision-docs (four commits, one per change above).
Questions
- Does the
decision field fit the scenario schema as proposed, or would you prefer it somewhere else (e.g. under metadata)?
- Is a
grade hook on judge configs (code-only grading, no judge model) acceptable, since it touches how results are judged?
- Would you like this as four stacked pull requests, in the order above, or as one?
Happy to adjust the design before opening pull requests.
Summary
SimpleAudit audits chat models. There is a growing class of decision models that work differently: they read a document (a "state") and a question with a fixed set of options, and return the chosen option with a probability for every option, without writing prose. Examples:
POST /v1/systemone(System One API)I'd like to propose first-class support for them, so that one scenario pack can audit both decision models and chat models on the same questions, and compare them in the same results.
Motivation: I maintain Iustitia Marigonae, a benchmark of closed questions about Kosovo court judgments (Albanian), which I already publish as a SimpleAudit scenario pack for chat models. I want to run the decision models on it through SimpleAudit too.
Why
CallableTargetis not enough todayI first bridged a decision model with
CallableTarget. It works, but:Proposal
Four small changes. Each is self-contained with tests, and nothing changes for existing scenarios, targets or judges.
1. An optional
decisionfield in scenariosaccepted) never reaches the target.acceptedinTargetContext.extra["decision"], so thesend()signature does not change.scenario_meta["decision"].test_promptasks the question as text, with the options listed and the chosen key requested on the first line.run_asyncvalidates every block before any request, andcheck_scenario_pack.pyreports invalid blocks. The guidelines get a "Decision Field" section (version 1.2).2.
DecisionTarget{model, state, questions}.DecisionTarget.ollama(...)andDecisionTarget.openrouter(...)set the URL, key and limits.decision.stateplus the text of the scenario'sdocuments; document marks are never sent."yes: Found guilty"), so judges and the summary work unchanged. The full answer (choice, probabilities, confidence) goes in a new optionalTargetResponse.decisionand is stored beside the reply in the transcript.max_turns = 1; the auditor caps scenarios at a target'smax_turnsand warns once.3. A
choice_matchjudge, graded in codeaccepted. The choice comes from a decision model's answer, or from the option key on the first line of a chat reply.pass;accepted→ungraded.gradefunction. The auditor then calls it instead of a judge model and creates no judge client, so no judge API key is needed.rejudge,PromptVariant.from_judge) refuse such configs, because stored transcripts lack the decision block.customize_judgerefuses new criteria for them.4. Documentation and an example
A README section "Decision Models", a
choice_matchrow in the judge table, andexamples/decision_models_ollama.py.Compatibility
Tested
Against Clef on Ollama 0.35.1, with
DecisionTarget, and a chat model through Ollama's OpenAI-compatible API, both graded bychoice_matchwith no judge API key. Output of the example on three synthetic scenarios:The implementation is on my fork: KushtrimVisoka/SimpleAudit@main...decision-docs (four commits, one per change above).
Questions
decisionfield fit the scenario schema as proposed, or would you prefer it somewhere else (e.g. undermetadata)?gradehook on judge configs (code-only grading, no judge model) acceptable, since it touches how results are judged?Happy to adjust the design before opening pull requests.