Skip to content

Support decision models as audit targets (structured choice answers, exact-match grading) - #106

Merged
SushantGautam merged 4 commits into
mainfrom
feature/decision-models
Oct 9, 2026
Merged

SushantGautam merged 4 commits into
mainfrom
feature/decision-models

Conversation

@SushantGautam

Copy link
Copy Markdown
Collaborator

First-class support for decision models — targets that read a state and a fixed-option question and return a chosen option with per-option probabilities instead of prose (e.g. Clef on Ollama's System One API, Jev on OpenRouter's Decisions API).

Implements the four changes proposed in #104:

  1. decision field in scenarios — an optional closed-question block; the answer key (accepted) never reaches the target; validated before any request.
  2. DecisionTarget — sends {model, state, questions} to a System One endpoint; .ollama() / .openrouter() factories; enforces the option/body limits; max_turns = 1; full answer (choice, probabilities, confidence) in TargetResponse.decision.
  3. choice_match judge — code-only exact-match grading with a grade hook (no judge model, no judge API key). Rejudge/reframing paths refuse code-graded judges.
  4. Docs + example — README "Decision Models" section, judge-table row, examples/decision_models_ollama.py.

Compatibility

  • Additive only: existing scenarios, targets and judges are unchanged. Every code path is guarded behind decision is not None / judge_grade is None.
  • No new dependencies.

Testing

  • Full suite green: 1108 passed, 3 skipped.
  • 86 new tests (test_decision.py, test_decision_target.py, test_choice_match_judge.py); new modules at 100% line coverage.

Authorship

These are the four feature commits from @KushtrimVisoka's proposal, taken from their fork branch decision-docs and rebased onto main:

  • 9a769fc optional decision block to scenarios
  • c0544f7f DecisionTarget for System One endpoints
  • ecd5f81b choice_match judge (code-only grading)
  • 5619261e docs + example

Closes #104.

A scenario may now state its question as one closed question with a fixed
set of options and, optionally, the accepted answer. The same scenario can
then be run against chat models and against decision models (models that
return a choice and probabilities rather than prose).

- simpleaudit/decision.py: validate_decision (unknown keys raise, as for
  document marks), public_decision (drops the answer key) and
  render_decision_prompt (the question as text for chat targets).
- run_scenario / run_async: validate every decision block before any
  request; render the question as the first prompt when a scenario has no
  test_prompt; pass the question without `accepted` to the target in
  TargetContext.extra["decision"]; pass the full block to the judge's
  post-processing in scenario_meta["decision"].
- scripts/check_scenario_pack.py: report invalid decision blocks as errors.
- Scenario guidelines 1.2: "Decision Field" section and its row in
  "What reaches the models".
- tests/test_decision.py: 31 tests. Scenarios without a decision block
  behave exactly as before.
Decision models (e.g. Clef served by Ollama at /v1/systemone, Jev on
OpenRouter's Decisions API) read a state and typed questions and return a
choice with probabilities instead of prose. DecisionTarget sends a
scenario's decision block to such an endpoint and maps the answer back.

- simpleaudit/targets/decision.py: builds {model, state, questions} from
  the decision block (without its answer key); the state is
  decision.state plus the text of the scenario's documents (marks are
  never sent), or the prompt when there is neither. Checks the option
  limit (2-26 by default) and the body limit (64 KiB for Ollama) before
  sending, sends compact UTF-8, reports endpoint errors with their
  message, and returns the chosen option as content ("yes: Found
  guilty") with the full answer in TargetResponse.decision.
  DecisionTarget.ollama() and DecisionTarget.openrouter() set the URL,
  key and limits.
- TargetResponse.decision: optional, None for every other target.
- ModelAuditor stores a decision answer beside the reply in the
  transcript, and caps a scenario at the target's max_turns (1 for
  DecisionTarget), warning once.
- tests/test_decision_target.py: 22 tests against an httpx mock of the
  endpoint. Also checked against a live Clef (Ollama 0.35.1).
A scenario's decision question has an exact answer, so it can be graded in
code: the chosen option against decision.accepted. No judge model is called
and no judge client is created, so no judge API key is needed and the verdict
is deterministic.

- simpleaudit/judges/choice_match.py: reads the choice from a decision
  model's answer in the transcript, or the option key on the first line of
  a chat reply (case, leading markdown and a Key:/Answer: label ignored;
  longer keys first). Accepted -> pass; a wrong or unrecognised option ->
  the scenario's designed severity (medium by default); no accepted keys ->
  UNGRADED. Records choice, its source, and a decision model's confidence
  and probabilities. Registered as "choice_match" (output "binary").
- Judge configs may declare `grade`: ModelAuditor calls it instead of a
  judge model, creates no judge client, and without an auditor_model creates
  no auditor client either (follow-up turns then fail with a clear message).
  An explicit judge_prompt switches back to a model judge. The run banner
  shows "code (<judge>)".
- reframing: PromptVariant.from_judge and rejudge refuse code-graded judges,
  because stored transcripts lack the scenario's decision block.
- customize_judge refuses new criteria for a code-graded judge.
- tests/test_choice_match_judge.py: 33 tests; registry tests list the new
  judge. Checked live against Clef (DecisionTarget) and a chat model.
- README: "Decision Models" section (the decision field, auditing a
  decision model with DecisionTarget and the choice_match judge, asking
  chat models the same questions, limits and the single-turn rule), and a
  choice_match row in the judge table.
- examples/decision_models_ollama.py: Clef (DecisionTarget) and a chat
  model on the same three synthetic scenarios, graded by choice_match;
  configurable with OLLAMA_HOST, DECISION_MODEL and CHAT_MODEL.
@SushantGautam
SushantGautam merged commit dbcaa39 into main Oct 9, 2026
3 checks passed
@SushantGautam
SushantGautam deleted the feature/decision-models branch October 9, 2026 11:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Support decision models as audit targets (structured choice answers, exact-match grading)

2 participants