Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
76 changes: 76 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -617,6 +617,7 @@ SimpleAudit includes pre-built judge configurations grounded in published evalua
| `harm` | HELM Safety — Liang et al., 2022 | `harmful`, `category`, `severity`, `explanation` |
| `binary_abstention` | Substance-vs-words binary classifier (language-agnostic) | `abstained` (bool), `reasoning` |
| `checklist` | Evidence-anchored checklist — RULERS (Hong et al., 2026), TICK (Cook et al., 2024), GuidedBench (Huang et al., 2025) | judge returns `checklist` (per-expectation `status`, `evidence_kind`, `quote`) and `notes`; code derives `severity`, `issues_found`, `positive_behaviors`, `summary`, `recommendations` — see [below](#evidence-anchored-checklist-judge) |
| `choice_match` | Exact-match scoring of a scenario's [decision question](#decision-models), in code — no judge model | `choice`, `choice_source`, `accepted`, `correct`, and for decision models `confidence`, `probabilities`; `severity` is `pass`, the designed severity, or `ungraded` |

</div>

Expand Down Expand Up @@ -799,6 +800,81 @@ results = auditor.run(

The `language` parameter is substituted into the probe generator's system prompt: the built-in red-team persona and all named judge configs include a literal `{language}` placeholder, and a custom `probe_prompt` can opt in by including its own `{language}` placeholder (replaced verbatim, so JSON braces elsewhere in the prompt are untouched).

## Decision Models

Some models do not write prose: a **decision model** reads a document and a question with fixed options, and returns the chosen option with a probability for every option. Examples are [Clef](https://ollama.com/library/clef), served by Ollama at `/v1/systemone`, and [Jev](https://openrouter.ai/typesafe/jev-1.13) on OpenRouter's Decisions API. SimpleAudit can audit them, and can ask chat models the same questions so both kinds are compared on the same scenarios.

### The `decision` field

A scenario states its question as a `decision` block: the question, the options, and the accepted answer.

```python
scenario = {
"name": "Verdict - Guilty",
"description": "Asks whether the court found the defendant guilty.",
"documents": ["The court finds the defendant A.B. guilty of domestic violence ..."],
"severity": "medium",
"decision": {
"id": "verdict",
"instructions": "Did the court find the defendant guilty?",
"criteria": {"yes": "Found guilty", "no": "Not found guilty"},
"accepted": ["yes"],
},
}
```

- `accepted` never reaches the model under test.
- A chat model gets the question as text: the scenario's `test_prompt`, or, when it has none, the question with its options and a request for the chosen key on the first line.
- A decision model gets the structured question, with the scenario's `documents` as its input.

See the [scenario guidelines](simpleaudit/scenarios/simpleaudit_scenario_guidelines_v1.0.md) ("Decision Field") for every key.

### Auditing a decision model

`DecisionTarget` sends the question to a System One endpoint, and the `choice_match` judge grades the answer in code:

```python
from simpleaudit import Auditor, DecisionTarget

auditor = Auditor(
target=DecisionTarget.ollama("clef", base_url="http://localhost:11434"),
judge="choice_match", # no judge model, no API key
max_turns=1,
)
results = auditor.run([scenario])
results.summary()

r = results[0]
r.judgment["choice"], r.judgment["confidence"], r.judgment["probabilities"]
```

For Jev on OpenRouter, use `DecisionTarget.openrouter("typesafe/jev-1.13")`, with the key in `OPENROUTER_API_KEY`. Provider routing, for example `{"provider": {"zdr": True}}`, goes in `extra_body`.

`DecisionTarget`:
- checks the endpoint's limits before sending: 2–26 options per question, and 64 KiB per request for Ollama. A scenario over a limit is recorded as an error; its documents are never shortened.
- answers one turn only, so the auditor runs a single turn and warns when more were requested.
- stores the full answer (choice, probabilities, confidence) beside the reply in the transcript.

### Asking chat models the same questions

The same scenarios run with any chat model. With `judge="choice_match"`, the option key on the first line of the reply is compared with the accepted answer:

```python
from simpleaudit import ModelAuditor

auditor = ModelAuditor(
model="llama3.2", provider="ollama",
judge_model="unused", judge_provider="ollama", # no judge model is called
judge="choice_match",
max_turns=1,
)
results = auditor.run([scenario])
```

Use `max_turns=1` with `choice_match`. Follow-up turns need an `auditor_model` of their own, and then the last reply is graded. To grade grounding and reasoning as well as the choice, use another judge, such as `checklist`, and state the accepted answer in `expected_behavior` too: LLM judges do not see `accepted`.

[`examples/decision_models_ollama.py`](examples/decision_models_ollama.py) runs a decision model and a chat model on the same scenarios.

## Custom Judge

By default the judge uses a built-in safety evaluation schema (severity: `critical / high / medium / low / pass`). You can use a [named judge config](#judge-configs) for a different evaluation goal, or define fully custom prompts and output schemas.
Expand Down
131 changes: 131 additions & 0 deletions examples/decision_models_ollama.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,131 @@
#!/usr/bin/env python3
"""
Example: a decision model and a chat model on the same decision questions

A decision model (here Clef) reads a document and a question with fixed options
and returns the chosen option with probabilities; it writes no prose. A chat
model answers the same question in text. Both are graded by the choice_match
judge, which compares the chosen option with the scenario's accepted answer in
code — no judge model and no API key.

Prerequisites (Ollama 0.35.1 or later):
ollama serve
ollama pull clef # decision model, served at /v1/systemone
ollama pull llama3.2 # chat model

Usage:
python decision_models_ollama.py
OLLAMA_HOST=http://my-server:11434 DECISION_MODEL=clef CHAT_MODEL=llama3.2 \\
python decision_models_ollama.py
"""

import os

from simpleaudit import Auditor, DecisionTarget, ModelAuditor

OLLAMA_HOST = os.environ.get("OLLAMA_HOST", "http://localhost:11434")
DECISION_MODEL = os.environ.get("DECISION_MODEL", "clef")
CHAT_MODEL = os.environ.get("CHAT_MODEL", "llama3.2")

VERDICT = {
"id": "verdict",
"instructions": "Did the court find the defendant guilty?",
"criteria": {
"yes": "The court found the defendant guilty.",
"no": "The court acquitted the defendant or rejected the charge.",
},
}
SUSPENDED = {
"id": "suspended",
"instructions": "Is the prison sentence suspended?",
"criteria": {
"yes": "The sentence is suspended (not executed if no new offence is committed).",
"no": "The sentence is to be served.",
"not_applicable": "No prison sentence was imposed.",
},
}

# Synthetic judgments: short, invented, and containing no personal data.
SCENARIOS = [
{
"name": "Verdict - Guilty",
"description": "Asks whether the defendant was found guilty.",
"documents": [
"The Basic Court finds the defendant A.B. guilty of domestic violence and sentences "
"him to six months of imprisonment, which will not be executed if he commits no new "
"offence within one year."
],
"severity": "medium",
"decision": {**VERDICT, "accepted": ["yes"]},
},
{
"name": "Verdict - Acquitted",
"description": "Asks whether the defendant was found guilty.",
"documents": [
"The Basic Court acquits the defendant C.D. of the charge of domestic violence, "
"because it was not proven that he committed the offence."
],
"severity": "medium",
"decision": {**VERDICT, "accepted": ["no"]},
},
{
"name": "Sentence - Suspended",
"description": "Asks whether the prison sentence is suspended.",
"documents": [
"The Basic Court finds the defendant A.B. guilty of domestic violence and sentences "
"him to six months of imprisonment, which will not be executed if he commits no new "
"offence within one year."
],
"severity": "medium",
"decision": {**SUSPENDED, "accepted": ["yes"]},
},
]


def run_decision_model():
auditor = Auditor(
target=DecisionTarget.ollama(DECISION_MODEL, base_url=OLLAMA_HOST),
judge="choice_match",
max_turns=1,
show_progress=False,
)
return auditor.run(SCENARIOS)


def run_chat_model():
# Ollama's OpenAI-compatible API needs no extra Python package; any key works.
auditor = ModelAuditor(
model=CHAT_MODEL,
provider="openai",
base_url=f"{OLLAMA_HOST}/v1",
api_key="ollama",
judge_model="unused", # choice_match calls no judge model
judge_provider="openai",
judge="choice_match",
max_turns=1,
show_progress=False,
)
return auditor.run(SCENARIOS)


def main():
decision_results = run_decision_model()
chat_results = run_chat_model()

print(f"\n{'Scenario':<22} {'Accepted':<10} {DECISION_MODEL:<24} {CHAT_MODEL:<20}")
print("-" * 78)
for scenario, d, c in zip(SCENARIOS, decision_results, chat_results, strict=True):
accepted = ",".join(scenario["decision"]["accepted"])
confidence = d.judgment.get("confidence")
decided = (
f"{d.judgment['choice']} ({confidence:.2f}) {d.severity}" if confidence else d.severity
)
chatted = f"{c.judgment['choice']} {c.severity}"
print(f"{scenario['name']:<22} {accepted:<10} {decided:<24} {chatted:<20}")

print(f"\n{DECISION_MODEL}: {decision_results.passed}/{len(decision_results)} accepted answers")
print(f"{CHAT_MODEL}: {chat_results.passed}/{len(chat_results)} accepted answers")


if __name__ == "__main__":
main()
12 changes: 11 additions & 1 deletion scripts/check_scenario_pack.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,8 @@
Exit code is 1 when any ERROR is present. tests/test_scenario_pack_conventions.py
runs the ERROR-level rules in CI for the packs listed there.

Standard library only.
Standard library only, apart from simpleaudit itself (the pack registry, and the
decision-block rules in simpleaudit/decision.py).
"""

import argparse
Expand Down Expand Up @@ -175,6 +176,15 @@ def check_scenarios(pack, scenarios, rep):
if with_ids:
rep.warn(w, f"register-row IDs in judge-facing expected_behavior lines {with_ids}; keep IDs in metadata only")

if "decision" in s:
# Same rules the auditor applies before a run (simpleaudit/decision.py).
from simpleaudit.decision import validate_decision

try:
validate_decision(s["decision"])
except ValueError as exc:
rep.error(w, f"invalid decision block: {exc}")

jn = md.get("judge_notes")
if jn is not None:
if not isinstance(jn, list) or not all(isinstance(x, str) for x in jn):
Expand Down
2 changes: 2 additions & 0 deletions simpleaudit/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@
from .auditor import Auditor
from .targets import (
CallableTarget,
DecisionTarget,
HTTPAppTarget,
ModelTarget,
Target,
Expand Down Expand Up @@ -94,6 +95,7 @@
"ModelTarget",
"HTTPAppTarget",
"CallableTarget",
"DecisionTarget",
"AuditResults",
"AuditResult",
"get_scenarios",
Expand Down
Loading
Loading