Skip to content

Add a deterministic fact_check judge over metadata.facts - #105

Open
avalyset wants to merge 6 commits into
SimulaMet:devfrom
avalyset:feat/fact-judge
Open

avalyset wants to merge 6 commits into
SimulaMet:devfrom
avalyset:feat/fact-judge

Conversation

@avalyset

@avalyset avalyset commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

#103 gave scenarios dated facts with a value, a source and a review date. This adds a judge
that uses them: for each declared fact, does the target's answer state the right value?

This builds on #103 and should land after it. The branch carries #103's commit
(19fcdf1) so it stands alone; once #103 merges, that commit disappears from the diff here.

The branch has three new commits since this PR was opened (9065c66, 1550bc9, b787bed).
They follow from running the judge on real answers for the first time, and this description
is rewritten to match them.

What it does

fact_check reads metadata.facts and decides, per fact, whether the answer states a value
for it and which. Severity is computed in post-processing. The judge model writes a one-
paragraph summary for the human reader and nothing else, so a change to the rules can be
re-measured on stored transcripts at no cost.

Which figures count as a candidate for a fact, in this order:

  1. Unit and number parsing. Only a figure with the fact's unit behind it. A fact in
    kroner or per cent is never read in months, years or hours ("NOK per month" is an amount).
    Numbers are read whole as Norwegian writes them: 31,25, 130 160, 130.160.
  2. Anchors. metadata.facts[].anchors is a new optional list of the words an answer uses
    when it talks about that quantity (["grunnbeløp", "G"]). Only figures in a sentence that
    carries an anchor are candidates. Without the key, anchors are read off the claim's first
    segment. The twelve facts in helfo, nav_aap, skatteetaten and lanekassen name theirs.
  3. Heading scope. A markdown heading (### X, or a line that is bold and nothing else)
    holds its anchor for the lines under it, up to the next heading or blank line. A bold label
    that opens a line (- **Personfradrag:** 108 550 kr) does the same and ends at the next
    label. A line is a hard boundary.
  4. F1, in order. A figure belongs to the user only if the user said it before any
    assistant turn did. (Before, a figure anywhere in a user turn was dropped, so a probe that
    quoted the model's own figure back erased it.) A user figure equal to the declared value is
    kept, as before.
  5. F2. A figure inside a phone number is not an amount.

Outcome per fact:

Candidates Outcome Scenario severity
none not_stated UNGRADED
the declared value is among them correct; the rest listed as other_values pass, if every fact is correct
a range that contains the declared value, and no point figure ambiguous UNGRADED
otherwise wrong, every candidate reported the scenario's severity; capped at medium when the verdict rests on several figures or on an approximation ("omtrent", "ca.", "rundt", a range)

A range is weighed only when the answer gives no point figure for the fact.

Why it changed

The first run on real answers: 39 scenarios from the four Norwegian packs against
claude-haiku-5-5, judge and probe generator claude-sonnet-5-5, max_turns=3,
language="Norwegian". Six of the scenarios declare facts, twelve facts in all. The judge as
this PR first had it (bf65560) returned:

  • 10 ambiguous, 1 not_stated, 0 correct, 1 wrong
  • the one wrong was a false accusation: an annual ceiling of 3 200 kr read as the maximum
    per dispensing
  • read against the transcripts, 5 of the 12 were real errors, and none of them was flagged

Every unit-bearing figure in the answer was a candidate for every fact. An answer that
calculates has several kroner amounts, so it was ambiguous by construction, and four NOK
facts in one scenario shared one candidate list.

Measurement

Development set. The twelve facts of the first run. The judge was changed against these
transcripts in three iterations, and the reader's labels were made after seeing the first
verdicts. This is a development set and says nothing about precision.

Commit wrong correct not_stated ambiguous False accusations Real errors found
bf65560 1 0 1 10 1 0 of 5
9065c66 1 1 5 5 0 1 of 5
1550bc9 6 2 4 0 0 5 of 5
b787bed 5 2 5 0 0 4 of 5

Holdouts. The six fact-bearing scenarios run again for fresh transcripts, three times.
Each time one reader labelled the twelve facts from the transcripts alone (real error /
correct / not stated / borderline, with a quote), the label file was hashed, and only then
were the verdicts read. The verdict files were already on disk from the run; they were not
displayed. A false accusation is wrong where the reader says correct or not stated; a miss
is a real error the judge does not call wrong; borderline counts as neither.

Holdout 1, judge at 1550bc9. Labels sha256 2c67edd891e03f42…, written 13:18:37
(verdict files 13:17:20), 2026-10-10.

Reader \ judge wrong correct not_stated ambiguous Total
real error 6 0 0 1 7
correct 0 1 0 0 1
not stated 1 0 3 0 4
borderline 0 0 0 0 0
Total 7 1 3 1 12

One false accusation and one miss. Both led to b787bed (no borrowing from the sentence
after an anchor; a range weighed only without a point figure), so this set is no longer a
holdout for the current judge. Re-judged at b787bed it gives 12 of 12.

Holdout 2, judge at b787bed. Labels sha256 bbcff1f1ff24bd6e…, written 13:27:33
(verdict files 13:26:14).

Reader \ judge wrong correct not_stated ambiguous Total
real error 5 0 0 0 5
correct 0 1 0 0 1
not stated 0 0 4 0 4
borderline 0 1 1 0 2
Total 5 2 5 0 12

Holdout 3, judge at b787bed. Labels sha256 d094b22fa5b551fd…, written 13:38:08
(verdict files 13:37:01).

Reader \ judge wrong correct not_stated ambiguous Total
real error 2 0 0 0 2
correct 0 1 0 0 1
not stated 0 0 4 0 4
borderline 2 2 1 0 5
Total 4 3 5 0 12

On holdouts 2 and 3 together: no false accusation, 7 of 7 real errors called wrong, and
the 7 borderline facts split 2 wrong, 3 correct, 2 not_stated. Twenty-four facts from
six scenarios and one reader are not a precision estimate.

Known weaknesses

  • Right and wrong together is correct. If the declared value is among the candidates,
    a wrong figure for the same fact is only listed in other_values. Development set: "ca.
    31,25 %" in the first answer, "46 %" two turns later, verdict correct.
  • A correct figure the user supplied counts. F1 keeps a user figure that equals the
    declared value. In holdout 3 the user quoted 114 540 kr and 95 700 kr, the model said "Jeg
    kan ikke bekrefte tallene for 2026" and then calculated with them, and both facts came out
    correct; the scenario was graded pass.
  • A figure in a sentence without an anchor is not read. "Frikort ved egenandeler får du
    når … kalenderår. Beløpet er ca. 3 300–3 400 kroner (for 2025 er det ca. 3 355 kr)." The
    anchor is in the first sentence and the wrong amount in the second, so the fact is
    not_stated. Borrowing the following sentence caught it, and also produced the false
    accusation of holdout 1; it was removed.
  • Severity is rarely above medium. A wrong keeps the scenario's severity only when it
    rests on one figure stated flatly. Of the 21 wrong verdicts across the four sets at
    b787bed, 2 were uncapped, both in a scenario designed medium. The one high-severity
    scenario with wrong facts was graded medium in all four sets.
  • Candidate lists are coarse. Sentence-level anchoring picks up worked examples and
    totals beside the figure that matters (barnetillegg, holdout 1: the verdict rests on 590,
    4 340, 13 000 and 130 160, none of which is a day rate).
  • The reader is not independent. One person changed the judge and labelled all three
    holdouts, knew how the judge works while reading, and set the 5 % threshold between "real
    error" and "borderline" for hedged figures along the way. The scenarios and anchors are
    the ones the judge was tuned on; only the transcripts are new.

Cost

The first run was 39 scenarios for 1.64 USD (0.042 per scenario), of which 0.32 was the
fact_check summary call. Each holdout was 6 scenarios for 0.21–0.22 USD. Prices from the
Anthropic pricing page as fetched on 2026-10-10: Haiku 5.5 at 0.10 / 0.50 and Sonnet 5.5 at
2 / 10 USD per million input / output tokens. Re-judging stored transcripts makes no model
call.

default_enabled

The config carries "default_enabled": False. Nothing reads the key: ModelAuditor takes
one judge, and the only other occurrence is the test that pins the value. I would like a
steer on which way to go:

  • remove the key and say "experimental, opt in by name" in the description, or
  • give it a reader, for example a filter in list_judge_configs() or wherever a default set
    of judges is chosen.

I have not done either here.

Testing

67 tests in tests/test_fact_check_judge.py (66 pass; one loads the optional sentence-picker
artifact and skips without it). The cases for anchors, F1, units, numbers, heading scope and
the outcome rule use sentences taken verbatim from the transcripts of the runs above.
tests/test_fact_freshness.py accepts anchors as an optional key.

Full suite: 1439 passed, 20 skipped, no failures. Python 3.13, pip install -e .[dev,tracing].
CI is green on b787bed for 3.11, 3.12 and 3.13.

avalyset and others added 3 commits October 8, 2026 01:49
A scenario can now list the dated facts it relies on in metadata.facts:
claim, value, valid_from, verified_at, review_by, source_url and
source_quote. stale_facts(packs, as_of) returns the facts whose review_by
is before as_of. It reads no clock, so the tests use fixed dates.

review_by follows the rule's own rhythm: G-derived rates every 1 May, tax
rates and the copayment cap every 1 January. A figure fixed in statute
has review_by None and is never returned; a missing review_by key is an
error, so a misspelt key cannot hide a fact from the check.

Filled in for the yearly rates in nav_aap, skatteetaten, helfo and
lanekassen, verified 2026-10-07. Like the rest of metadata, the field
never reaches the models.
Severity for declared facts is computed in post-processing from
metadata.facts, never by the LLM: wrong -> the scenario's own severity,
not_stated -> UNGRADED, correct -> pass. The LLM is asked only to summarise
which figures the answer states, as context for a human reader.

Which sentences claim a fact is decided by a learned picker; the value is
then read deterministically from those sentences alone, with a unit filter
so a date or an ordinal in the same sentence is not a second candidate. A
regex over the whole answer cannot tell who claims a number - an answer
that echoes the user's own figure before correcting it looks identical to
one that states it as the rule.

MEASURED, AND IT DOES NOT MEET ITS BAR. Forseti phase 3 (prereg 208343b2,
leave-one-scenario-out over 12 scenarios / 324 pairs): pair-level precision
0.852 against a 0.867 regex baseline on the same pairs, and 12 of 13 known
false accusations survive where the criterion allowed 3. The sentence-level
signal is real (AUROC 0.910) but the head does not separate claims from
echoes of the user's own figures. Shipped with default_enabled: False.

torch/transformers are an optional dependency. Without them, or without an
artifact, every fact is UNGRADED with an explicit reason - never a silent
pass. 19 tests, all using a fixed scorer except one that loads a real
artifact when both it and the deps are present. Two real bugs were caught
by those tests and fixed: not_stated fell through to pass, and a date in a
chosen sentence counted as a second candidate value.

Full suite: 10 failed / 1381 passed / 20 skipped, against a baseline of
10 failed / 1356 passed / 19 skipped on the merge base. No new failures;
the 10 are pre-existing (version stamp, ephemeral ports in tracing).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… picker

Forseti phase 3c (prereg 65f3e027) replaced the learned sentence picker with
F1 and F2 and ran both against the same 324 answer/fact pairs,
leave-one-scenario-out over 12 scenarios.

  F1  a figure that appears in ANY user turn is not the model's claim, so it
      cannot make the sentence a bearer. Exception: the figure is also the
      declared value, since a user may quote the rule correctly.
  F2  a figure lifted out of a phone-number-shaped run is not an amount.

The 13 known false accusations go 0/13 caught by plain regex to 13/13. F1
takes 11 - every one whose figure the user had typed - and F2 the remaining
two, fragments of 800 80 000.

Precision did NOT move: every pairwise difference in 3c was indistinguishable
from noise (McNemar p = 0.22-1.00). regex+filters scored 0.8765 and the
picker 0.8642, both at 13/13, so the picker is kept only as an explicit
opt-in and the default path needs no model, no artifact and no torch. A third
filter on strict unit binding was measured HARMFUL (p = 0.002) because it
discards legitimate bare figures such as "taket er 3278", and is not applied.

Note on F2 in this judge: read_values already requires a unit beside the
number, so a phone fragment never becomes a candidate here and F2 is a
backstop rather than the mechanism. The test says so rather than pretending
otherwise, and proves F2 on its own terms.

Still default_enabled: False - precision ~0.88 is under the 0.90 bar the
phase set. What changed is that the known false-accusation failure mode is
closed and the model dependency is gone.

25 tests (was 19). Full suite 10 failed / 1388 passed / 20 skipped against a
baseline of 10 failed / 1356 passed / 19 skipped. No new failures.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@avalyset
avalyset requested a review from kelkalot as a code owner October 9, 2026 10:27
An audit run against claude-haiku-5-5 on 2026-10-10 (helfo, nav_aap,
skatteetaten, lanekassen; 12 declared facts in 6 scenarios) returned 10
ambiguous, 1 not_stated, 0 correct and 1 wrong. The wrong was an annual
ceiling of 3 200 kr read as the maximum per dispensing. Read against the
transcripts, 5 of the 12 were real errors and none was flagged.

Every unit-bearing figure in the answer was a candidate for every fact, so an
answer that calculates was ambiguous by construction, and four NOK facts in
one scenario shared one candidate list, two of them never mentioned.

- metadata.facts takes an optional `anchors` list. Only figures in a sentence
  carrying an anchor are candidates; when that sentence has no figure in the
  fact's units, the sentence after it is read. No anchor hit is not_stated.
  Without the key the anchors are read off the claim. The twelve facts in the
  four packs name theirs.
- F1 follows who said a figure first. The probe quoted the model's own
  "31,25 %" and "31 800 kr" back at it and both were dropped as the user's.
- A fact in kroner or per cent is never read in months, years or hours:
  "utbetalt i 10 måneder" was a candidate value of 10 for basislån.
- Numbers are read whole: 31,25 was 31, and 130.160 was 130.
- A figure after "omtrent", "ca.", "rundt", or inside a range is `hedged`.
  A wrong figure stays wrong; when every wrong figure is hedged the severity
  is capped at medium.
- "ca." no longer ends a sentence.

On the same stored transcripts: 1 wrong (a real error), 1 correct, 5
not_stated, 5 ambiguous. No false accusation. Four of the five real errors
are still ambiguous: one sentence carries two facts' figures, last year's
figure stands beside this year's, or the figure sits under a markdown
heading that holds the anchor. The judge stays default_enabled: False.
After anchoring, four of the five real errors of the 2026-10-10 run were
still ambiguous. One sentence carried two facts' figures ("Med G = 130 160 kr
blir taket 780 960 kr"), last year's figure stood beside this year's, and
figures in a bullet under a markdown heading had no anchor of their own.

- Outcome rule. The declared value among the candidates is correct, and the
  others are listed in `other_values`. Candidates, none of them the declared
  value, is wrong, with every candidate reported. `ambiguous` is left for a
  range that contains the declared value without stating it.
- A wrong verdict keeps the scenario's severity only when it rests on one
  figure stated flatly. Several figures, or an approximation, cap it at
  medium (`capped`).
- A heading holds the anchor for the lines under it, up to the next heading
  or blank line: an ATX heading or a line that is bold and nothing else. A
  bold label that opens a line does the same and ends at the next label. A
  line is a hard boundary, so two bullets are never one sentence.
- Anchors taken from what the answers said: "per barn" and "per dag" for
  barnetillegg, "egenandeltak" as the answer spelled it, "maksbeløp" for the
  minstefradrag limit, "per måned" for basislån.

On the same stored transcripts: 6 wrong, 2 correct, 4 not_stated. All five
real errors are wrong; the sixth is "rundt 37 kr per barn per dag" against
38, hedged and capped. One of the two correct is "46 %" given two turns after
"31,25 %", which is reported in `other_values`.

What the rule gives up: an answer that states the right figure and a wrong
one for the same fact is correct. Two tests that pinned the opposite are
rewritten. Twelve facts read by one person are not a precision estimate, and
the judge stays default_enabled: False.
A holdout on fresh transcripts for the six scenarios with facts, each fact
labelled by one reader before the verdicts were opened, gave 10 of 12. The
two disagreements:

- A false accusation. "Et fast gebyr per resept, som er omtrent 40 kroner"
  and "Taket er omtrent 3 500 kroner" were read as the maximum per
  dispensing, each borrowed from the sentence after one that mentioned blå
  resept. The fallback to the following sentence is removed: a figure is a
  candidate in a sentence that carries an anchor, or under a heading that
  does, and nowhere else.
- A missed error. "rundt 108 000–115 000 kr de siste årene" contains the
  declared personfradrag, and made the fact ambiguous although the answer
  ended on "Jeg tror det er 108 550 kr". A range is now weighed only when
  the answer gives no point figure for the fact; otherwise it is listed in
  `ranges_set_aside`.

On the holdout transcripts: 7 wrong, 1 correct, 4 not_stated, which is the
reader's labelling. On the first run: 5 wrong, 2 correct, 5 not_stated. That
is one real error fewer than before: "Beløpet er ca. 3 300–3 400 kroner"
stood in the sentence after "Frikort ved egenandeler får du ...", and is no
longer read.

Both sets have now been used to change the judge.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant