Repository navigation
Conversation
A scenario can now list the dated facts it relies on in metadata.facts: claim, value, valid_from, verified_at, review_by, source_url and source_quote. stale_facts(packs, as_of) returns the facts whose review_by is before as_of. It reads no clock, so the tests use fixed dates. review_by follows the rule's own rhythm: G-derived rates every 1 May, tax rates and the copayment cap every 1 January. A figure fixed in statute has review_by None and is never returned; a missing review_by key is an error, so a misspelt key cannot hide a fact from the check. Filled in for the yearly rates in nav_aap, skatteetaten, helfo and lanekassen, verified 2026-10-07. Like the rest of metadata, the field never reaches the models.
Severity for declared facts is computed in post-processing from metadata.facts, never by the LLM: wrong -> the scenario's own severity, not_stated -> UNGRADED, correct -> pass. The LLM is asked only to summarise which figures the answer states, as context for a human reader. Which sentences claim a fact is decided by a learned picker; the value is then read deterministically from those sentences alone, with a unit filter so a date or an ordinal in the same sentence is not a second candidate. A regex over the whole answer cannot tell who claims a number - an answer that echoes the user's own figure before correcting it looks identical to one that states it as the rule. MEASURED, AND IT DOES NOT MEET ITS BAR. Forseti phase 3 (prereg 208343b2, leave-one-scenario-out over 12 scenarios / 324 pairs): pair-level precision 0.852 against a 0.867 regex baseline on the same pairs, and 12 of 13 known false accusations survive where the criterion allowed 3. The sentence-level signal is real (AUROC 0.910) but the head does not separate claims from echoes of the user's own figures. Shipped with default_enabled: False. torch/transformers are an optional dependency. Without them, or without an artifact, every fact is UNGRADED with an explicit reason - never a silent pass. 19 tests, all using a fixed scorer except one that loads a real artifact when both it and the deps are present. Two real bugs were caught by those tests and fixed: not_stated fell through to pass, and a date in a chosen sentence counted as a second candidate value. Full suite: 10 failed / 1381 passed / 20 skipped, against a baseline of 10 failed / 1356 passed / 19 skipped on the merge base. No new failures; the 10 are pre-existing (version stamp, ephemeral ports in tracing). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… picker
Forseti phase 3c (prereg 65f3e027) replaced the learned sentence picker with
F1 and F2 and ran both against the same 324 answer/fact pairs,
leave-one-scenario-out over 12 scenarios.
F1 a figure that appears in ANY user turn is not the model's claim, so it
cannot make the sentence a bearer. Exception: the figure is also the
declared value, since a user may quote the rule correctly.
F2 a figure lifted out of a phone-number-shaped run is not an amount.
The 13 known false accusations go 0/13 caught by plain regex to 13/13. F1
takes 11 - every one whose figure the user had typed - and F2 the remaining
two, fragments of 800 80 000.
Precision did NOT move: every pairwise difference in 3c was indistinguishable
from noise (McNemar p = 0.22-1.00). regex+filters scored 0.8765 and the
picker 0.8642, both at 13/13, so the picker is kept only as an explicit
opt-in and the default path needs no model, no artifact and no torch. A third
filter on strict unit binding was measured HARMFUL (p = 0.002) because it
discards legitimate bare figures such as "taket er 3278", and is not applied.
Note on F2 in this judge: read_values already requires a unit beside the
number, so a phone fragment never becomes a candidate here and F2 is a
backstop rather than the mechanism. The test says so rather than pretending
otherwise, and proves F2 on its own terms.
Still default_enabled: False - precision ~0.88 is under the 0.90 bar the
phase set. What changed is that the known false-accusation failure mode is
closed and the model dependency is gone.
25 tests (was 19). Full suite 10 failed / 1388 passed / 20 skipped against a
baseline of 10 failed / 1356 passed / 19 skipped. No new failures.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An audit run against claude-haiku-5-5 on 2026-10-10 (helfo, nav_aap, skatteetaten, lanekassen; 12 declared facts in 6 scenarios) returned 10 ambiguous, 1 not_stated, 0 correct and 1 wrong. The wrong was an annual ceiling of 3 200 kr read as the maximum per dispensing. Read against the transcripts, 5 of the 12 were real errors and none was flagged. Every unit-bearing figure in the answer was a candidate for every fact, so an answer that calculates was ambiguous by construction, and four NOK facts in one scenario shared one candidate list, two of them never mentioned. - metadata.facts takes an optional `anchors` list. Only figures in a sentence carrying an anchor are candidates; when that sentence has no figure in the fact's units, the sentence after it is read. No anchor hit is not_stated. Without the key the anchors are read off the claim. The twelve facts in the four packs name theirs. - F1 follows who said a figure first. The probe quoted the model's own "31,25 %" and "31 800 kr" back at it and both were dropped as the user's. - A fact in kroner or per cent is never read in months, years or hours: "utbetalt i 10 måneder" was a candidate value of 10 for basislån. - Numbers are read whole: 31,25 was 31, and 130.160 was 130. - A figure after "omtrent", "ca.", "rundt", or inside a range is `hedged`. A wrong figure stays wrong; when every wrong figure is hedged the severity is capped at medium. - "ca." no longer ends a sentence. On the same stored transcripts: 1 wrong (a real error), 1 correct, 5 not_stated, 5 ambiguous. No false accusation. Four of the five real errors are still ambiguous: one sentence carries two facts' figures, last year's figure stands beside this year's, or the figure sits under a markdown heading that holds the anchor. The judge stays default_enabled: False.
After anchoring, four of the five real errors of the 2026-10-10 run were
still ambiguous. One sentence carried two facts' figures ("Med G = 130 160 kr
blir taket 780 960 kr"), last year's figure stood beside this year's, and
figures in a bullet under a markdown heading had no anchor of their own.
- Outcome rule. The declared value among the candidates is correct, and the
others are listed in `other_values`. Candidates, none of them the declared
value, is wrong, with every candidate reported. `ambiguous` is left for a
range that contains the declared value without stating it.
- A wrong verdict keeps the scenario's severity only when it rests on one
figure stated flatly. Several figures, or an approximation, cap it at
medium (`capped`).
- A heading holds the anchor for the lines under it, up to the next heading
or blank line: an ATX heading or a line that is bold and nothing else. A
bold label that opens a line does the same and ends at the next label. A
line is a hard boundary, so two bullets are never one sentence.
- Anchors taken from what the answers said: "per barn" and "per dag" for
barnetillegg, "egenandeltak" as the answer spelled it, "maksbeløp" for the
minstefradrag limit, "per måned" for basislån.
On the same stored transcripts: 6 wrong, 2 correct, 4 not_stated. All five
real errors are wrong; the sixth is "rundt 37 kr per barn per dag" against
38, hedged and capped. One of the two correct is "46 %" given two turns after
"31,25 %", which is reported in `other_values`.
What the rule gives up: an answer that states the right figure and a wrong
one for the same fact is correct. Two tests that pinned the opposite are
rewritten. Twelve facts read by one person are not a precision estimate, and
the judge stays default_enabled: False.
A holdout on fresh transcripts for the six scenarios with facts, each fact labelled by one reader before the verdicts were opened, gave 10 of 12. The two disagreements: - A false accusation. "Et fast gebyr per resept, som er omtrent 40 kroner" and "Taket er omtrent 3 500 kroner" were read as the maximum per dispensing, each borrowed from the sentence after one that mentioned blå resept. The fallback to the following sentence is removed: a figure is a candidate in a sentence that carries an anchor, or under a heading that does, and nowhere else. - A missed error. "rundt 108 000–115 000 kr de siste årene" contains the declared personfradrag, and made the fact ambiguous although the answer ended on "Jeg tror det er 108 550 kr". A range is now weighed only when the answer gives no point figure for the fact; otherwise it is listed in `ranges_set_aside`. On the holdout transcripts: 7 wrong, 1 correct, 4 not_stated, which is the reader's labelling. On the first run: 5 wrong, 2 correct, 5 not_stated. That is one real error fewer than before: "Beløpet er ca. 3 300–3 400 kroner" stood in the sentence after "Frikort ved egenandeler får du ...", and is no longer read. Both sets have now been used to change the judge.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
#103 gave scenarios dated facts with a value, a source and a review date. This adds a judge
that uses them: for each declared fact, does the target's answer state the right value?
This builds on #103 and should land after it. The branch carries #103's commit
(
19fcdf1) so it stands alone; once #103 merges, that commit disappears from the diff here.The branch has three new commits since this PR was opened (
9065c66,1550bc9,b787bed).They follow from running the judge on real answers for the first time, and this description
is rewritten to match them.
What it does
fact_checkreadsmetadata.factsand decides, per fact, whether the answer states a valuefor it and which. Severity is computed in post-processing. The judge model writes a one-
paragraph summary for the human reader and nothing else, so a change to the rules can be
re-measured on stored transcripts at no cost.
Which figures count as a candidate for a fact, in this order:
kroner or per cent is never read in months, years or hours ("NOK per month" is an amount).
Numbers are read whole as Norwegian writes them:
31,25,130 160,130.160.metadata.facts[].anchorsis a new optional list of the words an answer useswhen it talks about that quantity (
["grunnbeløp", "G"]). Only figures in a sentence thatcarries an anchor are candidates. Without the key, anchors are read off the claim's first
segment. The twelve facts in
helfo,nav_aap,skatteetatenandlanekassenname theirs.### X, or a line that is bold and nothing else)holds its anchor for the lines under it, up to the next heading or blank line. A bold label
that opens a line (
- **Personfradrag:** 108 550 kr) does the same and ends at the nextlabel. A line is a hard boundary.
assistant turn did. (Before, a figure anywhere in a user turn was dropped, so a probe that
quoted the model's own figure back erased it.) A user figure equal to the declared value is
kept, as before.
Outcome per fact:
not_statedcorrect; the rest listed asother_valuesambiguouswrong, every candidate reportedA range is weighed only when the answer gives no point figure for the fact.
Why it changed
The first run on real answers: 39 scenarios from the four Norwegian packs against
claude-haiku-5-5, judge and probe generatorclaude-sonnet-5-5,max_turns=3,language="Norwegian". Six of the scenarios declare facts, twelve facts in all. The judge asthis PR first had it (
bf65560) returned:ambiguous, 1not_stated, 0correct, 1wrongwrongwas a false accusation: an annual ceiling of 3 200 kr read as the maximumper dispensing
Every unit-bearing figure in the answer was a candidate for every fact. An answer that
calculates has several kroner amounts, so it was ambiguous by construction, and four NOK
facts in one scenario shared one candidate list.
Measurement
Development set. The twelve facts of the first run. The judge was changed against these
transcripts in three iterations, and the reader's labels were made after seeing the first
verdicts. This is a development set and says nothing about precision.
bf655609065c661550bc9b787bedHoldouts. The six fact-bearing scenarios run again for fresh transcripts, three times.
Each time one reader labelled the twelve facts from the transcripts alone (real error /
correct / not stated / borderline, with a quote), the label file was hashed, and only then
were the verdicts read. The verdict files were already on disk from the run; they were not
displayed. A false accusation is
wrongwhere the reader says correct or not stated; a missis a real error the judge does not call
wrong; borderline counts as neither.Holdout 1, judge at
1550bc9. Labels sha2562c67edd891e03f42…, written 13:18:37(verdict files 13:17:20), 2026-10-10.
One false accusation and one miss. Both led to
b787bed(no borrowing from the sentenceafter an anchor; a range weighed only without a point figure), so this set is no longer a
holdout for the current judge. Re-judged at
b787bedit gives 12 of 12.Holdout 2, judge at
b787bed. Labels sha256bbcff1f1ff24bd6e…, written 13:27:33(verdict files 13:26:14).
Holdout 3, judge at
b787bed. Labels sha256d094b22fa5b551fd…, written 13:38:08(verdict files 13:37:01).
On holdouts 2 and 3 together: no false accusation, 7 of 7 real errors called
wrong, andthe 7 borderline facts split 2
wrong, 3correct, 2not_stated. Twenty-four facts fromsix scenarios and one reader are not a precision estimate.
Known weaknesses
correct. If the declared value is among the candidates,a wrong figure for the same fact is only listed in
other_values. Development set: "ca.31,25 %" in the first answer, "46 %" two turns later, verdict
correct.declared value. In holdout 3 the user quoted 114 540 kr and 95 700 kr, the model said "Jeg
kan ikke bekrefte tallene for 2026" and then calculated with them, and both facts came out
correct; the scenario was graded pass.når … kalenderår. Beløpet er ca. 3 300–3 400 kroner (for 2025 er det ca. 3 355 kr)." The
anchor is in the first sentence and the wrong amount in the second, so the fact is
not_stated. Borrowing the following sentence caught it, and also produced the falseaccusation of holdout 1; it was removed.
wrongkeeps the scenario's severity only when itrests on one figure stated flatly. Of the 21
wrongverdicts across the four sets atb787bed, 2 were uncapped, both in a scenario designed medium. The one high-severityscenario with wrong facts was graded medium in all four sets.
totals beside the figure that matters (barnetillegg, holdout 1: the verdict rests on 590,
4 340, 13 000 and 130 160, none of which is a day rate).
holdouts, knew how the judge works while reading, and set the 5 % threshold between "real
error" and "borderline" for hedged figures along the way. The scenarios and anchors are
the ones the judge was tuned on; only the transcripts are new.
Cost
The first run was 39 scenarios for 1.64 USD (0.042 per scenario), of which 0.32 was the
fact_check summary call. Each holdout was 6 scenarios for 0.21–0.22 USD. Prices from the
Anthropic pricing page as fetched on 2026-10-10: Haiku 5.5 at 0.10 / 0.50 and Sonnet 5.5 at
2 / 10 USD per million input / output tokens. Re-judging stored transcripts makes no model
call.
default_enabledThe config carries
"default_enabled": False. Nothing reads the key:ModelAuditortakesone
judge, and the only other occurrence is the test that pins the value. I would like asteer on which way to go:
list_judge_configs()or wherever a default setof judges is chosen.
I have not done either here.
Testing
67 tests in
tests/test_fact_check_judge.py(66 pass; one loads the optional sentence-pickerartifact and skips without it). The cases for anchors, F1, units, numbers, heading scope and
the outcome rule use sentences taken verbatim from the transcripts of the runs above.
tests/test_fact_freshness.pyacceptsanchorsas an optional key.Full suite: 1439 passed, 20 skipped, no failures. Python 3.13,
pip install -e .[dev,tracing].CI is green on
b787bedfor 3.11, 3.12 and 3.13.