A database of findings about machine learning models: third-party claims about how a specific
model behaves, made after the fact. The YAML files under data/ are the only source of truth;
the graph, the site and the CSV exports are all derived from them and can be deleted and rebuilt
at any time.
This README is a manual. What every command does, in the order you would normally run them.
python3 -m venv .venv && .venv/bin/pip install -r requirements.txtsource .venv/bin/activateActivate once per shell and every command below works as written. PyYAML is all the build and
the site need. pypdfium2 is used to read PDFs and openreview-py only by harvest.py, which is
the one script that touches the network.
You will run these most often. None of them touches the network.
| Command | What it does |
|---|---|
python3 run_tests.py |
Runs all 308 tests across five suites. No arguments. |
python3 build.py |
Validates data/, writes out/graph.json, prints an audit. Exits non-zero and writes nothing if validation fails, so run it before every commit. |
python3 render.py |
Builds the static site into site/ from out/graph.json. |
python3 export.py |
Writes one CSV per node type plus edges.csv into out/csv/. |
The usual loop after editing any YAML file:
python3 build.py && python3 render.pyThen open site/index.html by double-clicking it. Links are relative, so no server is needed and
the whole folder can be zipped and sent as one thing.
build.py prints more than a pass or a fail. It lists how many records came from manual versus
automatic extraction, how many findings and how many models each concept reaches, which entities
more than one finding reaches, which findings have no dataset link, which no concept covers,
which registry entries lack an anchor, and which registry entries nothing reaches at all.
| Command | What it does |
|---|---|
python3 check.py path/to/candidate.yaml |
Schema errors and link resolution for a candidate finding that is not in data/ yet. Blames only the candidate, never the existing registries. |
python3 verify.py data/findings/ID.yaml source.pdf |
Locates the record's numbers, linked entity names and stable identifiers in the PDF, and finds the page sharing most of the caveat's vocabulary. |
verify.py changes nothing on disk. Missing items are suspicions for a human, not corrections.
It exits 1 when a check is blocking or when the record offered nothing to check at all — an
empty finding must not report 0 blocking and look like a pass.
harvest.py is the only script that talks to the network. Run it with .venv/bin/python, because
openreview-py lives there. OPENREVIEW_USERNAME and OPENREVIEW_PASSWORD must be set in the
environment — never in the repository.
| Command | What it does |
|---|---|
harvest.py doctor |
Offline. Checks the interpreter, the installed packages and that the API client still has the methods we call. No account needed. |
harvest.py preflight [venue_id] |
Everything doctor does, plus a real login. With a venue id it also checks the group, the submission invitation and a sample paper's fields. Run this once before harvesting a new conference. |
harvest.py venues [substring] |
Lists venue identifiers, e.g. harvest.py venues ICML. |
harvest.py meta <venue_id> [--all] |
Fetches metadata only — no PDFs — screens each paper and appends a row per paper to corpus/manifest.jsonl. Resumable: papers already in the manifest are skipped. --all includes rejected submissions. |
harvest.py stats |
Tier breakdown of the manifest, how many PDFs and texts are on disk, and which screening rules produced each row. |
harvest.py pdfs [--tier a,b] [--limit N] [--ids FILE] |
Downloads PDFs for the chosen tiers into corpus/pdf/. Defaults to strong,possible. --ids takes a file with one identifier per line and overrides --tier. |
harvest.py text |
Extracts text from every PDF into corpus/text/, skipping files already done. |
Screening never rejects a paper, it only sorts it into strong, possible or weak. Downloading
is a separate step so a bad screening rule costs nothing but a rerun.
extract.py drives the extraction pipeline. Everything here reads corpus/text/, never the PDFs.
| Command | What it does |
|---|---|
extract.py prompts [paper,paper] |
Builds one extraction prompt per paper from corpus/text/ into corpus/prompts/. With no argument, every paper. |
extract.py collect <directory> |
Reads model answers from a directory, repairs common YAML damage, matches each answer to its paper by content and saves it into corpus/answers/. Add file.txt=<paper> to assign an answer that cannot be matched. |
extract.py verify |
Checks every citation the model wrote against the text of its own paper and writes corpus/entities.jsonl. Exits 1 if any citation is rejected. |
extract.py propose [N] |
Lists entities the answers name that no registry holds, reaching N papers or more. Also reports which concepts the model refused, proposed or silently skipped. Writes corpus/proposed.jsonl. |
extract.py tags [all] |
Writes one small tagging prompt per finding that carries no concept. all re-tags every finding instead. |
extract.py split [--write] [--force] |
Turns collected answers into records under data/findings/. Reports only by default. --write creates files but never overwrites; --force overwrites. |
A candidate from the entity linker is never accepted automatically. It is a suggestion for a human: on four real suggestions, two were wrong.
build.py YAML -> validate -> assemble -> out/graph.json, then the audit
render.py out/graph.json -> site/
export.py out/graph.json -> out/csv/*.csv
check.py a candidate finding -> schema errors plus link resolution
verify.py a finding + its PDF -> evidence locations and suspicions
harvest.py OpenReview -> corpus/; the only script that uses the network
extract.py corpus/ -> prompts, answers, proposals, data/findings/
run_tests.py runs all five suites
modelpedia/ the library; imported, never run
graph.py node and edge types, the NODE_TYPES table
schema.py the finding schema: link fields, vocabularies, regexes
paths.py every filesystem location; the only file that derives the root
graph_io.py load/dump out/graph.json with the format_version guard
atomic.py write-then-rename; the single .part convention
record_keys.py string constants for keys inside records
console.py console output primitives
build/ data/*.yaml -> out/graph.json
database.py the only YAML reader
validate.py Database -> error strings; creates nothing
assemble.py Database -> the graph dict; validates nothing
report.py the audit build.py prints
site/ out/graph.json -> HTML
ingest/ papers -> candidate findings
text.py PDF -> normalised searchable text
link.py entity name -> hit / candidates / miss
screen.py title and abstract -> score and tier
manifest.py the corpus manifest: validation, reading, selection
openreview.py everything that knows the OpenReview API
report.py console reports for extract.py
verification.py the evidence checks verify.py runs
data/ vocabularies.yaml, registries/*.yaml, findings/*.yaml
corpus/ harvested papers and model answers; not tracked
out/, site/ build artifacts; not tracked, delete and rebuild freely
Never search a PDF except through modelpedia/ingest/text.py. Extracted text breaks words
across lines, splits small-capital headings and mangles ligatures. A plain grep misses them and
reports absence that is not real. This has already cost the project one wrongly deleted citation.
Identity lives on the node, role lives on the edge. dataset:terramesh is the same entity
whether one paper trained on it and another evaluated on it; [train] or [eval] belongs on the
link, never in the registry.
Gaps are stated, not guessed. Where a source names no dataset or prints no URL, the field
stays empty and build.py lists it in the audit. A visible gap beats false precision.
Every artifact is replaced in one step or not at all. Each writer stages its output beside the target and renames it into place, so a crash or a full disk leaves the last good output untouched.
62 findings across 5 registries; 388 nodes and 755 edges. Nine were written by hand
(extracted_by: manual-extraction); the other 53 came from ICLR 2025 through automatic extraction
and have not been read against their sources.
extracted_by is the only record-level field and it states origin, nothing more. An earlier
review_status field was removed because reading the sources found errors in 5 of the 7 records
that carried verified — the label recorded that someone had checked, not that the check was
good. A record's presence in this database is not evidence that anyone verified it against its
source.