A loan-file copilot for home-loan (mortgage) officers: upload a borrower's packet, get an automatic review of the application against its supporting documents, record a decision on every flag, and ask anything else with page-cited answers. Under the hood it is a multi-user RAG (Retrieval-Augmented Generation) application for PDF documents. Built with a React frontend and a FastAPI backend, featuring accounts, persistent per-document conversations, open-source document extraction and hybrid sparse-dense search.
One account → many chats → exactly one document each. A chat session's id is its Pinecone namespace, so chat ↔ document ↔ namespace is 1:1. Neon/Postgres owns identity, ownership and conversation history; Pinecone holds only vectors.
- Accounts & persistent chats: JWT auth over bcrypt, DB-backed login lockout, and conversations that survive a restart — including their source citations. A login lasts as long as the tab is open: tokens live in
sessionStorage, so a reload keeps you signed in, closing the tab signs you out, and a new tab asks you to log in. - Session rehydration: the in-process retriever is a pure cache. On a miss it rebuilds from Neon + Pinecone in ~11s instead of re-ingesting the document (~180s), by persisting the fitted BM25 encoder and recomputing centroids from the index.
- Asynchronous ingestion: upload returns
202and processes in the background with a pollable status, because extraction runs for minutes on a real packet. - Pluggable Extraction: AWS Textract (TABLES + FORMS, default) or PyMuPDF (local, no-AI, text-layer only) via
EXTRACT_METHOD. Contextual chunking attaches each chunk's document identity so entity-specific queries stay unambiguous. - Hybrid Search Engine: A single Pinecone sparse-dense index holding
gemini-embedding-2embeddings (768-dim) alongside BM25 sparse vectors, fused by a tunablealpha(0.0 = pure keyword → 1.0 = pure semantic). - Conversational follow-ups: a follow-up like "and when does it lock?" is condensed into a standalone question before retrieval, because retrieval runs before any LLM sees a prompt. History is read server-side from Neon.
- Answer Generation: gemini-2.5-flash with thinking capped at 2048 — on an ambiguous multi-candidate question it enumerates candidates with sources instead of guessing. Thinking tokens bill at the output rate, so the cap bounds the tail (dynamic permits 24,576) without touching the ~300-token median.
- Two models, split on measured need: classification and per-page boundary detection run on gemini-2.5-flash-lite (closed-set label, yes/no answer — and boundary detection fires once per page, making it the volume driver of ingest cost). Answers and query rewriting stay on flash.
- Automatic file review: once a packet is ingested, the application's claims (income, employer, balances, declarations, loan terms) are checked against the pay slips, bank statement, Loan Estimate and title report. Each check shows both sides' values with their pages, in a panel beside the chat (a drawer on narrow screens) so the officer can ask about a flag without leaving it. See Automatic file review.
- Bring your own Gemini key: ⚙ Settings (a popup holding the key and the search parameters) lets a user add their own Gemini API key. It is checked with Gemini first, kept only in that browser tab, sent with each question and never stored on the server. With a key set, the daily credit limit is skipped, those questions don't use up credits, and the header shows "Your API key" instead of the credit count. The key pays for the question rewrite and the answer; search and uploads stay on the app's keys.
- Flag decisions with an audit trail: every mismatch or review flag can be accepted (a note is required) or confirmed as an issue. Decisions are append-only, record who decided and when, and the sidebar counts open flags per borrower.
- Several PDFs per borrower: upload up to 20 files at once; they are merged, in order, into the file's one document.
- Grounded answers: values are quoted exactly with their document and page, never computed or rounded; a question the file cannot answer gets "not in the documents", not a guess. Summaries use their own rules.
- Semantic Routing: Automatic query routing to specific document sections via embedding centroids — no extra LLM call.
Note on reranking: a cross-encoder reranking stage (BAAI/bge-reranker-base) was built and evaluated on 250 questions, then removed — it changed answer quality by a statistically indistinguishable amount while costing a 3× over-fetch and a GPU round-trip per query. It remains a reasonable optional addition under conditions this corpus does not meet. See Design FAQ Q2 for the measurements.
A loan officer's core check is stare and compare: the application holds the borrower's claims, the supporting documents are the evidence. The review does that comparison for every file, and it runs inside the same background task as the ingest, right after the document turns ready.
- Classify and split the packet into documents (16 types, including Loan Application).
- Extract a fixed set of fields per document type: first from Textract's FORMS label/value pairs, then one flash-lite call per document for anything missing and for lists (deposits, debits, liabilities, liens, easements).
- Grounding guard: a model-extracted value is kept only if its exact text is found in that document, and its page is located in code, never taken from the model.
- Compare with rules in plain Python (thresholds at the top of
core/review.py), not with the LLM.
| Check | Compares |
|---|---|
| Income | stated monthly income vs pay slips (pay frequency from the slip or its period dates), 5% tolerance |
| Employer | application vs pay slip |
| Checking balance | stated vs statement ending balance, 10% tolerance |
| Payroll deposits | statement deposits vs pay-slip net pay |
| Large deposits | non-payroll deposits above 25% of monthly income vs the declaration |
| Loan amount, purchase price | application vs Loan Estimate |
| Stated debt payments | stated monthly payments vs statement debits |
| Liens, easements | title report; "payoff required" read from the report's own wording |
| Missing documents | required document types absent from the file |
Each check is mismatch, review, missing, match or info. A review that fails is stored as failed and never fails the upload: the document stays searchable.
Try it: samples/synthetic_borrower/whitfield_loan_packet.pdf is an 8-page synthetic packet (or the same pages as 5 separate PDFs) with planted issues. answer_key.md lists the expected review and 35 test questions. The packet is regenerated with python samples/synthetic_borrower/make_samples.py.
- Page numbers are 1-based everywhere a person or the model sees them. Pages are stored 0-based; citations, the answer prompt and the eval prompt add 1. Before this, answers cited one page early. Pages are pages of the merged file when several PDFs were uploaded.
- The grounding guard ignores table cell separators, so a value quoted from a table row (
07/15 | DIRECT DEP … | +$2,845.31) still matches, and list values the model wraps as{"value": …}are unwrapped rather than dropped. - With PyMuPDF, page-boundary detection reads table rows without their
|separators: with them, the model called a pay slip and the next bank statement "the same document" every time. - Provider calls retry 429s, 5xx and timeouts up to 3 times with full-jitter backoff; 4xx errors are never retried. Without this a rate-limited minute during page splitting silently kept pages together and merged documents.
Six layers, read top to bottom. Each arrow is a hand-off between layers; the ingest and retrieval pipelines flow left-to-right within their own band and the bands stack one below the other, while the shared services (data, external AI) are reached once per layer rather than by every stage, so the flow stays legible. Observability is cross-cutting.
graph TD
User(["👤 User"])
subgraph CLIENT ["1 · Client layer — React / Vite"]
Land["🛬 Landing page"]
UI["💬 Chat UI · upload gate · live ingest stepper · streamed answers"]
end
subgraph APP ["2 · Application layer — FastAPI (main.py)"]
Auth["🔐 Auth · bcrypt · access + refresh JWT · lockout · per-IP rate-limit"]
REST["🗂️ Chat endpoints · POST /message → SSE · 202 async upload + polling"]
end
subgraph INGEST ["3 · Ingest pipeline — background task"]
direction LR
Ext["📄 Extract · Textract / PyMuPDF"] --> Split["🏷️ Classify + split · flash-lite"] --> Chunk["✂️ Chunk · tables atomic · contextual"] --> Emb["🧬 Embed · 768d"] --> Up["📤 BM25 fit + Pinecone upsert"]
end
subgraph DATA ["4 · Data layer — Neon Postgres + Pinecone"]
PG[("🐘 Postgres · accounts · chats · messages · bm25_params")]
Pine[("🌲 Pinecone · one namespace per user")]
end
subgraph QUERY ["5 · Retrieval + answer layer"]
direction LR
RW["📝 Rewrite follow-up → standalone"] --> Hyb["🔍 Hybrid query · α·dense + (1−α)·sparse"] --> Ans["🤖 gemini-2.5-flash · streamed, cited"]
end
subgraph EXT ["6 · External AI services"]
Gem["☁️ Google Gemini · flash / flash-lite / embeddings"]
Tex["☁️ AWS Textract · TABLES + FORMS"]
end
OBS["📈 Observability · cross-cutting<br>Langfuse (LLM) + Grafana (HTTP · metrics · dashboard)"]
User --> CLIENT
CLIENT -->|HTTP + JWT| APP
APP -->|identity · ownership| DATA
APP -->|upload| INGEST
INGEST -->|vectors + bm25_params| DATA
DATA -->|hybrid search| QUERY
INGEST -->|extract · classify · embed| EXT
QUERY -->|rewrite · answer| EXT
APP -.->|HTTP traces · metrics| OBS
INGEST -.->|LLM traces| OBS
QUERY -.->|LLM traces| OBS
- Node.js: For the React frontend.
- Python 3.10+: For the FastAPI backend.
- Neon (or any Postgres): Accounts, chat sessions and messages.
- Pinecone: Serverless index,
dimension=768,metric=dotproduct. - Google AI API key:
gemini-embedding-2embeddings andgemini-2.5-flashanswers. - AWS account (default extraction path): AWS Textract API (
AmazonTextractFullAccess— or scoped totextract:AnalyzeDocument). Skip if you setEXTRACT_METHOD=pymupdf(local, no AWS calls).
- Navigate to the backend directory:
cd backend - Install dependencies:
pip install -r requirements.txt
PDF extraction runs on AWS Textract (TABLES + FORMS features), billed at $0.065 per page — about $3.25 for a 50-page packet. FORMS returns label/value pairs, which the file review uses before asking an LLM.
EXTRACT_METHOD=pymupdf reads the PDF's text layer locally, for free. It detects tables with PyMuPDF's find_tables() so they stay whole in chunking, but it cannot read scanned pages and returns no FORMS pairs. A PDF with no readable text on any page is rejected with a clear message; pages without text are listed in the file header ("No readable text on page 9. Not searched or reviewed.") instead of being indexed blank. Answers come from the Gemini API directly, so no self-hosted LLM server is required.
- Create an IAM user with
AmazonTextractFullAccess(or a scoped policy grantingtextract:AnalyzeDocument). - Generate an access key pair for that user.
- Set the credentials in
backend/.env(see section E).
Create a serverless index with dimension=768 and metric=dotproduct — dotproduct is required for sparse-dense hybrid queries, and the dimension must equal EMBED_DIM in llm/llm_router.py, which also drives the embedding call itself. The backend verifies both at startup and refuses to run on a mismatch, rather than failing minutes later at upsert.
Create a database and copy its pooled connection string (the -pooler host — PgBouncer multiplexes many client connections onto few backends, which a free-tier compute needs). Tables are prefixed drs_ so this schema can share a database with other projects.
Create them with either:
# from backend/
python -c "from db.database import engine, Base; import db.models; Base.metadata.create_all(engine)"
# ...or paste migrations.sql into the Neon SQL editor# Vector store
PINECONE_API_KEY=your_key
PINECONE_INDEX_NAME=your_index
PINECONE_HOST=https://your-index-xxxxx.svc.region.pinecone.io
# Text generation provider — GEMINI or CLOUDFLARE. Required, no default.
LLM_MODEL=GEMINI
# Embeddings (always Gemini) + generation when LLM_MODEL=GEMINI
GEMINI_API_KEY=your_gemini_key
# Cloudflare Workers AI — only required when LLM_MODEL=CLOUDFLARE
CLOUDFLARE_ACCOUNT_ID=your_account_id
CLOUDFLARE_API_TOKEN=your_workers_ai_token
# Extraction — pluggable. Default textract; set to "pymupdf" for local text-layer read (no AWS calls).
EXTRACT_METHOD=textract
# AWS Textract (only required when EXTRACT_METHOD=textract)
AWS_ACCESS_KEY_ID=your_aws_access_key
AWS_SECRET_ACCESS_KEY=your_aws_secret
AWS_REGION=us-east-1
# Database + auth
DATABASE_URL=postgresql://user:pass@ep-xxx-pooler.region.aws.neon.tech/dbname?sslmode=require
JWT_SECRET=<python -c "import secrets; print(secrets.token_urlsafe(48))">
# Optional
ALLOWED_ORIGINS=https://your-frontend.onrender.com,http://localhost:5173 # CORS allow-list (comma-separated)
MAX_UPLOAD_MB=3 # upload cap (MB), enforced while streaming
DAILY_MESSAGE_CAP=5 # REQUIRED — messages per account per day (1 credit = question + answer)
GEMINI_THINKING_BUDGET=2048 # fixed ceiling (default); 0 = off, -1 = dynamic
GEMINI_FAST_MODEL=gemini-2.5-flash-lite # classification + boundary detection
CONTEXTUAL_CHUNKING=1 # attach per-document identity to each chunk (default 1, set 0 to disable)
TOKEN_TTL_HOURS=24
# Observability (optional — all tracing stays off unless these are set)
LANGFUSE_PUBLIC_KEY=pk-lf-...
LANGFUSE_SECRET_KEY=sk-lf-...
LANGFUSE_HOST=https://us.cloud.langfuse.com
GRAFANA_OTLP_ENDPOINT=https://otlp-gateway-prod-<region>.grafana.net/otlp
GRAFANA_OTLP_AUTH=Basic <base64> # the full Authorization header value
OTEL_SERVICE_NAME=document-retrieval-system
- Navigate to the frontend directory:
cd frontend - Install dependencies:
npm install
- Run Dev Server:
npm run dev
- Open
http://localhost:5173, create an account, then start a chat and upload a PDF.
Point a build at a deployed backend with VITE_API_URL:
VITE_API_URL=https://your-backend.onrender.com npm run buildEvery endpoint except signup and login requires Authorization: Bearer <token>.
| Method | Path | Notes |
|---|---|---|
POST |
/api/auth/signup |
→ { token, user_id, username } |
POST |
/api/auth/login |
5 failures / 15 min locks the username |
GET |
/api/chats |
Sidebar list, newest first |
POST |
/api/chats/new |
Reuses an existing empty chat |
GET |
/api/chats/{id} |
Chat + full message history, the file review, the latest decision per flag and the decision history |
POST |
/api/chats/{id}/document |
202 — ingests in the background. One PDF or up to 20 under the same file field, merged in order |
GET |
/api/chats/{id}/status |
Poll while processing |
POST |
/api/chats/{id}/message |
Ask a question. Returns question_asked and question_searched so a rewritten follow-up is diagnosable |
PATCH |
/api/chats/{id} |
Rename |
DELETE |
/api/chats/{id} |
Drops the namespace and the rows. Flag decisions are kept (audit trail) |
GET |
/api/account/credits |
{cap, used, remaining} for today; questions asked with the user's own key are not counted |
POST |
/api/account/check-key |
{key} — validates a Gemini API key without storing it (reads model metadata, bills nothing). 400 if Gemini rejects it |
POST |
/api/chats/{id}/review/decisions |
{check_key, decision: accepted|confirmed, note}. Accept requires a note. 409 if the file has no finished review, 422 for an unknown flag |
POST /message also accepts an optional X-Gemini-Key header: the user's own key for that one question — used for the rewrite and the answer, skips the daily cap, never stored or logged.
Chat lifecycle: awaiting_document → processing → ready | failed
A chat that belongs to another user returns 404, not 403 — a 403 would confirm the id exists.
The system's performance is validated using the Ragas evaluation framework, focusing on faithfulness, relevancy, and retrieval quality.
Current pipeline — k=6, 250 questions (AWS Textract extraction + contextual chunking; generator gemini-2.5-flash-lite, judge gemini-3.5-flash-lite)
All three retrieval modes over the same corpus and models — raw per-question output in results/ragas_k_6_{hybrid_new, vector_6, sparse_6}.csv:
| Metric | Hybrid (α=0.4) | Vector (α=1) | Sparse (α=0) |
|---|---|---|---|
| Answer Correctness | 0.941 | 0.890 | 0.909 |
| Context Precision | 0.847 | 0.805 | 0.880 |
| Context Recall | 0.980 | 0.972 | 0.968 |
| Faithfulness | 0.979 | 0.954 | 0.963 |
| — | |||
| Fully correct (AC = 1) | 222 / 250 | 207 / 250 | 210 / 250 |
| Partial (0 < AC < 1) | 21 / 250 | 25 / 250 | 27 / 250 |
| Wrong (AC = 0) | 7 / 250 | 18 / 250 | 13 / 250 |
Hybrid wins on Answer Correctness, Context Recall, Faithfulness, and the fully-correct count, and cuts the wrong-answer bucket to 7/250 — sparse alone leaves 13 wrong, vector alone leaves 18. Sparse edges hybrid only on context_precision (0.880 vs 0.847), because BM25 tends to fetch a tighter set of exact-keyword matches; the α=0.4 blend trades a bit of that precision for large gains everywhere else.
Note
Evaluation was performed on a multi-document mortgage packet (21 logical documents across 50 pages). Raw per-question output for all configurations lives in results/.
Model provenance for these numbers: the 250-question RAGAS runs above used gemini-2.5-flash-lite as the answer generator (and gemini-3.5-flash-lite as the judge) — chosen to keep the eval affordable. The deployed app, however, answers real user questions with gemini-2.5-flash (see GEMINI_CHAT_MODEL in backend/llm/llm_router.py).
backend/load_test.py spawns the real app with only the retrieval + LLM boundary stubbed by default, so a local run is free and takes seconds — it exercises the async endpoints, the connection pool, JWT auth and SSE, not the model. With --base it instead targets a running API: the --ramp --mix read mode only sends GET browse requests and does not call the model. Idle-vs-saturated phases, a --ramp capacity sweep, and --calibrate for a few real messages.
Live Render re-test, 2026-09-29 — --base https://document-retrieval-system-5gqx.onrender.com --ramp --mix read --levels 5,25,50,100 --duration 15 --ramp-stop-pct 1. The four 15-second levels ran against the deployed main API (not a locally spawned stub). A single dedicated account supplied the token; each "client" is a concurrent synthetic browse loop, not a distinct user or an AI answer. The mix calls /api/health, /api/chats and /api/chats/{id}. Raw results: ramp_2026-09-29_11-36-45.json.
| Concurrent browse clients | p50 API | p95 API | req/s | errors / throttled |
|---|---|---|---|---|
| 5 | 578ms | 672ms | 10 | 0% / 0 |
| 25 | 719ms | 1,641ms | 30 | 0% / 0 |
| 50 | 1,828ms | 3,079ms | 26 | 0% / 0 |
| 100 | 2,907ms | 4,735ms | 30 | 0% / 0 |
Measured healthy ceiling: ~25 concurrent browse clients under the harness SLO (<1% errors and p95 <3× the 5-client baseline). At 50 and 100, latency crosses that SLO even with no errors; throughput plateaus near 30 req/s. The p95 combines successful responses from the three GET endpoints and does not measure question-to-answer latency. This is a single short run, not a sustained capacity guarantee.
An earlier, unsaved Render run reported 8,640ms p95 at 100 clients, with no error-rate data file. That figure is historical, not the newly measured value. Earlier A/B pool work found that increasing connections from 15 to 30 roughly doubled read throughput (~18 → ~33 req/s); streaming /message still competes for DB connections with browse reads (_prepare / _save). The older ~3× /health degradation did not reproduce in local re-runs (32 → 31ms); the pool contention on /api/chats remains the relevant signal.
Measured on one dev box, the runs behind backend/load_test_results/. Absolute numbers are remote-Neon-from-a-laptop and mean little on their own; co-located on Render the floor collapses.
Saturated — 15 streaming answers, 5 browse clients, read mix (api_2026-08-31_20-44-10.json):
| endpoint | idle p95 | saturated p95 |
|---|---|---|
chats |
968ms | 2109ms |
chat |
1672ms | 1688ms |
health |
32ms | 31ms |
45 answers streamed, 0 errors. Capacity ramp (ramp_2026-08-31_20-46-39.json), read mix:
| clients | 3 | 5 | 10 | 15 | 25 | 40 |
|---|---|---|---|---|---|---|
| p95 | 1219ms | 1172ms | 1406ms | 1313ms | 1188ms | 2063ms |
| req/s | 4 | 7 | 13 | 20 | 33 | 39 |
0% errors at every level; estimated ceiling ~40.
Note
Quote the saturated figure, not the multiplier. It is tempting to compress the first table into a ratio (968 → 2109 is ×2.2), but the idle baseline wanders between runs on an unloaded box while the saturated value is stable — so the ratio moves even when nothing changed, and a lower ratio can mean a slower idle run rather than a faster loaded one. Repeat runs on this harness have disagreed on the multiplier while agreeing on the saturated milliseconds. Compare 2109ms, not ×2.2.
OpenTelemetry over OTLP, wired programmatically (not the opentelemetry-instrument wrapper). openinference's GoogleGenAIInstrumentor auto-traces every Gemini call:
- LLM spans → Langfuse + Grafana. Each message is one
chat-messagetrace with the rewrite and answer generations nested under it, tagged with user + session. - HTTP spans → Grafana (a separate provider, so Langfuse stays LLM-only).
chat_messages_totalmetric → Grafana, with a paste-importable dashboard and a muted error-rate alert (backend/grafana/).
Everything is a no-op unless the env vars are set, and nothing raises — tracing must never break a request. Set LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY / LANGFUSE_HOST, and GRAFANA_OTLP_ENDPOINT / GRAFANA_OTLP_AUTH (the full Basic <base64> header) / OTEL_SERVICE_NAME.
.github/workflows/ci.yml runs on every push and PR:
- backend —
ruff,compileall,load_test.py --selftestandcore/review.py(file-review rules and grounding guard), all with no network. - frontend —
npm ci,npm run lint,npm run build. - docker — build both images (no push), so a broken Dockerfile fails here, not at deploy.
- deploy — only after all three pass, only on push to
main: POSTs the Render deploy hooks (RENDER_DEPLOY_HOOK_*secrets), skipping gracefully if they're unset.
docker compose up --build runs the stack locally: backend/Dockerfile (python:3.12-slim + uvicorn, non-root, /api/health probe) and frontend/Dockerfile (Vite build → nginx). Extraction runs on AWS Textract, so the backend image needs no GPU/GL libraries.
The backend runs with --forwarded-allow-ips * (in the Dockerfile CMD) so slowapi's per-IP rate limits key on the real client (X-Forwarded-For) behind a proxy/balancer rather than the proxy's own IP — otherwise every user shares one rate-limit bucket. On a non-Docker deploy, set FORWARDED_ALLOW_IPS=* in the service env instead (uvicorn reads it).
document-retrieval-system/
├── backend/
│ ├── core/ # Processing & retrieval engine
│ │ ├── document_store.py # Orchestration + rehydrate()
│ │ ├── retriever.py # Hybrid search, routing, namespaces
│ │ ├── pdf_processor.py # AWS Textract integration (TABLES + FORMS)
│ │ ├── chunker.py # Structure-aware chunking
│ │ ├── document_classifier.py # Doc-type & boundary detection
│ │ ├── query_rewriter.py # Follow-up → standalone question
│ │ ├── answer_generator.py # Grounded answer prompt
│ │ ├── review.py # Automatic file review: fields, grounding guard, rules
│ │ └── models.py # Core dataclasses
│ ├── db/ # Neon / Postgres
│ │ ├── database.py # Engine + session factory
│ │ └── models.py # accounts, chat_sessions (+ review), messages, review_decisions
│ ├── llm/
│ │ └── llm_router.py # Gemini answers + embeddings
│ ├── eval/ # Measurement harnesses
│ │ ├── model_sweep.py # Generated questions, no judge (trust this)
│ │ ├── model_eval.py # flash vs flash-lite, LLM-judged
│ │ └── prompt_eval.py # Prompt-change A/B
│ ├── grafana/ # Grafana dashboard + alert provisioning
│ ├── auth.py # JWT + bcrypt + refresh tokens
│ ├── observability.py # OpenTelemetry → Langfuse + Grafana
│ ├── main.py # API entry point
│ ├── load_test.py # Capacity / responsiveness harness
│ ├── Dockerfile # python:3.12-slim + uvicorn
│ ├── migrations.sql # Schema + housekeeping queries
│ ├── requirements.txt
│ └── .env # Keys, DB URL, worker URLs
├── frontend/
│ ├── src/
│ │ ├── api.js # API client, owns the JWT
│ │ ├── Landing.jsx # Animated landing page
│ │ ├── Login.jsx # Login / signup
│ │ ├── ChatPanel.jsx # Upload gate, live ingest stepper, messages, review panel
│ │ ├── App.jsx # Shell, auth gate, chat rail
│ │ └── App.css # Design tokens & styles
│ ├── Dockerfile # Vite build → nginx
│ └── nginx.conf # SPA routing
├── .github/workflows/ci.yml # lint · build · docker · gated deploy
├── docker-compose.yml # local api + frontend stack
├── ruff.toml
├── samples/synthetic_borrower/ # Synthetic loan packet + answer key for the review
├── results/ # Ragas metrics (k=6 CSVs; results_old/ = pre-migration baselines)
└── README.md
MIT License