Technical map of this repository: what each stage does, how data flows
between them, and what's planned next. AGENTS.md is the conventions/how-to-run
guide; this file is the "what is this and why is it shaped this way" guide.
This repository is a staged model-training workspace, split by pipeline stage rather than by model:
pre-training/ PDF corpus -> OCR text, summaries, layout, synthetic QA (CSVs)
fine-tuning/ text/summary pairs -> trained LoRA adapters
serving/ trained adapters -> inference (FastAPI)
training/ reserved for from-scratch / non-LoRA training (not yet built)
Each leaf folder (pre-training/, fine-tuning/<pipeline>/,
serving/<pipeline>/) is an independent uv project: its own
pyproject.toml, uv.lock, .python-version, and pinned dependency set (in
particular, its own CUDA torch build). Nothing is shared at runtime between
folders — a pipeline can be deleted or reworked without touching its
siblings. One uv binary at the root drives all of them via
uv run --directory <folder> ... (see the root README.md's uv Commands
section for the current, verified list); there is deliberately no root-level
Python project or shared virtualenv, since the folders pin conflicting
dependency versions (e.g. different torch builds) that a single shared
resolution would fight.
Repo-wide hygiene is intentionally centralized rather than per-folder: one
.gitignore at the root (unanchored patterns match every project's drop-zone
folders at any depth), and one AGENTS.md/CLAUDE.md pair at the root
covering conventions for the whole repo. No subfolder should have its own
copy of any of these.
Every project pins .python-version to 3.12. Without it, uv run picks
the newest CPython it can find (e.g. 3.14), and some pinned dependencies —
pillow==10.4.0 in particular — have no prebuilt wheel for that new a
version yet, so uv falls back to a from-source build that fails on Windows
(missing zlib headers). Pinning 3.12 (already installed and known to have
wheels for every pinned dependency across all four projects) is what makes
uv run --directory <folder> ... work reproducibly from a clean checkout.
This was verified by reproducing the failure and fixing it during this pass —
see the Verified working list below.
Turns a PDF corpus into training data. Local, GPU-first, Surya OCR +
Gemma 3 (unsloth/gemma-3-4b-it, an ungated mirror — no HF_TOKEN needed).
Five steps (exec_1.bat … exec_5.bat, or main.bat for an interactive
menu): PDF → PNG pages → OCR CSV → summary CSV / layout CSV / synthetic-QA
CSV, all written per-run to outputs/[timestamp]_[dataset]/.
Step 1 (PDF → PNG) is not a standalone Python entry point — it's
scripts/convert_pdf_to_png.ps1, which shells out to poppler
(pdftoppm/pdfinfo on PATH) and a small compress_png_max.py helper.
Run it via exec_1.bat or the .ps1 directly, not uv run ... python scripts/convert_pdf_to_png.py — that file doesn't exist (an earlier version
of this doc incorrectly assumed it did; fixed here). Steps 2–5
(ocr_detection_png.py, summarize_ocr_gemma.py, describe_layout_gemma.py,
generate_qa_gemma.py) are genuine argparse Python scripts and do run via
uv run --directory pre-training python scripts/<name>.py.
Status: out of scope for active work — left as-is, beyond the
.gitignore/.python-version consolidation described above.
Two example pipelines, same transformers+peft pattern, both LoRA-based,
both sized for a single RTX 3090 (24GB). (A third, Axolotl-based
axolotl-ocr-summary/ pipeline existed earlier but was removed by the repo
owner — it only resolved its uv environment on Linux/WSL, never natively
on Windows, since axolotl[deepspeed] depends on triton.)
| Pipeline | Framework | Base model | Data shape |
|---|---|---|---|
fine-tuning/vicuna-7b-lora/ |
transformers + peft (manual Trainer loop) |
lmsys/vicuna-7b-v1.5 — LoRA on q_proj/v_proj, loaded directly via AutoModelForCausalLM/AutoTokenizer. No LLaVA checkpoint, no vision encoder, no multimodal projector anywhere in the dependency graph. |
JSONL with text / summary fields |
fine-tuning/qwen25-3b-lora/ |
transformers + peft (same pattern as vicuna-7b-lora/) |
Qwen/Qwen2.5-3B-Instruct — LoRA on q_proj/v_proj (same target modules as Vicuna; Qwen2ForCausalLM uses the same separate Q/K/V/O naming, confirmed via peft's own default LoRA target-module table). ChatML prompt format instead of Vicuna's USER:/ASSISTANT: (verified against the tokenizer's chat_template/eos_token). |
JSONL with text / summary fields |
Naming/loading history: vicuna-7b-lora, previously llava15-lm-lora,
originally llava15-lora. Two renames, each fixing a real overstatement:
llava15-lora→llava15-lm-lora: the pipeline only ever LoRA'd the language-model backbone (q_proj/v_proj) ofllava-hf/llava-1.5-7b-hf— never the vision encoder or multimodal projector, never an image. The-loraname alone overstated that as a VLM fine-tune.llava15-lm-lora→vicuna-7b-lora: even loading LLaVA's checkpoint at all was unnecessary once the vision half was never used — it still downloaded the full ~14 GB multimodal weights viaLlavaForConditionalGeneration/AutoProcessorto get to a submodule that is, in substance, Vicuna-7B. This pass switched to loadinglmsys/vicuna-7b-v1.5directly viaAutoModelForCausalLM+AutoTokenizer— ~13 GB instead of ~14 GB, no vision-related code path in the dependency graph at all, same LoRA config/target modules. Trade-off:lmsys/vicuna-7b-v1.5is the checkpoint LLaVA 1.5 was later visually-instruction-tuned from, not the LLaVA-tuned weights themselves — different starting point, not directly comparable to the oldllava15-lm-lorarun's results, but a cleaner, smaller, honestly-named base for a language-model-only LoRA.serving/llava15-lorawas renamed toserving/vicuna-7b-lorain step, with matching model-loading changes (functionally required: a Vicuna-7B-trained adapter's parameter names don't matchLlavaForConditionalGeneration'slanguage_model.*prefix, so serving would fail to load it otherwise). A plannedllava15-full-lorasibling, trained on image+text pairs and actually exercising the vision encoder/projector, remains the natural first real VLM fine-tune in this repo — that one should load the full LLaVA checkpoint. Full reasoning infine-tuning/vicuna-7b-lora/README.md.
vicuna-7b-lora/ is a generic text-summarization LoRA, not OCR-specific —
its interface, data source, and default prompt were all cleaned up this pass
to reflect that:
- JSONL field is
text(wasocr_text);build_vicuna7b_dataset.pyandgenerate_vicuna7b_lora.py's flags are--source-csv/--text/--text-file(were--ocr-csv/--ocr-text/--ocr-text-file). build_vicuna7b_dataset.pyonly builds from a CNN/DailyMail Parquet dump now (--cnn-dailymail-dir, required) — the earlier dual-source mode that also read pre-training's image-linked OCR/SUMMARIES CSV pair (normalize_image_key/resolve_image_path/load_summaries) was removed entirely, not just renamed, since it's not needed for this pipeline's current use (generate_vicuna7b_lora.py's--source-csvbatch-eval mode still accepts any generic CSV with atextcolumn, unrelated to that removed ingestion path).train_vicuna7b_lora.py/generate_vicuna7b_lora.py'sDEFAULT_INSTRUCTIONis now the CNN/DailyMail news-article wording (was "Summarize this scanned document page... UAP-related content") — since that's the only source this pipeline builds from,--instructionno longer needs to be passed explicitly for the common case.
qwen25-3b-lora/ is a near-clone of vicuna-7b-lora/ — same dataset
builder logic, same trainer/generator structure, same CLI shape. Two
verified differences (not assumed): the ChatML prompt wrapper (see table
above), and no protobuf/sentencepiece dependency needed (Qwen2.5-3B-Instruct
ships a ready tokenizer.json, unlike Vicuna's raw SentencePiece tokenizer).
No serving/qwen25-3b-lora/ yet — serving/vicuna-7b-lora/ is Vicuna-specific
(ChatML wrapper differs), so a sibling serving folder would be needed if this
adapter goes to production.
Downloaded locally to C:\Users\luisarandas\Desktop\cnn_dailymail\3.0.0\
(outside the repo, outside the root folder, gitignored regardless). Measured
against the actual files:
| Split | Rows | Size |
|---|---|---|
| train (3 shards) | 287,113 | ~772 MB |
| validation | 13,368 | ~35 MB |
| test | 11,490 | ~30 MB |
| Total | 311,971 | ~799 MB |
article (avg ~3,950 chars) → text, highlights (avg ~260 chars) →
summary. build_vicuna7b_dataset.py --cnn-dailymail-dir ... --max-samples 2000 was run end-to-end against the real files and produces valid JSONL
records; full details and commands in
fine-tuning/vicuna-7b-lora/README.md.
Before the switch to loading Vicuna-7B directly, a 2,000-sample run on the
old llava15-lm-lora pipeline (1,800 train / 200 val, 1 epoch, 450 steps,
~31 min on a single RTX 3090) showed loss dropping 1.66 → ~1.12 in the first
~50 steps then plateauing in a ~1.0–1.2 band, with eval_loss (1.11)
tracking train loss closely (no overfitting) — a normal curve for a
rank-16, 2-projection adapter on a small dataset, not evidence of a broken
run. That adapter and its data/hf_cache were deleted as part of this pass's
switch to lmsys/vicuna-7b-v1.5 (different base weights, not compatible
with the old adapter) — the numbers above are illustrative of the expected
curve shape, not a claim about the current pipeline's untrained state.
What carries forward: loss alone doesn't say whether summaries are actually
good, so generate_vicuna7b_lora.py has a --jsonl-eval <path> --num-samples N reconstruction-test mode (added the same pass as the old
run above) — it replicates the trainer's train/val split (same seed/ratio)
and prints genuinely held-out source/reference/generated triples with
token-F1, instead of requiring a manual --text string. Run against the old
adapter it produced coherent, on-topic, correctly-bulleted CNN/DailyMail
summaries (avg token-F1 0.357 across 5 samples) despite the plateaued loss —
this is the tool to use to judge the next real run on the current pipeline,
not the loss curve. See the
pipeline README.
One serving/<pipeline>/ folder per fine-tuning pipeline that has a serving
story. Currently: serving/vicuna-7b-lora/,
a FastAPI service (app.py) that loads the base Vicuna-7B model + trained
adapter (or a fused/merged model) once and serves a JSON API plus a
dataset-browser front-end. Deliberately decoupled from
fine-tuning/vicuna-7b-lora/ — it only reads the trained adapter directory
(../../fine-tuning/vicuna-7b-lora/runs/vicuna7b_lora/final_adapter), never
imports its training code. uv run --directory serving/vicuna-7b-lora python app.py --help verified working.
Reserved for from-scratch / non-LoRA training of other models, as distinct
from adapting an existing checkpoint (fine-tuning/).
training/adult-income-logreg/ is
the first pipeline here: logistic regression on the UCI
Adult / Census Income
dataset, implemented with raw numpy rather than scikit-learn — the sigmoid,
binary cross-entropy loss, gradient derivation, and gradient-descent update
loop are all written out by hand in train_logreg.py so the math stays
visible, and build_income_dataset.py parses the raw CSV files and does
one-hot/z-score encoding without pandas. Own uv project like every other
pipeline folder, but no torch/CUDA dependency at all — just numpy.
Verified end-to-end against the real dataset: 30,162 train / 15,060 test
rows after dropping "?" rows (matches the cleaned-variant counts in
adult.names exactly), 300 epochs of batch gradient descent, 84.6% test
accuracy (in line with the 84–86% published for tree-based methods on this
same cleaned split — a from-scratch linear model landing close to that is
the expected sanity-check result, not a target to beat).
None of the fine-tuning pipelines ship data — DATASET/, data/, runs/,
output/ etc. are all git-ignored, drop-zone folders (via the single root
.gitignore). Both vicuna-7b-lora/ and qwen25-3b-lora/ train on
CNN/DailyMail. Example small, permissively-licensed public datasets are
listed in the root README.md's Datasets section.
2,000-sample JSONL (1,800 train / 200 val), 2 epochs, 900 steps, ~61 min on
a single RTX 3090. Train loss 1.76 → 0.99, eval_loss essentially flat across
epochs (1.100 → 1.094 — a mild overfitting signal in isolation, train loss
kept falling while eval_loss didn't). What matters: reconstruction-test avg
token-F1 rose from 0.357 (an earlier 1-epoch/1,800-sample run on the
predecessor llava15-lm-lora pipeline) to 0.467 on this run, and the
generated summaries reproduced exact figures from source text correctly
(e.g. "383-41" and "70-26" vote counts). Confirms the earlier lesson again:
eval_loss plateauing is not itself a stop signal — the reconstruction test
is what actually shows whether a further epoch helped.
uv run --directory <folder> python <script> --help, and further real
executions where noted, actually run, not assumed:
fine-tuning/qwen25-3b-lora—build_qwen3b_dataset.py,train_qwen3b_lora.py,generate_qwen3b_lora.py, plus a real 40-sample smoke train against the actual downloadedQwen/Qwen2.5-3B-Instructweights, confirmingtrainable params > 0(LoRA genuinely attached toq_proj/v_proj) rather than trustingpeft's target-module table alone.fine-tuning/vicuna-7b-lora— real 2,000-sample/2-epoch training run (see above), executed by the repo owner, not just a smoke test.serving/vicuna-7b-lora—app.py --help(also fixed aSyntaxWarningfrom unescaped backslashes in three docstrings:app.py,merge_adapter.py,inspect_weights.py).
Not executed this pass (no PDFs/poppler set up in this environment):
pre-training/exec_1.bat/scripts/convert_pdf_to_png.ps1
fine-tuning/vicuna-7b-lorahas a verified-good real run (see above) — reasonable next moves are more samples (the eval_loss plateau suggests more epochs on this same 1,800-row set has limited further upside), judged by the reconstruction test, not loss alone.fine-tuning/qwen25-3b-lorais smoke-tested but not yet trained for real — same next step as Vicuna's first run: build a few-thousand-sample JSONL, train, then judge with--jsonl-eval.training/(from-scratch, non-LoRA) is still undesigned.fine-tuning/llava15-full-lora(planned, not started): the first real VLM fine-tune in this repo — image+text pairs, vision encoder/projector actually in the training graph, unlikevicuna-7b-lora/qwen25-3b-lora.- A
phi35-mini-lorasibling (discussed, not started) would needtarget_modules=["qkv_proj"]instead of["q_proj", "v_proj"]— Phi-3 fuses Q/K/V into one linear layer (confirmed by readingPhi3Attention's source), so thevicuna-7b-lora/qwen25-3b-loratarget-module config would silently attach to nothing on that model.