Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
27 commits
Select commit Hold shift + click to select a range
c7d788a
traces: recover tool schemas; more_plus: route on every phrasing in a…
pradipta-lyzr Aug 3, 2026
52af845
synth: the fourth inlet — training data for traffic that doesn't exis…
pradipta-lyzr Aug 3, 2026
209bc73
synth: reach it from the shell, the server, and the studio
pradipta-lyzr Aug 3, 2026
2271c78
docs: the synthesizer in the README, CLAUDE.md, and two runnable exam…
pradipta-lyzr Aug 3, 2026
32ea312
synth: fix what a live teacher broke that a fake teacher couldn't
pradipta-lyzr Aug 3, 2026
8ecdcaf
synth: stop the preference teacher from answering where the user shou…
pradipta-lyzr Aug 3, 2026
b20c517
synth: self-review — survive a flaky teacher, account for every row, …
pradipta-lyzr Aug 4, 2026
b837a55
synth: end saved JSONL with a newline
pradipta-lyzr Aug 5, 2026
cff6cfb
make: say what to install instead of 'No such file or directory'
pradipta-lyzr Aug 5, 2026
f5fca29
studio: rebuild _static with the Synthesize tab
pradipta-lyzr Aug 5, 2026
323ed26
synth: refuse a keyless OpenAI teacher at construction
pradipta-lyzr Aug 5, 2026
80454cd
synth: report progress per finished job, and stop serialising the doc…
pradipta-lyzr Aug 5, 2026
77588d7
methods: sdft + sdpo — self-distillation trainers for SFT and RL
ghanshyam-lyzr Aug 5, 2026
c632d97
Merge branch 'main' into feature/sdft-sdpo
ghanshyam-lyzr Aug 10, 2026
cd46861
synth: make a run meterable, stoppable, and durable
pradipta-lyzr Aug 10, 2026
e7ac46d
Merge upstream main into the data synthesizer
pradipta-lyzr Aug 10, 2026
3931f6a
docs: untangle the synthesis paragraph the merge left run-on
pradipta-lyzr Aug 11, 2026
909d199
serve: projects, evaluations, frontier model, capture, deployments
patel-lyzr Oct 8, 2026
3c8d5ac
studio: the fine-tuning cockpit, with Ctrl agent, on a restructured f…
patel-lyzr Oct 8, 2026
91ba09a
studio: rebuild the shipped UI
patel-lyzr Oct 8, 2026
36bc7de
docs: product context, the design system, and the cockpit's direction…
patel-lyzr Oct 8, 2026
3547756
Merge remote-tracking branch 'origin/main' into feat/cockpit-redesign
patel-lyzr Oct 8, 2026
58e2b43
Merge PR #14 (data synthesizer) into the cockpit redesign
patel-lyzr Oct 8, 2026
c291468
Merge PR #13 (sdft + sdpo) into the cockpit redesign
patel-lyzr Oct 8, 2026
d97053e
serve: synthesis writes a project's examples
patel-lyzr Oct 8, 2026
e1dfc4e
studio: the cockpit writes examples and offers SDFT
patel-lyzr Oct 8, 2026
9b3e98e
studio: rebuild the shipped UI
patel-lyzr Oct 8, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -21,3 +21,7 @@ frontend/node_modules/

# local secrets
.env

# Impeccable working files: review captures and decision-page logs
.impeccable/review/
.impeccable/questions/
492 changes: 492 additions & 0 deletions .impeccable/design.json

Large diffs are not rendered by default.

32 changes: 32 additions & 0 deletions .impeccable/surfaces/frontend-src-features-cockpit-cockpit-tsx.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
---
version: 1
slug: "frontend-src-features-cockpit-cockpit-tsx"
primary_target: "frontend/src/features/cockpit/Cockpit.tsx"
related_targets: ["frontend/src/features"]
---

# Fine-tuning cockpit

Scope: the studio's surface for making and improving one fine-tuned model (a project). Visitor mode: Operate. Both Business and Research modes use it; Research gets full settings, logs and every number in the inspector.

Audience and job: a business user turns their knowledge, examples or agent traffic into a proven model they can deploy; a researcher iterates versions on evidence. Success: from a sentence to an evaluated fine-tune in one sitting, then each next version justified by the scorecard.

Constraints: the opencontroller console visual system stays (tokens, shadcn primitives, light and dark); embeddable in opencontroller; every action Ctrl agent takes is visible, reversible where possible, and has an SDK/CLI equivalent; no invented metrics.

## Direction contract

THESIS: Ctrl agent and the loop are one instrument: a conversation proposes each next step as an approvable action, and a live map of the whole loop (data, fine-tune, evaluate, deploy) shows the consequence the moment you approve. It refuses the category default of a settings form followed by a separate runs page.

OWN-WORLD: The console's Paper & Ink: white canvas, warm hairlines, slate-blue tint for what can be acted on and what is live, earthy good/warning/destructive tints for verdicts. Action cards are bordered panels with one tinted primary and one quiet edit. Loop nodes are four hairline-framed stations on a single rail, the active one tinted, the rail drawn as one continuous line whose segment flows while work runs.

STORY: The visitor says what the model should do; Ctrl agent answers with a plan they approve; the map lights station by station; the scorecard lands in the Evaluate station and Ctrl agent proposes the next version or the deploy. They leave knowing whether their model is ready and why.

FIRST VIEWPORT: Left 7/12: the loop rail across the top (Data → Fine-tune → Evaluate → Deploy, each station with its status line and version chips), the selected station's inspector filling the rest (examples, live loss curve, scorecard, endpoint). Right 5/12: Ctrl agent's conversation, header with the project's name and stage, thread of Ctrl agent's messages and action cards, composer pinned at the bottom with suggested next actions as chips (stacked on a phone, Ctrl agent comes first). The primary action is always the newest action card's tinted button in the conversation.

FORM: Copilot cockpit (the agent named Ctrl agent, placed on the right by the user after the build), position 1 of 7 in the ranked list (dealt second); seed key f7febb6b. Signature interaction: approving an action card lights its station and runs the rail segment into it; hovering a message highlights its station; clicking a station scrolls the thread to its latest message. Motion grammar: 150–250 ms ease-out state changes, a dash-flow on the active rail segment while a job runs, all of it off under reduced motion.

FINISH: unreviewed and undocumented is unfinished; this build ends with the finish review, the verdict, DESIGN.md, and every shipping raster carrying its provenance

## Unresolved

- Free-text understanding is rule-based (intents and goal classification) in this build; planning with the user's frontier model is a follow-up.
73 changes: 68 additions & 5 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
## What this is

ShadowLM Trainer is a fine-tuning SDK: load any open model, train it with any of
13 methods, on any hardware, then own the weights. The headline use case is
15 methods, on any hardware, then own the weights. The headline use case is
"shadowing" — moving one task off a rented frontier model onto a small model you
own, by capturing real agent traffic (`slm.capture()`), judging episodes, and
training on them — without modifying the agent (the model API is the only
Expand Down Expand Up @@ -87,21 +87,57 @@ touches one file and no others.

### The shadowing / agent-tuning loop

There are four ways data gets in, and they all end at a `Trajectory`:
`capture.py` (live), `traces.py` (already ran), `synth/` (doesn't exist yet), and
`Dataset.from_*` (you have a file).

- `capture.py` — `slm.capture(model)` is a drop-in OpenAI-compatible proxy that
records an unmodified agent's traffic, reconstructing message-level
trajectories (calls that extend a prior call's message prefix merge into one
episode; use an `x-session-id` header to disambiguate interleaved conversations).
- `traces.py` — the offline sibling of `capture.py`: ingests OpenTelemetry GenAI
spans (OTLP JSON, event/spec/OpenInference/raw-wire shapes), groups them into
conversations, and `to_dataset()`s them. For agents already instrumented —
no proxy in the path.
spans, groups them into conversations, and `to_dataset()`s them. Four dialects
are read (OTel spec `{role,parts}`, OpenInference, indexed OpenLLMetry,
OpenAI-wire blobs). For agents already instrumented — no proxy in the path.
- `rl.py` — `Trajectory` / `TrajectoryGroup` / `judge_group` (LLM-judge scoring),
fed into `method="grpo"`.
- `apo.py` — `optimize_prompt()`: optimize the prompt instead of weights, same
capture/judge front end, no GPU.
- `eval.py` — `slm.evaluate()` / `shadowlm eval`: score a model on a held-out set
(exact / contains / numeric / JSON / LLM-judge scorers).

### The synthesizer (`synth/`)

The fourth inlet — for traffic that doesn't exist yet (cold start, amplifying a
handful of episodes, covering cases production never hit). Two orthogonal axes
again, mirroring backends × methods:

- **seeds** (`seeds.py`) — where scenarios come from: a plain-English `task`, a
`document` (chunked, facts extracted, answers judged against the passage), or
real `episodes` to vary. Any seed composes with any mode.
- **modes** (`generate.py`) — what gets written per scenario: a `conversation`,
a `preference` pair, or `paraphrases`. Generation is always **taxonomy first,
instances second** — that structure, not prompt wording, is what stops mode
collapse.

Everything converges on `Trajectory` (the same type capture and traces produce),
then `emit.py` renders it into the shape the consumer takes. **Output shape is
chosen from the method's spec, never its name** (`resolve_output`). `to_otlp` is
the exact inverse of `traces._spec_message` — change one and you must change the
other; `tests/test_synth_otlp_roundtrip.py` is what holds them together.

`quality.py` validates (the "must end on an assistant turn" rule is load-bearing
— see `torch.py:_train_dataset`), deduplicates, and gates on a judge score.
Nothing is dropped silently: `SynthReport` reconciles exactly, and
`report.balanced` asserts it.

A run costs money, so it is meterable and stoppable. Tokens come from the
provider's own `usage` block — never estimated, and there is deliberately no
price table to go stale. `token_budget=` and `should_stop=` end a run early
while **keeping** what it produced; both gate *generation* only, because
leaving already-generated rows unscored fails them at the gate and wastes the
whole spend. Studio runs persist under `work_root/synth/` and are cancellable.

### Signature methods (MoRE)

`more.py` / `more_plus.py` implement "mixture of retrieval experts" — facts fused
Expand All @@ -128,13 +164,40 @@ routing; its run progress is one step per unit (see `resolve_total_steps`).
Same protocol backs `backend="remote"` and ShadowLM Studio. `SHADOWLM_API_URL`
may list several servers; `pick()` binds to the least-busy reachable one for
the session (client-side routing, deliberately not a scheduler).
- `frontend/` — React 19 + Vite + Tailwind v4 studio. `npm run build` outputs to
- `frontend/` — React 19 + Vite + Tailwind v4 studio. Layout: `src/app/` (shell,
`router.ts` — one typed route table over the URL hash, the format the embed
bridge speaks), `src/features/<feature>/` (one folder per surface: cockpit,
projects, datasets, models, train, runs, evaluate, deployments, playground,
machines, overview), `src/components/` (ui primitives + shared pieces),
`src/lib/` (`queries.ts` is the react-query data layer: one hook per
resource, polling only while something runs). The **cockpit**
(`features/cockpit/`) is where a model is made: the Ctrl agent conversation
on the right (`copilot.ts` narrates state and proposes the next action; rule-based, no
model call) drives a live map of the loop on the left (`model.ts` folds every resource
into Data → Fine-tune → Evaluate → Deploy stations). Its direction contract
lives in `.impeccable/surfaces/`. `npm run build` outputs to
`../shadowlm/_static` (the wheel ships the compiled UI; end users never need
node). `frontend/src/api.ts` is the typed mirror of the remote protocol. The
pages (Dashboard · Datasets → Models → Train → Runs → Playground · Machines)
are the capture→train→own loop as a UI. Auth has three modes — `password`,
`apikey`, or `none` (`GET /v1/auth` reports which) — plus long-lived, hashed,
individually-revocable **machine tokens** that workers authenticate with.
The studio has two modes over the same objects (`frontend/src/lib/mode.ts`,
asked once on first visit, switched in the side panel): **Business** walks a
*project* (`/v1/projects`: one fine-tune for a job, goal knowledge / task /
takeover) through Data → Fine-tune → Evaluate → Deploy, with base model and
method picked by `lib/recipe.ts` and shown; **Research** keeps the object
pages (datasets, models, runs) plus **Evaluate** (`/v1/evals`: several
targets scored on the same questions, run as a background job and persisted
under `<work>/evals/`). Chat runs as a background task too
(`/v1/tasks/chat`, polled by `chatAsync`), because a held request dies at the
proxy's 100 s. The user's **frontier model** (settings: OpenAI-compatible
base URL, key, model; `shadowlm/frontier.py`) is an eval baseline, the
`judge` metric, and the upstream for **agent capture** (`/v1/capture/<id>`
passes calls through and records them; `capture.reconstruct` turns them into
episodes → a dataset). **Deployments** serve a fine-tune at `/openai/v1`
(OpenAI-compatible, a hashed per-deployment key, no studio login). Product
context and principles live in `PRODUCT.md`.
The UI uses the opencontroller console's design system (shadcn primitives in
`frontend/src/components/ui/`, tokens in `index.css`, light + dark) and can
run **embedded** in a host console over the `oc-embed/1` postMessage bridge
Expand Down
Loading
Loading