Skip to content

Free the playground's GPU models before training; show why a run failed - #17

Merged
patel-lyzr merged 3 commits into
mainfrom
fix/train-frees-playground-vram
Oct 7, 2026
Merged

patel-lyzr merged 3 commits into
mainfrom
fix/train-frees-playground-vram

Conversation

@patel-lyzr

Copy link
Copy Markdown
Collaborator

Why

Two Qwen3-8B MoRE runs on the studio box failed at step 0 — one with CUDA out of memory (44.37 of 44.39 GiB in use), one with Cannot copy out of meta tensor. The playground had Qwen3-8B cached on the GPU, and training loaded a second copy next to it. The same config trains fine (200 steps) on a free card.

What

  • serve.py: before the local runner loads a base, it drops the inference cache and releases the allocator (shared with clear_vram), and logs unloaded N playground models to free the GPU for training in the run's console.
  • Runs.tsx: a failed run shows Why it failed above the tabs; it was only under Artifact.
  • Rebuilt _static.

Testing

  • make test: 224 passed (2 new in tests/test_serve_vram.py).

A big base trained next to a cached inference copy of itself ran the GPU out
of memory (Qwen3-8B MoRE on a 48 GB L40S: OOM at step 0, or accelerate
offloading to meta and the load failing). The local runner now drops the
inference cache and releases the allocator before it loads the base, and
says so in the run's console.
The error only showed under the Artifact tab, so a failed run looked like one
with no data yet.
@patel-lyzr
patel-lyzr merged commit 74df4e6 into main Oct 7, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant