Skip to content

Free the GPU after a run finishes; one waiting row in shadow mode - #18

Merged
patel-lyzr merged 3 commits into
mainfrom
fix/release-gpu-after-runs
Oct 7, 2026
Merged

patel-lyzr merged 3 commits into
mainfrom
fix/release-gpu-after-runs

Conversation

@patel-lyzr

Copy link
Copy Markdown
Collaborator

Why

  • After lyzr-assistant (Qwen3-8B MoRE) finished, the studio box kept ~39 GB on the GPU. Clean VRAM unloaded the playground cache but couldn't touch it: the runner loop's locals (be, result, callbacks) held the trained model until the next job.
  • In shadow mode, once the shadow answered, the playground drew a second pair of waiting panes, so one question looked like two.

What

  • serve.py: after each run is persisted, drop the runner's references and release the allocator's cache (_release_gpu_cache, shared with the inference-cache drop).
  • Playground.tsx: draw the waiting pair only until the shadow answers; after that its row carries the base's wait.
  • Rebuilt _static.

Testing

  • make test: 225 passed. New test_a_finished_run_lets_go_of_its_model runs a job on a fake backend and asserts the backend is garbage-collected after the run; it fails without the fix.

The runner's locals kept the trained model alive until the next job, so the
GPU stayed full between runs (39 GB after a Qwen3-8B MoRE run) and Clean VRAM,
which only drops the inference cache, couldn't free it. Drop the references
and release the allocator's cache after each run.
Once the shadow answered, the playground still drew a second pair of waiting
panes under the row already waiting on the base, so one question read as two.
@patel-lyzr
patel-lyzr merged commit f76686f into main Oct 7, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant