Skip to content

Cpu cache for unload - #507

Draft
brycehenson wants to merge 10 commits into
remsky:masterfrom
brycehenson:cpu_cache_for_unload
Draft

Cpu cache for unload#507
brycehenson wants to merge 10 commits into
remsky:masterfrom
brycehenson:cpu_cache_for_unload

Conversation

@brycehenson

@brycehenson brycehenson commented Aug 17, 2026

Copy link
Copy Markdown

adds a cpu cache of the model to reduce load speed

      - MODEL_UNLOAD_STRATEGY=cpu_cache

Benchmark

Case Time
Fresh model load, average 1.692s
CPU RAM to CUDA, average 0.193s

depends on #506 which should go in first

TODO

  • remove notes.md

@RBEmerson970

Copy link
Copy Markdown

adds a cpu cache of the model to reduce load speed

Reduce load speed, or reduce load time? I don't understand why reducing load speed would be a good thing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants