|
I saw my memory usage went to ~98% during LLM model load to GPU (I had not much free RAM at that moment) and then "Terminated" in terminal. So I've added MMAP option in koboldcpp and it loaded, albeit longer. The delay was on "Model warm up" message in terminal. But it took like ~20 seconds to generate 1st token. Smaller model that loaded without MMAP worked equally fast loaded with and without MMAP. With MMAP RAM usage for smaller model was lower all the way (after load too). For larger model memory usage was about same as for smaller (both when loaded with MMAP). Larger worked fine (fast) before - when free RAM was abundant. Questions: 1) why larger model so slow if should be fully in VRAM after load? If not in VRAM, why? Can I fully put LLM into VRAM with koboldcpp (or if not with other tool)? 2) seems after load into GPU (Vulkan on NVIDIA) RAM taken to model load is not released. Why? 3) What does it mean in koboldcpp GUI MMAP help says "model will not be unloadable"? How to upload models if MMAP is not used? |
Replies: 2 comments 1 reply
|
|
This memory-pressure behavior is something I’ve also been experimenting with. I’m developing an experimental C++ GGUF runtime called MemVanta that deliberately prioritizes resident-memory reduction rather than maximum throughput. On one OpenLLaMA 7B v2 Q4_0 CPU-only comparison using the same GGUF:
In a separate cgroup-v2 test with swap disabled, MemVanta completed at a tested 3584 MiB memory ceiling where the pinned llama.cpp comparison was OOM-killed. The cost is throughput — llama.cpp is substantially faster at present. The approach uses mmap-backed tensor access, bounded caching and paged KV state. I would be particularly interested to know whether KoboldCpp users working on RAM-constrained machines see similar page-cache/RSS trade-offs. If anyone wants to independently reproduce or challenge the result, the raw evidence and reproduction instructions are here: |
Loading speed largely depends on your storage (e.g. SSD vs HDD).
yes you can do a full offload if you have enough VRAM. Check in task manager. If you share your GPU we can advise.
if you use mmap to load, the model will stay in memory until koboldcpp.exe closes. This may not be desired if you plan to swap model at runtime e.g. using the admin endpoint, but if you don't change model while using this is not an issue.