Skip to content
Discussion options

You must be logged in to vote
  1. Loading speed largely depends on your storage (e.g. SSD vs HDD).

  2. yes you can do a full offload if you have enough VRAM. Check in task manager. If you share your GPU we can advise.

  3. if you use mmap to load, the model will stay in memory until koboldcpp.exe closes. This may not be desired if you plan to swap model at runtime e.g. using the admin endpoint, but if you don't change model while using this is not an issue.

Replies: 2 comments 1 reply

Comment options

You must be logged in to vote
0 replies
Answer selected by alex-ie
Comment options

You must be logged in to vote
1 reply
@alex-ie
Comment options

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
3 participants