> Unlike Ollama, [llama.cpp] serves one model at a time: the GGUF you launched it with.
FWIW: llama-server's supported multiple models for a while now:
https://github.com/ggml-org/llama.cpp/tree/master/tools/serv...
> Unlike Ollama, [llama.cpp] serves one model at a time: the GGUF you launched it with.
FWIW: llama-server's supported multiple models for a while now:
https://github.com/ggml-org/llama.cpp/tree/master/tools/serv...