I tried
curl -LsSf https://llama.app/install.sh | sh
and then llama serve -hf unsloth/Qwen3-4B-GGUF:Q4_0
Then I get: W load: control-looking token: 128247 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden
Terminated
And the web interface says Server unavailable
Maybe it gets killed by the OS because it uses too much RAM?When I try
llama serve -hf unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
It seems to work. Nice.