llama.cpp is like the ffmepg of AI, and one of the reasons I so greatly dislike ollama is that the latter completely obfuscates that they're a rebrand of the former. Georgi Gerganov and team did all the hard work; ollama is langchain-like VC-bait with a HF download wrapper.
llama.cpp will happily download models from hugging face, btw.
ICYMI, llama.cpp was also VC funded. Search for "ggml" on this page:
https://aigrant.com