Issue is that llama.cpp is the best way to run models on hardware that isn't nvidias.
Except when they have less than 16 gb of ram?
Except when they have less than 16 gb of ram?