What about Bonsai 2? You can fit Qwen3.8 27B on an 8GB GPU with it, and upstream llama.cpp support is already being worked on (they just got System1 support too).
The ternary model? Hopefully those are worth a damn in a few years, but currently just an interesting toy from what I understand.
The Bonsai models are really bad when you actually use them for more than short responses.
Their marketing made it look like a breakthrough, but in my experience it’s just the next step down from the Q2 quants in both size and quality.
Q2 quants are already not very useful in my experience. The Bonsai models are even worse.
If you only need 80% plausible outputs that don’t need to reference a lot of context they can be useful. If you try to use them for real tasks it feels like time warping back to 2023 when you LLMs were barely useful if you babysat every word of the output.