logoalt Hacker News

0xbadcafebee • today at 1:51 PM • 2 replies • view on HN

Lol, sure, if you quant it to hell (Q2) it'll go real fast...

They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.

It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.


Replies

snehesht • today at 2:18 PM

You're right but for simple use cases its useful. Someone pointed to ds4 + qwen3.8 with Q4_K will try that out.

sigbottle • today at 2:20 PM

It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?

➕ show 3 replies