logoalt Hacker News

fsiefken • today at 3:20 PM • 4 replies • view on HN

I wonder if a higher Qwen3.8-27b quant could beat or match these lower < 16/24/48/64G Qwen3.8-Flash Next quants given similar quality.

What speed are you willing the sacrifice to debug/program for more complex jobs faster?

Then there are also these quants; https://huggingface.co/IsValorum/Qwen3.8-35B-A3B-Distill-MLX...


Replies

happycube • today at 7:52 PM

Maybe, but those higher quants would need a large GPU accessible memory space - and the obvious candidates such as DGX Spark and Strix Halo don't have the bandwidth to run 27B at high quality quickly.

With Flash Next you only have ~6B active parameters so you can toss experts up into VRAM and/or run them on a CPU if you have enough RAM and bandwidth.

latentsea • today at 4:17 PM

The benchmark indicates the IQ3_XXS quant beats 27B. I've switched to that now and am ditching 27B. Genuinely better results so far.

zkmon • today at 3:28 PM

I'm not going to knock off my 27B-Q_6 for this. Good to to experiment though.

xreborn • today at 3:22 PM

from my experience dense models like 27b suffer less from quantization compared to large MoEs