logoalt Hacker News

petutoday at 3:51 PM0 repliesview on HN

Qwen3.6 is very token inefficient with it's thinking. Quantized versions often get into loops.

Glimmer is trained with 4 effort levels, not just thinking on/off. Maybe it's more token efficient in general. There's official 4 bit quantizations with reported 1% loss across 15 benchmarks -- so quants probably work good.

IMO that alone is worth trying for, even if they're otherwise equal.