logoalt Hacker News

Havoctoday at 11:01 AM1 replyview on HN

The favourable comparisons to Gemma 4 and qwen3.6 look promising!


Replies

cmrdporcupinetoday at 11:28 AM

Those two offer MoE variants, this doesn't seem to.

Dense model makes it dog slow on anything without HBM. Max 15tok/sec on decode on DDR5 systems like a Spark or a Strix Halo -- and that's at 4 bit quant.

show 3 replies