logoalt Hacker News

cmrdporcupinetoday at 11:28 AM3 repliesview on HN

Those two offer MoE variants, this doesn't seem to.

Dense model makes it dog slow on anything without HBM. Max 15tok/sec on decode on DDR5 systems like a Spark or a Strix Halo -- and that's at 4 bit quant.


Replies

EddieRingletoday at 12:28 PM

Dense models run at a very usable speed (Qwen 3.6 was running at ~50t/s last I looked) on my dual 7900 XTX desktop. (And before anyone brings it up, I did not buy them for this purpose, so the up-front cost is irrelevant in my case.)

Havoctoday at 1:00 PM

The benchmark comparison is against the dense variants not MoE

petutoday at 11:38 AM

3090/4090 probably would do 40 t/s, for 5090 75 t/s is shown in the blog.