logoalt Hacker News

nacstoday at 4:09 PM1 replyview on HN

People don't buy Sparks and M5 Ultras to run a 27B model - you buy it to run an MoE model like Qwen Next which this M5 excelled at.


Replies

ProllyInfamoustoday at 5:56 PM

Exactly; when I first got my RTX 5070 Ti (16gb, to game with!!!, upgrading from VEGA56), I loaded then-latest Qwen3.6 (~30B, cannot remember exactly). My only prior LLM experience was with models <8gb, primarily llama3.1.

My technical-expert twin played around with these LLMs, for about an hour, and then correctly reasoned "it's able to be WRONG, faster."

This seems apt. My next LLM machine will be closer to 96gb+ vRAM.

show 1 reply