logoalt Hacker News

beastman82today at 4:04 PM8 repliesview on HN

can confirm.

I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.


Replies

fhubtoday at 8:44 PM

How are you deciding which work to send to the 5090 vs a frontier model, or making the two work together nicely?

Correct is much more important than fast for me, but if I could get correct and fast, that would obviously be amazing.

_hugerobots_today at 4:50 PM

Have a 5090, and yes it's very fast. But it's like the worst ADHD team member and requires constant supervision and review from larger models. It's context size on-card is good for super, suuuuuper shallow precision work. The gb10/spark on top of it, that thing can refactor enormous monorepo architecture. The time it takes the 5090 to compact, reiterate and execute a plan is often the same time as the gb10.

show 1 reply
tomega2134today at 5:52 PM

Is a 5090 still cost efficent when it is (currently) unobtainable? Or when obtainable only at current prices (min. $6500 USD)?

show 1 reply
nacstoday at 4:09 PM

People don't buy Sparks and M5 Ultras to run a 27B model - you buy it to run an MoE model like Qwen Next which this M5 excelled at.

show 1 reply
throwaway27448today at 4:35 PM

A) the macos value add is enormous if you have any investment in the ecosystem, B) for me at least a GPU is completely useless for anything but being a token generator.

show 1 reply
Eisensteintoday at 4:41 PM

A 5090 has a 1.79TB/s memory bandwidth. Qwen 3.8 27B NVFP4 is 22GB. You cannot generate tokens faster than the weights can traverse the GPU memory, so that makes max generation speed without MTP to be 81T/s. Say MTP is giving you 0.5 acceptance rate (very good), that is 1.5 * 81 is 121T/s. Even with a perfect acceptance rate you would only get 162T/s.

show 3 replies
cyanydeeztoday at 6:51 PM

Qwen3.8-Flash-Next is pretty damn worth the extra ram you need.

mathisfun123today at 4:08 PM

same reason they spend huge amounts of money on rolexes when seikos work better (the tech crowd isn't immune from vanity).

show 1 reply