logoalt Hacker News

peri-cltoday at 4:12 PM2 repliesview on HN

Surprisingly, the Reddit crowd are reporting 50–60 tokens/s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth,

https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f...

(Note it's a sparse MoE with only 6B active).


Replies

bitexplodertoday at 7:53 PM

I have a 3 year old gaming system. RTX 4080 w/128GB of DDR5. It runs Qwen 38 Flash around 44-40 t/s with 128K context. It is on a specialized build that caches MoE experts and uses an optimized 3bit quant that basically is within a few points of the full 8 bit quant. In general, in casual benchmarking with Alibaba's endpoint I could not tell much of a difference. Overall this model is very good on long horizon agentic work. The main pain point for it is that its input processing speed is slow. Regardless, it gets meaningful work done.

I paid $500 for the RAM in Nov 2023 :)

show 1 reply
nacstoday at 4:24 PM

Good to know thanks.

That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.

show 1 reply