logoalt Hacker News

snehesht • today at 12:51 PM • 5 replies • view on HN

I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

https://huggingface.co/Qwen/Qwen3.8-Flash-Next


Replies

roscas • today at 2:03 PM

Coder version with 30t/sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I've seen and 3080 had its days of glory.

I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t/sec and Laguna.XS-2.0.

➕ show 1 reply
thatsabadlook • today at 1:45 PM

Why is this surprisingly well? It's 2.5x faster than anthropic models, you have data sovereignty, privacy,and that's a strong model. Sounds like a best case scenario to me

➕ show 1 reply
jacquesm • today at 9:02 PM

Speed is one thing, accuracy another. Have you benchmarked it against a reference? If so, what were the results? I tend to go for accuracy over speed because usually that means fewer round trips and fewer tokens wasted.

proc0 • today at 1:03 PM

Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.

➕ show 3 replies
notnullorvoid • today at 3:06 PM

Which quantization are you using to reach those numbers?