logoalt Hacker News

bitexploder • today at 5:05 PM • 1 reply • view on HN

That is what stood out to me as well. Qwen 3.8 Flash can run on very limited hardware as well and it is at least as good as Sonnet 5 in benchmarks like DeepSWE. With its n-gram design you can get flash next running on very limited GPU resources, as little as 16GB of VRAM.

People follow the latest frontier lab models with great attention and migrate to the next big model on their subscriptions. Meanwhile these local models have quietly gotten REALLY good. It is not even an exaggeration. It has happened in the last couple of months.

"Local model you can run at 40 t/s on a gaming machine that is better than Opus 4.6" is way less exciting than "OpenAI IS DOING CRIME!!! OpenAI SOLVED NAVIER STOKES. DARIO SAYS GLM 5.3 BAD! SLOW DOWN THE FRONTIER!".

(edit: also... totally ignore that 27B dense column over there where Qwen 3.8 27B beats Kolibri on nearly every single benchmark. Why would I choose to run this model?)


Replies

frumplestlatz • today at 7:06 PM

I think the simple reality is that if your goal is producing quality work output, you want the smartest possible model available.

What is the upper bound on the value of more intelligence applied to your problem domain?

➕ show 1 reply