logoalt Hacker News

Roark66 • today at 2:07 PM • 0 replies • view on HN

Bingo, DeepSeek (v4.1) is horribly overhyped. In all my personal benchmarks it sits below Glm5.3 Flash. Waaaay below Qwen3.8-Flash-Next a model less than half it's size.

No, the only open weight model that really makes sense for me is Qwen3.8-Flash-Next, but it is mainly because I can run it locally with reasonable speed (prefill between 650-1400t/s generation between 22-50t/s depending on number of slots/users I configure).

This is the first model that truly competes with Opus 4.8. I'd say it may be better than Opus 4.6 on programming.

But it is very verbose when it comes to reasoning tokens. The more difficult the task the more verbose it is. Certain very hard tasks that take opus 4.8 400k tokens take Qwen3.8-Flash-Next 2M tokens... But it finishes them.

And what you loose on the generation speed you get back on input caching you can keep on for weeks.

It really depends on the workload.