logoalt Hacker News

okamiueru • today at 4:39 PM • 1 reply • view on HN

Can't you infer the comparisons you would like from baseline results provided? There are better results in the sibling post on HN: https://aleph-alpha.com/en/blog/kolibri-has-landed-a-soverei...


Replies

spijdar • today at 5:09 PM

Maybe?

Qwen3.8 27B scored notably higher in most of the provided benchmarks, including the German-specific ones. The only "downside" is that inference is much more costly and slow, since it's a dense model.

Qwen3.8 Flash-Next appears to usually "benchmark higher" than 27B, while remaining fast.

I'm sure I could dig up the equivalent benchmarks for Flash and do the comparison myself, but as far as inference goes, it's messy. Consider that Qwen3.5 35B-A3B scores higher than Qwen3.6 on some of the German-specific benchmarks.

So it seems superficially plausible that Qwen3.8 Flash-Next might not be "27B but faster" in the ways that are important for this model. Or it could just "be superior" in all ways.

Either way, I don't think an LLM has to be "the best" at anything to be worthwhile, necessarily. And I kind of distrust benchmarks on top of that, so...