logoalt Hacker News

jchwtoday at 4:56 PM4 repliesview on HN

Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?

I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM.

"Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5...

This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.


Replies

sambusa_123today at 6:05 PM

Just read through some of the code "benchmarks", and I see why: https://reinvently.co.uk/tools/ed-o-meter/tests/

Most are extremely trivial tasks. I would be surprised if a model from 2 years ago failed these...

show 1 reply
rfgplktoday at 5:25 PM

> Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?

The only true benchmark for any of these models that I've discovered isn't if they can pass precanned SWE tests, but rather can they create something novel? This isn't even too difficult to test, just give it a seemingly impossible task let it spin and see where it ends up.

finaardtoday at 5:44 PM

Interesting - when Kimi K2.6 came out I switched over from Anthropic models, with at that time comparable to better results for me. I was using Anthropic via API, heavier months were roughly $400 worth of Anthropic tokens - I can get the same thing done via a $100 ollama subscription.

zuzululutoday at 7:18 PM

yeah that haiku really diminishes the claims behind the benchmarks. luna-max is significantly cheap and it is a strong performer but i dont see it on the benchmarks.

I just get a feeling this site started with an intent to elevate Chinese models above the rest so wouldn't be surprised if the whole prompt sail was set to that tune