logoalt Hacker News

321ahTtoday at 6:29 PM2 repliesview on HN

How is it possible that all models from xAI, OpenAI, Anthropic, Qwen etc. win all benchmarks on each release?

Tomorrow all of the above (except Anthropic of course) will bump version numbers and be at the top of HN winning all benchmarks.

Science breakthroughs incoming? First of all, you are already restricting science in Fable, secondly, we have been hearing the same for several years now.


Replies

pohltoday at 6:37 PM

There are hundreds of benchmarks. You just need to pick a favorable dozen on release day.