logoalt Hacker News

sottoltoday at 6:17 PM5 repliesview on HN

A lot of the benchmarks seem often near meaningless these days - really bench-maxxed to the hilt. I tend to still look at the Artificial Analysis rankings to get at least an idea on relative performance of models, is that still warranted?

What or other opinions on how representative the AA rankings are of real-world performance? Any better indicators?


Replies

re5i5tortoday at 6:39 PM

Have you tried it? I’d recommend doing so, it’s impressive in real use cases.

drob518today at 7:42 PM

IMO, ELO rating from The Intelligence company and arena.ai are more representative of rankings since they use humans to judge a head-to-head comparison between a couple models at a time. https://www.intelligence.ai/ http://arena.ai

aqme28today at 6:54 PM

Do you have any evidence that this model is bench-maxxed? I know that's particularly difficult to quantify. If there is an indicator of bench-maxxing, that just becomes the new benchmark to benchmax.

show 2 replies
deauxtoday at 7:00 PM

This one is very benchmaxxed, and you can tell from this page alone. Look at the huge variance in ranking per benchmark. Most models, including at that size, are much more consistent.

show 1 reply
Iolaumtoday at 6:22 PM

Yea and we are reaching the point where this benchmaxing is visible in the model's reported overthinking.

show 1 reply