logoalt Hacker News

seizethecheesetoday at 6:18 PM1 replyview on HN

These results don’t just contradict more serious benchmarks, they are wrong on an entirely different axis. This is a saturated benchmark. Haiku gets 96%. The results here are “not even wrong” and this being #1 on HN right now is a massive smell of either bots or massive ignorance or both.


Replies

uramstoday at 6:31 PM

People REALLY want the open models to be better than the frontier labs'.

show 1 reply