logoalt Hacker News

otterdudetoday at 5:22 PM2 repliesview on HN

Benchmarks saturate around 80-90%?

This is not "Acing" a test, this is hitting a wall.


Replies

scotty79today at 5:29 PM

Even on very small tests a fraction of questions might have wrong answers in the key.

If models can't get more than 90% of the benchmark right I think it's a strong indication that they were not trained on the answers and that benchmark itself is messy enough that <10% desired answers might be wrong or misleading.

show 1 reply