logoalt Hacker News

sunbumtoday at 10:56 AM2 repliesview on HN

the 2nd website is not official, just something someone slopped together for some reason.


Replies

Alifatisktoday at 11:11 AM

I have plenty of these websites, I can’t understand why someone is doing this.

show 1 reply
yorwbatoday at 11:26 AM

Even if it weren't slopped together, 65% vs 80% on 10 tasks just isn't a significant difference. For 80% power to distinguish at a significance level of 0.05, you'd need more like 140 samples, if those were the true success probabilities.

The number one problem in LLM benchmarking is that people try to draw conclusions from sample sizes far too small to conclude anything but "it works sometimes, it fails sometimes, hard to say which is better." (The number two problem is that people run benchmarks blindly without checking that they measure something meaningful.)