logoalt Hacker News

ckocagiltoday at 5:32 PM2 repliesview on HN

Because it's notoriously hard to benchmark LLMs. Ultimately every benchmark is different and measures different things. This is why companies that make LLM models have private benchmarks - they find the areas where the model is weak and make that their goal.


Replies

hellohello2today at 7:26 PM

Yes, I agree, this is why this post is interesting despite being clickbait. You get what you measure but its better than being blind etc.

irishcoffeetoday at 6:03 PM

The whole concept is kind of silly. We don’t “benchmark” humans. Or do we, via standardized tests? Why don’t we just use those? Or is that what the benchmarks are? I have no idea.