logoalt Hacker News

deauxyesterday at 6:22 PM1 replyview on HN

"Intelligence" being what, math? Coding? Unfortunately there's a billion use cases for LLMs whose performance is not at all captured by the popular benchmarks they're all trying to maxx.


Replies

whimsicalismyesterday at 6:33 PM

if you are relying on a model for a business process, it should be simple enough to benchmark on that process