logoalt Hacker News

Aissen • today at 9:58 AM • 0 replies • view on HN

> test several levels against your own evals

Of course, and this is the basics anyone should do when working with LLMs & agents; but with their high-variance, doing statistically significant benchmarking is very costly. Which is why the debates here on HN often talk about the "feelings" of degradation (or improvement!), but often without proofs. I'm not sure how to solve ạt; maybe inference providers should provide free benchmarking to anyone publishing results, along with the guarantee to never train on those sessions.