logoalt Hacker News

traceroute66yesterday at 9:34 PM4 repliesview on HN

So TL;DR benchmarking in a completely non-reproducible manner ?

"Model X performed great, but we can't possibly tell you anything about the code it was looking at apart from it was a large code base from an unknown company".

So basically pinky-promise benchmarking ?

I'm not sure I follow the value here ?


Replies

kadobanyesterday at 9:43 PM

If it builds up history and perceived reliability, this type of thing can be valuable. You're giving up transparency for it being harder to game.

show 1 reply
sigmaryesterday at 11:04 PM

Lots of private benchmarks already exist, where you have to trust the tester (ex Artificial Analysis, Arc-agi).

deepwoodsyesterday at 10:23 PM

In theory, as long as all the models are doing the same thing with the same tools, it's at least useful to see how they stack up against each other right now. It might not be great to track progress over time, as it can get benchmaxxed or the underlying resources may become obsolete.

demibabsyesterday at 9:47 PM

Doesn’t it ultimately have to be this way, to prevent saturation?