logoalt Hacker News

gertlabstoday at 5:13 PM2 repliesview on HN

I started thinking about this back after the Llama 4 release, and since then our team has put a lot of thought into designing evaluations that don't saturate, are resistant to contamination, and can scale. What has worked best for us is using multi-agent environments with open-ended cooperative or competitive goals. Mostly designed as multiplayer games. The results tend to align with our experience for coding better than any non-aggregator benchmark, and likely at lower cost to run.

Data at https://gertlabs.com/rankings


Replies

erikwiffintoday at 6:15 PM

I've developed a benchmark that I think should be resistant to saturation, is easily verifiable, and anecdotally correlates with desirable behavior (ability to not get confused while generating text with state).

I think it's interesting, I think other people would find it useful, but I don't want to spend a bunch of money running it against all the frontier models.

What's the best way to reach out to labs like yours to collaborate on something like that? Are there any labs that are more open to submissions from internet randos?

nwienerttoday at 5:25 PM

If you're ranking Opus > Fable you're ranking "do [clearly defined thing with easy to grade endpoint]" too much. Real world doesn't value that nearly as much and it's why benchmarks are maxxed.

show 1 reply