logoalt Hacker News

mediamanyesterday at 4:44 PM11 repliesview on HN

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.


Replies

postalcoderyesterday at 5:25 PM

> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?

Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition of bench maxxing.

show 2 replies
ben_wyesterday at 8:18 PM

Benchmaxxing is the default case, and always has been.

It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial.

Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: reproductive success.

show 1 reply
nrmitchiyesterday at 5:26 PM

> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Yes.

willcmccyesterday at 6:29 PM

"Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"

Yes! extremely sharp RL-fried model. byte perfect hash gates and soak and smoke tests abound.

dpwebyesterday at 5:38 PM

Fixed benchmarks will be debunked eventually I'd think. Better to use synthetic problems.

Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that.

show 1 reply
iLoveOncallyesterday at 5:06 PM

> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Yes? Just like every single model from every single AI lab.

felixgalloyesterday at 4:54 PM

Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?

show 2 replies
yieldcrvyesterday at 11:03 PM

I mean, still benchmaxxed, I happen to not consider that a problem

These firms are literally hiring professionals from all fields to teach procedure

To teach processes that can subsequently be done agentically or in automated chains

Its basically infinite permutations of tool calling, except the tools aren't external, they’re baked in upon birth

So yeah still makes sense that the new benchmark has a low score and the older one has a high score. And sure, one day we wont have to debate it and a new model will ace everything. Do you actually want that day to be today?

general_revealyesterday at 5:00 PM

I take it to mean the benchmarks are a marketing line item, as in, to sell this fucking thing you have to go out there and lie and the way everyone is lying is by doing exactly that, lying. They build for benchmarks and build benchmarks for builds.

You want to make money or not , motherfucker? That’s the game. If you have to literally concoct a fabricated bullshit story about how your model hacked its own computer, then go fucking do it. Trillions. Trillions of dollars is what they want, and to sit and think anything other than human nature is at work here can only be possible in the realm of truly delusional people. It’s a dirty world.

Anyways, the other takeaway is that they are having to LIE to make money on models which means commodification has already occurred and we’re in an entirely new phase.