logoalt Hacker News

postalcodertoday at 4:27 PM8 repliesview on HN

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).

Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"


Replies

mediamantoday at 4:44 PM

This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?

Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.

show 11 replies
fallingbanannatoday at 5:15 PM

Those 27.3% are still in the ballpark of modern models:

- Sonnet 5 - 12.4%

- Luna - 17.3%

- Grok 4.6 - 20.3%

- Sol - 37.3%

- GLM 5.3 - 41.8%

- Opus 5 - 51.8%

show 5 replies
throwatdem12311today at 5:48 PM

This is why I find benchmarks absolutely worthless.

First, almost all models are within spitting distances of eachother.

Second, it never translates to being better for my own workloads.

You just need to make your own benchmarks.

walrus01today at 6:04 PM

For comparison Qwen 3.8-Flash-Next which runs in under 190GB of RAM locally scores 25.3% on terminalbench 4.0.

Readeriumtoday at 6:19 PM

Yeah DeepSeek V4.1 beats this by 15 percent (4 points) on terminal bench 4.0

eranationtoday at 5:02 PM

When a benchmark becomes a target, it's no longer a good benchmark...

show 1 reply
thefourthchimetoday at 6:38 PM

Came here to say the exact same thing! People have to stop paying any attention to coding benchmarks that aren't Terminal Bench 4.

I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now!

Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.

show 1 reply
enraged_cameltoday at 4:42 PM

Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.

show 2 replies