If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).
Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
Those 27.3% are still in the ballpark of modern models:
- Sonnet 5 - 12.4%
- Luna - 17.3%
- Grok 4.6 - 20.3%
- Sol - 37.3%
- GLM 5.3 - 41.8%
- Opus 5 - 51.8%
This is why I find benchmarks absolutely worthless.
First, almost all models are within spitting distances of eachother.
Second, it never translates to being better for my own workloads.
You just need to make your own benchmarks.
For comparison Qwen 3.8-Flash-Next which runs in under 190GB of RAM locally scores 25.3% on terminalbench 4.0.
Yeah DeepSeek V4.1 beats this by 15 percent (4 points) on terminal bench 4.0
When a benchmark becomes a target, it's no longer a good benchmark...
Came here to say the exact same thing! People have to stop paying any attention to coding benchmarks that aren't Terminal Bench 4.
I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now!
Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.
Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?
Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.