logoalt Hacker News

abejora • today at 6:17 PM • 8 replies • view on HN

Sonnet 5.5 scoring higher (70.6) than Opus 5.5 (66.4) in Terminal-Bench is interesting. I looked into this, because it felt strange.

Turns out that Opus had 10% of its trials answered by a fallback model due to safeguards; versus only 1.5% fallbacks for Sonnet. [1] So I would not read too much into this, just the difference in fall backs could probably explain the gap.

[1] Section 8.5 of the Sonnet 5.5 System Card


Replies

eli • today at 6:19 PM

Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails

➕ show 2 replies
subscribed • today at 7:27 PM

I disagree, I think we should read a lot from it, as it stands in this benchmark Opus performs worse than Sonnet, it doesn't really matter why.

Anthropic made it that way, and I'd say the lower score is accurate.

radlad • today at 6:26 PM

I believe you meant to cite the Opus 5.5 System Card which states:

> Claude Opus 5.5 scored 66.36% on Terminal-Bench 4.0 with safeguards enabled; requests flagged by the safeguards were answered by a fallback model following the default server-side fallback policy (2.5% of requests, affecting 10% of trials).

> https://www-cdn.anthropic.com/fc1b44717c85dc068bc6ba50242199...

I cannot find a Sonnet 5.5 system card.

➕ show 1 reply
Leary • today at 6:26 PM

And Sonnet 5.5 is more expensive than Opus 5.5 to hit that score on terminal bench!

oh_no • today at 6:50 PM

it could be that, it could also be that sonnet max looks to burn about 60% more tokens than opus max

AA intelegence index (agent harness doesn't have sonnet data yet) on max: Astra 27k Fable 5.1 78k (Sonnet 5) 118k Opus 5.5 119k Sonnet 5.5 193k

Opus 5 was previous record holder so hats off to Anthropic on blowing it away on token churn.

falcor84 • today at 7:58 PM

So I suppose the easy fix for Anthropic would be to have Opus 5.5 now fall back to Sonnet 5.5, right?

MadameMinty • today at 6:24 PM

That's frankly hilarious. What was the fallback for Opus 5.5? Was it Sonnet 5 or 5.5?

I suppose it also explains how FrontierCode scores seriously dip at Opus/Xhigh and Sonnet/Max?

➕ show 1 reply
manojlds • today at 6:45 PM

Isn't that a worry then that the same bench has so much difference in what triggered fallback for one model and what did not in another?

➕ show 1 reply