logoalt Hacker News

ramish94 • today at 5:59 PM • 3 replies • view on HN

In terms of benchmarks for agentic coding, it basically stacks up nearly 1:1 with Opus 5.5.

Terminal-Bench: 70.6 (Sonnet 5.5) vs. 66.4% (Opus 5.5)

FrontierCode: 52.1% (Sonnet 5.5 xHigh) vs. 54.4 (Opus 5.5)

CursorBench: 55.5% (Sonnet 5.5) vs. 57.8 (Opus 5.5)

Opus 5.5 might be the best model I've ever used and Sonnet 5.5 matches it and exceeds in some benchmarks. Clearly Anthropic have had some sort of breakthrough with not just performance but also cost with the 5.5 family


Replies

level87 • today at 6:12 PM

This is crazy, what is the point of all these equivalent models?

➕ show 2 replies
bbor • today at 6:21 PM

Yup. Recursive self improvement presented in hard numbers.

bigyabai • today at 6:01 PM

It's long overdue. Sonnet 5 was terrible API value for agentic coding, there were open models like GLM-5.3 Flash that blew it out of the water at 1/20th of the price.

OpenAI and Anthropic's lead is vanishingly small at this point.

➕ show 3 replies