> In Anthropic's testing, at its default "medium" effort the model matched or beat Claude Opus 5 at "high" effort on such tasks, in fewer steps and with fewer tokens
Opus 5.5 has been amazing, but I'm confused by how this is worded. It "matched or beat" Opus 5? There is no matching. There is only surpassing. By miles. Like Opus 5 was the biggest disappointment of the year. Opus 5.5 is even better than Fable. I do not understand why they're not acknowledging it for the leap that it is?
Underneath this means that you have say 50 tests and you grade each of them out of 10, then there was no test it did worse on.
The data doesn't support it being better on every test (sometimes the score will be the same imperfect one, sometimes both will have gotten a perfect score).
I don't get it either. Ditto for visual design capability. It's so far above Opus 5 and yet the announcement mentioned nothing about it.