“It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”
They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?
Anthropic knows that the benchmarks showing Opus 5 better than Fable 5.1 are measuring something that's less than entirely useful.
> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.