logoalt Hacker News

pookieinctoday at 4:36 PM2 repliesview on HN

“It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”

They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?


Replies

randomblock1today at 4:45 PM

> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

jbellistoday at 4:43 PM

Anthropic knows that the benchmarks showing Opus 5 better than Fable 5.1 are measuring something that's less than entirely useful.

show 1 reply