logoalt Hacker News

rfgplktoday at 6:18 PM1 replyview on HN

Opus 5 is better than Fable 5 except for creative programming work (like graphics). Fable 5 might be slightly better but the token cost isn't worth it.


Replies

bredrentoday at 6:27 PM

How are you evaluating the models?

On the Fable 5.2 eval summary, Opus 5 only beats Fable on SWE-bench multilingual and multimodal.

I primarily use the models via interactive sessions enhanced with custom tools and skill. For that Opus 5's benchmark superiority has not materialized into greater productivity and frankly has been quite a let down.

The outputs are too often unreadable even after adding recommended prompts. There is an ongoing problem with the heron_brook system prompt affecting orchestration. [1]

I've used Opus 4.8 since the second week Opus 5 was released.

Over this time, Fable 5 has been reliably fantastic. Both in planning and direct execution on complex changes across code and infra.

I'm a bit surprised that there doesn't (seem) to be a section discussing ~performance across different modalities. This system card and blog post too-often default to an API-based use case when the gander primarily experience Anthropic's models via interactive sessions.

I understand waiting to comment until Opus 5.1 is available and handles these problems, though I am hopeful that Anthropic will confront the elephant in the room on Opus 5's failure to delivery great interactive sessions and the widespread negative feedback on the release.

It would show the org is paying attention, taking steps to balance model evals between interactive and API use. Also, some empathy for customers that wasted time trying to make opus 5 work for them.

[1] https://github.com/anthropics/claude-code/issues/80988