What's new in Claude Fable 5.1 – https://platform.claude.com/docs/en/models/fable-5-1/whats-n...
System Card: https://www-cdn.anthropic.com/0339e6a7c5c7b87f5c07798616dc32...
“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.”
I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better.
I don’t think Anthropic realizes that humans have a token limit too and it can be exhausting to read Claude’s output. Prose density is not the same thing as succinctness.
The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M).
This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general.
Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is hard to see ANY improvement:
Terminal-Bench 4.0: Fable 5.1 is +3.5% vs Opus 5.
GDPval-AA v2: +1.5% vs Opus 5.
OSWorld 2.0: +2.5% vs Opus 5.
Humanity's Last Exam (with tools): +1.6%
Keep in mind that this is supposed to be an entirely higher tier of a model than Opus 5. For one tier up and one version up, these are not really improvements. Probably leaves no room to place Opus 5.1 anywhere. Combined with the fact that they are selling 'readability'... Has frontier progress finally stalled?
Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
I'm still waiting for effort max to finish.
EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30!
Excerpts from the reasoning trace:
> Adding pedal shapes near both feet, with the far foot on the second leg partially visible behind the frame. I'm considering whether to add a small scarf or cap for extra character, but leaning toward keeping it simple to avoid clutter.
> Now I'm debating a bicycle helmet on the head versus the pelican's signature crest—the beak and pouch already read clearly as "pelican," so a helmet could reinforce the bicycle theme without losing identity, though it might compete with the crest for visual space.
> I realize the beak at (484,84) would overlap with the dome helmet, so I need to shrink the helmet so it only covers the top of the head, adjusting its arc endpoints to sit higher and narrower so the beak can attach cleanly at the front without collision. [...]
> I'm adding a darker tip region to represent the primary feathers, then reconsidering the trailing edge to include scalloped feather curves instead of one smooth line for a more natural look. [...]
> Now I'm checking the vent line placements on the helmet, making sure they sit far enough inside the helmet's edge given the stroke width and rounded caps, and confirming each vent stays within the helmet's circular boundary. [...]
> I decide skipping a handlebar bell and tire highlights since they're unnecessary additions. Now I'm reconsidering the front fork's curve — the current control point pulls the shape backward when it should bow forward for a proper rake, so I need to shift the control point rightward to fix the fork's lean.
This is a notable result because most of the recent Claude models have been pretty bad at drawing pelicans, at least when compared to models in the Gemini or GLM series.
Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from other models verbatim (since it can see the decrypted version). I get that in their eyes it's an "exploit" but still kinda disappointing that they patched this
Going to hold off a few days until I adopt it, lets see what the general consensus develops as. Regretted jumping over day one for 5.0.
"Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%."
Glad to see this!
> We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.
I'm not an emdash hater but this isn't how you use them. It should be a comma.
Let me guess: it's the end of the world again. These new models are sooo powerful that will take over the world, just like the others before them.
Are they going to try the banned for export for a week marketing move too?
> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying.
Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.
To be honest, these frontier model releases have become boring for me. Opus 4.8 was already good enough for most of my use cases. I don't have any projects right now that I would use Fable for instead of Opus. So when I see announcements like this I just think "that's cool I guess" and then go back to using weaker/cheaper models.
What's far more exciting right now is models like DeepSeek V4 Flash and GLM 5.3 Flash. They have achieved good-enough-intelligence at extremely low prices and fast speeds. I don't have a use for Fable-level intelligence, but I do have uses for Opus-4.8-level intelligence that I can use as much as I want without worrying about the bill.
The breaking API changes are frustrating, especially the one that removes forced tool use.
I've recently been running these agent sessions on more and more long running tasks because these latest models can do a REALLY good job on big chunks of work, and i've been watching them way less. It's starting to occur to me the importance of alignment is a today problem, it's not a tomorrow problem.
In the past I watched and saw everything the model did, not a lot got past me. Today it does A TON of work while i'm busy on other tasks. It also has extensive access to my computer, other computers on my network, my internet. It's really helpful when you give it a lot of resources, but right now I have very autonomous, very smart agent running around more or less unattended with a lot of resources.
Instead of a new model that's going to have unreasonably shallow usage limits, I wish they would:
1) address the claude 20x plan usage being only 6-7x the ceiling of the claude pro plan
2) either fix opus 5, make it completely free, or delete it entirely
Cool. I’ve realized though that I don’t really need better models anymore. SOTA is good, I just want them faster/cheaper now.
Data retention still sounds bad: "Claude Fable 5.1 and Claude Mythos 5.1 carry 30-day data retention and aren't available under zero data retention unless expressly authorized by Anthropic."
Anyone know who the ZDR special treatment is available to?
All the benchmarks in the world don't matter if the model just straight up refuses to do mundane things. Claude has too much of an attitude.
Not until you stop being cheap and let pro users use fable under their existing paid subscriptions.
I've been building Cargo-for-C (https://github.com/tspader/spn), and the difference between Fable and Opus was already astounding. Fable was the first time that I could point a model at a piece of code I'd written and expect it to make it meaningfully better rather than a hard pattern match to whatever mistakes it had.
5.1 so far seems like another leap, which is really surprising. I threw it at a few bigger features I've been designing for a while, and it came back with some extremely thoughtful wrinkles in the design that I'd legitimately not considered. Which, OK, package managers and build executors and compiling C/C++ is pretty well trodden ground, but my thing is very different from everything that exists, and I was very surprised it was able to understand all that context so deeply and intuitively
"Cache reads now cost 75% less, or $0.25 per million tokens." For me, at a typical 95% cache hit rate, I think my optimal context window size before autocompaction goes from ~200K to ~400K tokens. Great for longer horizon tasks.
On both my work (Team Premium) and personal accounts (Max 20x), Fable 5.1 hit the 5-hour limit before it could finish the first task I gave it. On my work account, it took about 30 minutes, and on my personal account, less than an hour.
This has never happened to me before, but if this is normal behavior, Fable 5.1 is essentially unusable.
Fable 5.1 is actually more expensive than 5.0 when run on the Artificial Analysis suite:
I cancelled my pro max 20x subscription, tired of Opus stopping the work from time to time, or saying "this is 2 months of work"
Tbh with that price , not even willing to try . What are the benefits for a regular coding agent ? I barely have any errors already with 4.8 level , eg grok 4.6 , gpt 5.6 sol/terra behind router . Why do I need to pay so much money for this ? Any reason ?
Bit of a discount if you're using caching:
> same input and output prices, with cache reads at a quarter of the cost
This should impact any long-running agent since subsequent calls can benefit from cached reads for previous transcripts.
According to the FrontierCode Extended benchmarks in the system "card" (page 169-170), Fable 5.1 apparently does best on the medium effort level for this benchmark: "[...] at higher efforts, Fable 5.1 occasionally adds more small, unrequested changes [...]" Though Fable 5.1's medium is also lower than Fable 5's best score on the same benchmark, which uses xhigh.
"Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations"
Then why does it have separate datapoints for Terminal Bench, and score higher? Something doesn't add up here??
The thing with Fable-level models is that I will never feel comfortable using them for agentic tasks on a pay-as-you-go API pricing plan without monitoring them strictly, which becomes a chore.
I once caught Fable 5 spinning its wheels on a rendering issue, which evaporated 90% of my usage in a single prompt. I could never let Fable run free attached to a credit card without staring at it the whole time.
> This required us to add a watermark—a numerical way of determining the likelihood that Claude was involved in writing a piece of text—to the outputs of models released after August 2, 2026. As we recently explained, this watermark is invisible to anyone who does not have the detection API. It has no practical impact on the quality or content of Claude’s outputs and contains no information about the user, their organization, or their conversations with Claude.
How does this work if it doesn’t change the output?
I’m really excited to try this out. Fable and Opus 5 constantly wow me when working together. Unfortunately, I’m a little burned because of technical issues.
Anthropic accidentally over-billed my account, and when I reached out to the support bot, it downgraded my account to a Free account. It’s been impossible to get it resolved and I have almost $200 held hostage.
I don’t want to do a charge back. I’m one of the main advocates for Claude Code at work, I use this subscription to try out new features before it’s available at work.
The whole experience has been illuminating about our dependencies on these AI companies.
Sadly still not available for Pro subscription. At least they reset everyone's limits.
Why aren’t these models available on subscription plans?
I tried the old fable and it didn’t seem worth paying for. It still made errors like Opus does so I might as well use the included model…
Anthropic, the company employing "treat them mean, keep them keen" as a marketing tactic. Pass.
I'm so suspicious of this after Opus 5 benchmarks scored it higher than Fable 5, yet Opus 5 was untrustworthy (overconfident, error-prone).
There's now a 40X discount in the cache input pricing instead of 10X.
This seems to point to them having achieved some kind of optimization in attention mechanism perhaps along the lines of DeepSeek V4, which had a similarly high discount between cache input and normal input.
In real world use, the savings should be quite noticeable. For example, you can now use the model at 800K tokens context window at the same cost efficiency as the previous model at 200K tokens context window.
The most remarkable thing here is just how close Opus 5 is on most of these benchmarks.
Somewhat ironically, Fable 5.1 was flagged by the biology safeguards after I asked it to have a dig around the Fable 5.1 system card :)
> Data retention. Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention.
This is interesting. I wonder if customers will be allowed to create an auto expiry for their own data to prevent future subpoenas. That’d be a treasure trove for discovery.
Looks like the API is nerfed to mitigate some recent thinking extraction attacks.
I wonder to what extent this will make the automatic Fable-to-Opus downgrade give worse results.
I'm having a very hard time finding mention of token-generation speed.
Unless these people start offering free, unlimited inference for a cautionary period so we can test the new model without an up-front (re-)investment, I am not touching this load-bearing pile of neuralese spew with a ten thousand token pole.-
Am I alone in not prioritizing the quality of prose produced by my coding agent? My foremost and almost only concern is how well it can engineer software.
I use Claude Design heavily, I wish these charts show "10% better at picking a color" or laying out an app. Maybe it's hard to build a good visual design test. Claude's good at layouts but not the colors or smaller design details.
I don't know how I feel when all the documentations are written by AI for humans.
AI to AI doc share: sure, do what you please.
AI to human: please make it legible and flowly.
example, "Every thinking block records which model produced it, and it's preserved in one direction only: Claude Fable 5.1 reads earlier models' thinking blocks, and no earlier model reads Claude Fable 5.1's." is a very Claude-isk way of writing. Choppy, long, and lacking flow.
This coupled with verification primitives will be quite compelling. we really have to start reimagining existing systems and processes from the ground up.
"with cache reads at a quarter of the cost"
OK, I think that's what they meant when they suggested reduced extra promo usage will not sting this much.
(I work at Anthropic)
Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier.
Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.
[1] https://github.com/harbor-framework/terminal-bench-science