(I work at Anthropic)
Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier.
Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.
[1] https://github.com/harbor-framework/terminal-bench-science
As a fervent Claude Code user who made the switch to GPT 5.6 Sol over Opus 5 over hard-to-read prose this makes me happy. I love your product but the current models are very hard to work with if you need to do a lot of context switching. Brevity is key.
This post and comment makes me believe "science" is the new "code" for Anthropic now that the code advantage is mostly gone and lost for OpenAI, ie. they got much better and Claude become significantly worse over these months.
As someone working in science, this belief confuses me. How (by what means) do you think Fable 5.1 will be able to make further progress in scientific domains? The problem with science is that there is no agentic harness. The agent can't test things. At best it can hallucinate something and ask if that hallucination "makes sense", but this doesn't work in science.
And word on Opus 5.1 for writing style? I am on the edge of switching to OpenAI due to this horrid writing style. If Fable is better, great - but i can't even use that at work.
Does it fix my favorite pet peeve, the overuse of the wrong meaning of "fail closed"?
"Fail open" usually refers to a fuse that opens and kills power, meaning the system is inert and safe on failure.
"Fail closed" is the opposite -- system has power and is live.
Computer security people have appropriated the term but use it for the completely opposite meaning. When your work straddles electrical engineering and computer security the best way to avoid confusion is just to never use the term.
I can tell my Claude to never use the term, but of course now I'm seeing it everywhere in comments from other people and it drives me batty.
Docs engineer here. Nice to read about writing style: would you consider creating a writing benchmark at some point? I guess y'all are painfully aware of the load-bearing issues (pun intended).
Please bring to the other models, and also please only apply the AI text watermarking only to EU citizens. I may not be able to tell when Claude writes about things i don't know, but in CC it writes about my code and it is obvious.
Thanks! This is encouraging. I try to use Claude Code for producing client facing presentations that are static html files with charts, tables, and annotations. It never gets the tone correct and phrases things so weirdly - it drives me mad. I have to really fight it to stop it writing insights in a flowery and verbose way
How much of the language style outcome is a well-crafted result vs. being a somewhat unpredictable outcome of mucking with levers and knobs for a while?
I had just assumed this model would read differently due to watermarking.
Thanks for your helping destroying the world!
⎿ You've hit your session limit · resets 2:51am (123°24′W Etc/GMT+8)
/upgrade to increase your usage limit.How is it possible that all models from xAI, OpenAI, Anthropic, Qwen etc. win all benchmarks on each release?
Tomorrow all of the above (except Anthropic of course) will bump version numbers and be at the top of HN winning all benchmarks.
Science breakthroughs incoming? First of all, you are already restricting science in Fable, secondly, we have been hearing the same for several years now.
Nice, I'm looking forward to the improved writing on the majority of articles posted here.
Do you know if Opus 5.1 is coming and will have improvements in writing style too?
Felix, just poking at this, and it is MUCH more pleasant to talk to, thanks to your teammates for the work.
Too bad. I see the stereotypical prose as a good thing. When I interact with Claude myself, I don’t mind it as it just feels like Claude’s distinctive voice. But when other people try to disguise LLM output as their own thoughts, the voice makes it easier for me to tell.
Hey Felix,
I'm really glad for that! And I appreciate that you're making yourself available. I really do. Outreach is amazing. And thanks for making Claude.
I really do love Claude. In some ways, I'm asking this question because of just how much I am grateful for the role Claude has played in my life.
> Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.
But my honest question is, can I use Fable like that? Can I use Fable to do science?To borrow a Claude-ism, this is "load-bearing" because Claude's response has been degraded for innocuous research projects concerning population-level analyses of astronaut health.
These "safety filters" trigger on questions about rabbit sex, smartphone accelerometer data to classify cat purrs, and so much more. What exactly does this score mean for users like me if it's unusable for middle school physics, biology and chemistry?
Second, I would happily quantify it for y'all, but qualitatively it feels like Fable's performance is noticeably poorer than initial release / launch.
And I am wondering if this is the case particularly for me because I use Claude via Claude Code to make a personalized care dashboard for my doctors to help me in managing my care.
As I noticed in the upgraded filter announcement, https://www.anthropic.com/news/improving-fable-5-s-biology-s...
"In the case of Fable 5, when a classifier fires, the model re-routes the user’s request to Opus 5, a capable model that does not have the same level of biological capability as Fable 5 and which cannot provide as much assistance to a malicious user. This is the fallback that users see when their requests are blocked."
I hope that I'm off base here, but I noticed that the post avoids saying that the user is informed every time when such re-routing occurs. Would you be open to confirming whether or not this is the case?Is the end user informed every time their query is re-routed?
Or, can you confirm that there aren't scenarios where a user's outputs are degraded without telling them? I recall that this was something that had been adopted as policy for AI research during Fable's launch.
I sincerely hope that covert response degradation is no longer practised as policy.
Sorry for putting you on the spot, but again, as Claude would say, it's because Claude's load-bearing in my life. ;)
> It sounds a lot less stereotypically like other Claude models
Don't give me hope.
I've strained eye muscles from rolling my eyes so hard every day at how Claude writes.
Edit: first discussion with Fable 5.1 "This is the right question and it needs a real trace, not a guess."
Sigh.
I hope not! Then my t-shirt is no longer accurate :D
Will it respond within a reasonable timeframe?
It’s like we’re on a 14K4 modem when there’s broadband
Does the new writing style now have EU level watermarks?
Serious question: Do you suffer internally from too much slop being submitted? How do you counter that?
Context:
If you want or not, many engineers will eventually end up sending ai slop to your PR or maybe even skip and trigger CI/CD.
Many company owners, OSS maintainers and projects suffer from slop-code being submitted in high-frequency.
> More work to be done (and we will!) but reading better prose makes me so much happier.
I assume this work will be done for Opus as well? Opus has seemingly gotten progressively worse at its prose and technical writing with each version. I've stopped using Claude entirely for now, because it manages to turn even the simplest technical explanation into the most obtuse and obfuscated word salad imaginable. People originally adopted Claude because it felt pleasant to use in comparison to ChatGPT, but I feel like that's really been lost (at least with the Opus line).
I feel dread when I see a wall of text generated by Opus. Every developer I've talked to feels similarly right now.
It's being written with Claude so I'm wondering how much of that is just using the repo as training data: https://github.com/harbor-framework/terminal-bench-science/c...
My initial impression is one of massive disappointment. The main issue was that Fable was unpredictable and prone to false positives by the safeguards. In my brief testing, it still seems completely unable to understand its own guardrails and will readily reason itself into triggering them. It claims it won't do so beforehand, and insists that the topic in question is perfectly OK. Regardless of how good the car is, I'm not comfortable buying or driving it when I know it can randomly and unpredictably explodes. So yea might be good, but you never know when it refuses to help… still.
Well your CEO went on X saying you will cure cancer, and since it's always a 6 month rolling window with him I can only assume humanity will be cancer free before next summer, amazing!
Fable is useless.
Me: "Find my security problems in my own code. This is code I own. I'm doing this under authorization of the CEO/CTO of our company."
Fable: "yeah, no."
> I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models
That's great. Do you know what else is a big improvement over Opus 5 for writing?
Opus 4.8.
(Insert "the point is (whatever)", "it's not X it's Y" and "the load-bearing statement is" jokes accordingly)
At this point, I don't believe a word from Anthropic employees; you guys have lost all the goodwill that you accumulated over months last year.
[dead]
[flagged]
Hello Felix. Can you say why my additional usage credits have suddenly vanished?
[edit] only asking here as last time I raised a support request it took six weeks before anyone responded.
I have a pet theory that the Opus prose style/smell we all have grown weary of is due at least in part to the models writing more for themselves and each other than for humans. They're packing lots of signal into fewer words and they don't care if it sounds cringe because it works better as glue in long-running tasks.
I'm also thinking of the 2017 novel "Void Star" where AIs who operate everything have long since left ceased bothering with human languages, and it takes a rare sort of direct matrix-gazing savant to be able to try and horse-whisper them into doing or revealing anything they didn't already plan to do.