logoalt Hacker News

Claude Opus 5.5

642 pointsby km144today at 4:29 PM559 commentsview on HN

Comments

yipinwongtoday at 5:44 PM

I spent about $5 per sentence in my resume using Fable 5.1 (High) to verify accuracy, inconsistency, and edit.

Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.

Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.

show 1 reply
notduckrabbittoday at 4:58 PM

They purport 40% drop in costs due to lower token pricing (presumably aimed at winning back the many of us that switched providers in discovering Opus 5 unusable) and improved token efficiency.

show 1 reply
glubtoday at 4:46 PM

> For users with cybersecurity use cases that may be blocked by our cyber safeguards, we recommend accessing our models with reduced cyber blocking classifiers via our Cyber Verification Program. Claude Opus 5.5 will be available through this program in the near future.

Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.

Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?

louskentoday at 5:34 PM

Cost to Run Artificial Analysis Intelligence Index is higher than previous Opus, so still not cheaper

jacobgoldtoday at 4:41 PM

I use the other 50% of my $200/mo Claude subscription by having Fable run Opus subagents for a lot of work. That way I don't have to deal with Opus directly.

variety8675today at 4:32 PM

I hope this actually fixes the terrible writing style of Opus 5

show 2 replies
34679today at 5:20 PM

I don't care how good their models get, I won't sign up for one of their plans until they define "X" in their pricing. 5X of this plan, 20X of that plan means nothing when they never tell you what "X" is.

Maybe this model can finally figure it out for them.

breezybottomtoday at 5:26 PM

"Where Opus 5.5’s advantage is very clear is efficiency."

Not efficiency in writing, clearly.

Retr0idtoday at 5:38 PM

> Opus 5.5 (1M context)'s safeguards flagged this session. You may be seeing this for the first time on an Opus model: Opus 5.5 (1M context) is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages. Opus 4.8 is answering instead, or you can edit and retry with Opus 5.5 (1M context).

Yay, yet another model I can't use for anything interesting, even with CVP.

Foobar8568today at 5:04 PM

I have just switched to 5.5. First mistake was stale environment variable, didn't realize it was replaced, "oh my memory had stall data" and that's it. Second one, a powershell command had the wrong syntax. Great for my first two prompts.

isodevtoday at 5:18 PM

So is it cheaper? Are we AGI yet? Am I left behind? I didn't have patience for the intro animation on the website... maybe one day, Claude Code will understand accessibility but that day is not today.

show 2 replies
thatxlinertoday at 6:40 PM

So much for pacing the frontier

HarHarVeryFunnytoday at 6:18 PM

METR: Is it safe? Has it escaped confinement?

Ants: It's a good model, sir!

desmondltoday at 5:12 PM

I'll have to try 5.5 on my work's Cursor account. If they really solved the communication issues, I might consider moving my personal account from Codex back to Claude Code.

calibastoday at 4:36 PM

> We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in.

We can't test it properly because it knows it's being tested.

show 1 reply
blfrtoday at 5:18 PM

It's awesome that the apt packages for claude and claude-code are out right now. I can test-drive Opus 5.5 right away. Very cool, Anthropic.

somewhatjustintoday at 4:47 PM

> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

Nice. I was starting to think Haiku was going to be abandoned.

km144today at 4:41 PM

I think this release is really going to give them a hard time selling Fable:

> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.

show 3 replies
aurareturntoday at 5:03 PM

I found myself going back to Fable over and over again. At this point, I’m not sure if I’m just used to its style or it is truly more capable.

I tried Opus 5 and Astra.

__vivektoday at 5:49 PM

I'm only interested in the Opus series, if they fixed the talking issues.

datadrivenangeltoday at 5:11 PM

But have they made it any better at communicating clearly? I cancelled my personal subscription because Opus is so painful to read.

jidaigeisttoday at 4:50 PM

>Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude.

Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.

show 3 replies
alvistoday at 4:38 PM

$0.20 vs the old $0.5 cache read is pretty much 60% off

nickandbrotoday at 4:35 PM

Wow! Though need to see its token efficiency to better assess. Been hearing rumors it generates much more output tokens per task.

show 1 reply
Fizzadartoday at 6:28 PM

So is this AGI+ now?

tag2103today at 4:41 PM

Why would anyone reward bad behavior?

thibrantoday at 4:36 PM

Anthropic models are ridiculously expensive. I've stopped using any of their models months ago.

blurbleblurbletoday at 5:26 PM

Hopefully OpenAI throws us some more usage resets now.

garo-protoday at 5:49 PM

Opus 5.5 is now the recommended model in Claude Code's model picker, which is quite a claim, given how they struggled with capacity.

greenavocadotoday at 4:51 PM

Enjoy it for the next 2 weeks until its silently quanted to 4.8 level

iamsyrtoday at 4:47 PM

I don't yet have any reason to leave Haiku 4.5 and switch to Opus 5.5.

keeebatoday at 4:52 PM

Opus 5.1 came out about a month ago, what gives?

show 3 replies
aennassiritoday at 4:40 PM

Let's see how much they benchmaxxed their model!

woeiruatoday at 5:19 PM

So... why would you use Fable now?

richardjenningstoday at 4:46 PM

My 20x plan was set to end tomorrow. The writing style and insistence on word vomit just became too annoying. Is Opus 5.5 worth sticking around for ?

Yaboodtoday at 4:51 PM

Current models, especially Opus are almost unusable because they don’t respect instructions and their responses are infuriating. They are clearly designed for token consumption. I find myself wasting a lot of time just asking it to shorten or simplify its responses. I’ll give this new model a go, but I’m not holding my breath because the last model release was supposed to fix the very same issues and it didn’t.

firemelttoday at 5:43 PM

wow its really smarter than opus?

Lord_Zerotoday at 4:33 PM

The test they performed to port HAProxy from C to Rust is crazy.

cmrdporcupinetoday at 5:53 PM

GPT Sol 6 has also released today, but no official blog announcement yet

https://www.reddit.com/r/codex/comments/1wnggya/gpt_6_droppe...

🔗 View 33 more comments