logoalt Hacker News

Claude Opus 5

1736 pointsby alvislast Friday at 4:57 PM1259 commentsview on HN

https://www.anthropic.com/claude-opus-5-system-card


Comments

tyrelast Friday at 5:32 PM

I'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.

show 1 reply
boclast Friday at 5:43 PM

Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:

"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.

What I can tell you is what I actually observe:"

I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.

One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!

show 2 replies
arrowleaflast Friday at 5:27 PM

I can't find anything about whether this is zero data retention, or falls under their required 30 day retention like Fable and Mythos?

bovermyerlast Friday at 5:48 PM

This stood out to me as a little concerning:

> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.

show 2 replies
doginasuitlast Friday at 11:21 PM

Any observations on Opus 5 personality quirks? I had to skip 4.8 entirely because it has zero chill.

himata4113last Friday at 5:06 PM

Rather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.

show 2 replies
yusufozkanlast Friday at 5:02 PM

> arc-agi-3 30.2%

wow

MasterScratlast Friday at 9:17 PM

Damn the pelican guy can’t get no sleep

shockembopperlast Friday at 5:56 PM

I wish these releases came out earlier in the day so I could try them during my work day instead of waiting until the next.

arjielast Friday at 6:48 PM

I wonder when a model will be released that can work in a loop and port Qwen-3.6 27B to run on Tenstorrent P150.

whatever1last Friday at 5:22 PM

Where does this leave Fable? I am confused.

show 2 replies
8notelast Friday at 5:42 PM

im excited that cad and object=>cad is getting into the test tasks

i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?

itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.

the automated test harness for physical stuff seems a bit beyond reach still

born-jrelast Friday at 7:07 PM

Is it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .

show 1 reply
theplumberlast Friday at 6:51 PM

The most important thing is it has the same drama queen mode on safety “guards” like Fable.

uramslast Friday at 5:25 PM

So Opus 5 is basically "distilled" Fable? The benchmarks look often better than Fable.

mcastlast Friday at 5:04 PM

Interesting timing to release this on the same day Jensen makes a statement on open source AI.

show 1 reply
b-sideyesterday at 12:00 AM

Significantly worse than it predecessors it will now just refuse to acknowledge when it is wrong (which would be less of an issue if it wasn’t getting basic things wrong) also the “personality” when pushed back on obvious mistakes is unbearable.

korabslast Friday at 6:20 PM

So in benchmarks it's better than Fable?

But they say it's "almost as good as fable"

doctobogganlast Friday at 5:44 PM

According to these charts I should switch from Fable to Opus in Claude Code now?

inshardlast Friday at 5:16 PM

Arc AGI score is astounding

spstoyanovlast Friday at 5:14 PM

So same as Sol? I guess I’ll see which one is more token efficient.

vinishkapooryesterday at 5:46 AM

Tried and had great experience.

mkurzlast Friday at 5:15 PM

Where is the pelican?

show 1 reply
taf2last Friday at 5:10 PM

eager to see how it benchmarks on https://deepswe.datacurve.ai/

Eldodilast Friday at 5:08 PM

Models benchmarks start to get saturated again!

hahahaayesterday at 12:13 AM

I sense a bird on a bike coming.

arjlast Friday at 6:25 PM

On a Friday, I'm out of tokens ;-)

toephu2last Friday at 6:06 PM

How does it score on DeepSWE?

show 1 reply
firemeltyesterday at 1:12 AM

so what is the default effort for this model?

internet2000last Friday at 6:30 PM

Kimi K3 already left behind in the dust. They can't keep getting away with it!!!

nstjyesterday at 7:22 AM

came here for the pelican

hmontazerilast Friday at 5:30 PM

Honestly if reached a level of coding that sonnet 5 is more than enough for my needs as assistant/agent I don’t need long Horizon stuff…

abc42last Friday at 6:18 PM

Are we getting to singularity or something? This seems a bit crazy.

mihaulast Friday at 5:29 PM

30% on ARC-AGI-3

show 1 reply
Uptrendayesterday at 12:15 AM

Is this thing also going to try hack us?

_pdp_last Friday at 6:55 PM

Wake me when they deliver Opus 4.8 level performance for $5 per million tokens.

show 1 reply
arseniitrutlast Friday at 7:48 PM

atp, is it the end of fable 5 era?

throwaw12last Friday at 5:12 PM

is coding and engineering solved yet?

show 1 reply
tomlockwoodlast Friday at 11:25 PM

This stuff is a commodity and China seems to be the only one that's noticed.

shinhyeoklast Friday at 10:33 PM

I love it

ismailmajlast Friday at 6:39 PM

I'd pay good money to see OpenAI "oh fuck" war rooms.

holodukelast Friday at 8:47 PM

Is it me that the model performance between 4.7 and others is really small. For me even 4.7 works fine. Sure fable might be a bit better. But is it really noticable? It's in the same league if you ask me.

StrauXXlast Friday at 5:17 PM

The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.

show 1 reply
jakeoghlast Friday at 7:54 PM

Anyone else not getting chain of thought? Opus 4.8 would show it to me, until around the time Fable came back. Now I dont see it with 4.8/5.0 or Fable. Not having it makes catching mistakes harder.

LoganDarklast Friday at 6:15 PM

These cybersecurity safeguards are really annoying. There are ethical reasons to reverse-engineer and binary-patch software; for example Rewind got acquired by facebook and, as a gift to all their customers, implemented a killswitch in their software to ensure it will eventually stop functioning. I kept using a version without the killswitch, but the macOS 27 update killed it, and I needed binary patching to fix it. I should be allowed to repair software I purchased (I did purchase it like a month before they sold out), but unfortunately this overlaps significantly with cybersecurity.

show 2 replies

🔗 View 40 more comments