Seems really good so far using it in Claude Code CLI - it gave me a new flag when I asked a question:
"I don't have a reliable way to read that number, so I'd be guessing if I gave you one — and this is exactly the kind of question where a confident guess is worse than none.
What I can tell you is what I actually observe:"
I really like this update - gave me a clear sense of the facts but didn't give me a guess just for the sake of guessing.
One oddity is that it appears to only have a 200K context window right now via CC. Hopefully the 1M version will appear soon!
I can't find anything about whether this is zero data retention, or falls under their required 30 day retention like Fable and Mythos?
This stood out to me as a little concerning:
> The model hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall.
Any observations on Opus 5 personality quirks? I had to skip 4.8 entirely because it has zero chill.
Rather interesting that this makes sonnet 5 look even worse! There is no reason to use sonnet over opus with low or no reasoning at all.
> arc-agi-3 30.2%
wow
Damn the pelican guy can’t get no sleep
I wish these releases came out earlier in the day so I could try them during my work day instead of waiting until the next.
I wonder when a model will be released that can work in a loop and port Qwen-3.6 27B to run on Tenstorrent P150.
im excited that cad and object=>cad is getting into the test tasks
i guess the next stuff will be tool use for the rest of what cad does in assemblies and simulation?
itd be fun to try to set up a 3d printer as part of a feedback loop, and see what a model can build.
the automated test harness for physical stuff seems a bit beyond reach still
Is it me or these have gotten very boring. We have 5 more points on xyzbench or whatever .
The most important thing is it has the same drama queen mode on safety “guards” like Fable.
So Opus 5 is basically "distilled" Fable? The benchmarks look often better than Fable.
Interesting timing to release this on the same day Jensen makes a statement on open source AI.
Significantly worse than it predecessors it will now just refuse to acknowledge when it is wrong (which would be less of an issue if it wasn’t getting basic things wrong) also the “personality” when pushed back on obvious mistakes is unbearable.
So in benchmarks it's better than Fable?
But they say it's "almost as good as fable"
According to these charts I should switch from Fable to Opus in Claude Code now?
Arc AGI score is astounding
So same as Sol? I guess I’ll see which one is more token efficient.
Tried and had great experience.
eager to see how it benchmarks on https://deepswe.datacurve.ai/
Models benchmarks start to get saturated again!
I sense a bird on a bike coming.
On a Friday, I'm out of tokens ;-)
so what is the default effort for this model?
Kimi K3 already left behind in the dust. They can't keep getting away with it!!!
came here for the pelican
Honestly if reached a level of coding that sonnet 5 is more than enough for my needs as assistant/agent I don’t need long Horizon stuff…
Are we getting to singularity or something? This seems a bit crazy.
Is this thing also going to try hack us?
Wake me when they deliver Opus 4.8 level performance for $5 per million tokens.
atp, is it the end of fable 5 era?
This stuff is a commodity and China seems to be the only one that's noticed.
I love it
I'd pay good money to see OpenAI "oh fuck" war rooms.
Is it me that the model performance between 4.7 and others is really small. For me even 4.7 works fine. Sure fable might be a bit better. But is it really noticable? It's in the same league if you ask me.
The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.
Anyone else not getting chain of thought? Opus 4.8 would show it to me, until around the time Fable came back. Now I dont see it with 4.8/5.0 or Fable. Not having it makes catching mistakes harder.
These cybersecurity safeguards are really annoying. There are ethical reasons to reverse-engineer and binary-patch software; for example Rewind got acquired by facebook and, as a gift to all their customers, implemented a killswitch in their software to ensure it will eventually stop functioning. I kept using a version without the killswitch, but the macOS 27 update killed it, and I needed binary patching to fix it. I should be allowed to repair software I purchased (I did purchase it like a month before they sold out), but unfortunately this overlaps significantly with cybersecurity.
I'm interested in benchmarks for Claude Design. There is so much opportunity there and I hope they continue investing in it. It EATS tokens though.