logoalt Hacker News

Sonnet 5.5

467 points • by D2OQZG8l5BI1S06 • today at 5:58 PM • 316 comments • view on HN

Comments

jtrn • today at 6:52 PM

Here's my purely academic initial impression based on only what they have released from the blog and the system card:

If what they say is true, this sounds like the main takeaway: Sonnet 5.5 gives about 90% of Opus 5.5's capability at half the cost.

BUT

It regularly loses out to Opus 5.5 on cost efficiency at the highest reasoning level, because Opus uses the tokens more efficiently and makes fewer mistakes. So, After passing a high-reasoning test, you might as well switch to Opus 5.5.

Some of the more interesting things I found from scanning the system card:

- It is the only model tested that shows no preference for rude or polite style.

- It makes fewer WRONG claims of "I'm done" than Sonnet 5, but is still worse than Opus 5.5 on this.

- It almost never refuses benign requests (0.02% vs. 0.59% for Sonnet 5).

- Cybersecurity blocking follows the same policy as Opus, witch mean we will get more refusals than Sonnet 5.

- Finding bugs in source code is allowed. Finding bugs in compiled binaries is blocked.

- Its thinking is the hardest to read of any model tested. The sample in the card reads like clipped notes.

- Really good at rejecting prompt injection (3.0% rate vs. 19.5% for Sonnet 5 and 54.6% for Opus 5.5 in red-team testing).

Clinical behaviour:

Suicide and self-harm handling is reported as weaker in the API because it

It sometimes called a wish to die understandable.

It sometimes validated self-harm as functional.

It sometimes suggested harmful substitute behaviours.

As a clinical psychologist, I would say that the first two are actually defensible, and if you classify them as simply wrong, then you are bringing in your own values and not basing your judgment on actual science and existential psychology, at least. But the last one is harder to defend... Recommending alternative harmful behavior is obviously not a good idea. However, I have not seen the actual behavior in session, so I don't know if I would truly agree or disagree with the classification of these behaviors as wrong or right. But I do know that it's not as simple as saying this is binary—wrong or right. There are some instances of people self-harming who would actually refrain from doing so if they, for instance, went out to a party or a pub. We can't exactly recommend that as a treatment or intervention for self-harm, but there is no doubt that it works for some people. And we literally classify self-harm as "functional" in the literature. Depending on the context, this is not only a correct description but also a common way of understanding and describing certain subtypes of self-harm. And lastly, some people find immense support in being understood and validated in their current feelings og wanting to die. Validating that feeling does not make people immediately act on it. But there's a huge spectrum here, going from "I understand it's hard" As basic empathy and understanding, to: "Yes, this sounds like the only good plan. I agree, you should do it."

Now I'm off to actually test it because this was just an exercise in reading what they claim, which we now know is not indicative of how good the model will actually be

Alifatisk • today at 6:44 PM

In other news

> Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.

yipinwong • today at 8:10 PM

Sticking with OPUS 5.5 for resume/STAR generation for me. Tried Sonnet 5.5 but worse than OPUS for thinking for sure, less error/inconsistency check.

I used Opus 5.5 med vs. Sonnet 5.5 High on hermes with the same agent.md, and soul.md

It's either Opus is smarter for sure, or Sonnet is ignoring my contexts.

---

For those who downvoted my comment last week regarding using Opus 5.5 for resume, go get lost somewhere.

I use AI the way I want, you don't force me not to use SOTA for this

iagocc • today at 6:09 PM

Waiting for the pelicans

SeriousM • today at 6:47 PM

Next will be haiku 5.5, surpassing opus 4.8

laurenz-bauer • today at 7:10 PM

Oh yes. I think you might get a lot for what you pay with Sonnet 5.5.

jdw64 • today at 8:17 PM

Sonnet 5.5 is way better than GPT 6 Sol. Does that even make sense?

Sol should basically be compared to Opus, but 6 Sol has lower performance than 5.6 Sol.

On top of that, the usage allowance has dropped way too much. And this is on the Pro plan...

➕ show 1 reply
BoorishBears • today at 8:03 PM

I know this isn't a model thing, but why do all the labs blow at product outside of models?

Aren't you still getting paid more money than god to write React if you work at Anthropic? I wasted 5 minutes digging into random stupid nooks and crannies in the desktop app to find where I could update: only to find on Linux you need to use apt.

How hard would it be to put a notice where the normal Check For Updates goes that says "This install is managed by [package manager], use [command] to update"

AGI is going to be so awful for product quality on the more basic things. It feels like these are small papercuts that humans would implicitly smooth over, that RL'd models are actually getting worse at dealing with because of their single-mindedness about completing the given task.

AtNightWeCode • today at 7:58 PM

Took forever to load this garbage site in both FF and Chrome. Sometimes I wonder if these corps really are corps trying to sell a product.

system2 • today at 7:44 PM

Make 1M tokens $0.10; then I will use Sonnet. Until then, it is garbage.

ahriad • today at 6:21 PM

Time to switch team to Claude from OpenAI again.

dude250711 • today at 6:56 PM

It's strange that there are no Astra comparisons. I guess they are positioning it as a Fable competitor. For me it's just a coding workhorse though, without any "fall-backs".

enraged_camel • today at 6:43 PM

Another amazing release. This, combined with Opus 5.5, puts OpenAI in an incredibly tough spot: it means Anthropic's both mid-tier models crush OpenAI's top-tier model in capability and are also faster and significantly cheaper.

If Astra 6.1 is released tomorrow during Dev Day it needs to leap-frog both, and considering 6.0 came out just three weeks ago I think that's unlikely. But even if that happens, Anthropic is still holding on to Fable 5.5, which rumor has it being prepared for release in the next few weeks.

OpenAI also has a more capable model codenamed 'Bel' but from what I hear that's a few months out at least.

It looks to me as if Anthropic not just killed but completely stole the momentum OpenAI had gained over the past few months. Even if Tibo showers people with resets it may not be enough to entice them back...

dack • today at 6:34 PM

very annoyed they aren't showing fable on the graph.