logoalt Hacker News

Livenerf: Has Opus 5.5 been nerfed yet?

89 points • by bryan0 • today at 10:36 PM • 48 comments • view on HN

Comments

johnfn • today at 11:36 PM

"Nerf"ing models isn't real. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

I made a graphic to explain why people feel like the models get nerfed:

https://x.com/thesilenceturns/status/2103551351825543610

The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.

➕ show 5 replies
jug • today at 11:34 PM

We also have Nerf Bench:

https://www.bridgebench.ai/nerf-bench

They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.

This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.

➕ show 1 reply
xlayn • today at 11:43 PM

The only reason why claude fable is better than opus in my opinion is that it has more "criteria"... if you present a problem and then ask for his recommendation you can get an opinion on why and reasoning on why that one... Opus is going to vomit 10k lines of extremely dense prose in nerdify++ level.

Yesterday I fought claude fable to not just jump to make changes like a dog following a treat, that we were researching... at some point I introduced the word HAWAI... and only if I say HAWAI the thing can start making changes..

I was going to post here in HN just to have a "I knew this was the reason" when they release fable > 5.1

I had the exact same feeling every time they have a new big release

nico • today at 11:30 PM

Anecdata: I've been running a long-lived claude code session with Opus 4.6 for the last few days. Yesterday, almost right after the Sonnet 5.5 announcement, codex starting asking for permission to run things a lot more often

The quality of the output/work seems the same, but the speed at which it gets stuff done is a lot slower, because it's asking for permission so much more

I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output

➕ show 2 replies
judge2020 • today at 11:15 PM

I wonder if more organizations approving the model on a fast-tracked basis means Anthropic is straining for more compute and thus sheds a tiny bit to handle the increased demand, especially at peak times.

aabhay • today at 11:07 PM

Only ten day interval? I felt Astra got nerfed within a week

➕ show 2 replies
whs • today at 11:28 PM

I wonder if API is affected by this issue, especially Claude on public clouds? Would that means the subsidized rate just means they use cheaper quantized models and it's not comparable to API spending.

➕ show 1 reply
solfox • today at 11:26 PM

It seems as if this is based on demand. Whenever a new model is released, I'm guessing tens of thousands of us switch over to try the latest and greatest, which overloads the servers, leading to nerfing. It's 100% dishonest, but they realized they would lose users a lot quicker if they were honest and just said "our models are overloaded, come back later".

After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets.

I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again?

gr_norm • today at 11:28 PM

All this dishonesty and shadiness is part of why open models feel inevitable. Even if the total cost of ownership is higher (debatable; seems that way at small scales, but likely not as you grow), I'd rather have intelligence controlled by me that works for me.

The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word. The point isn't even whether they're nerfing the models (I don't think they are), but that people can't seem to trust them to do right.

bethekidyouwant • today at 11:53 PM

People just tend towards conspiracies you have to actively fight it.

LeoPanthera • today at 11:32 PM

n=1 is useless. The output is not deterministic.

octoberfranklin • today at 11:34 PM

OpenAI will simply set up a classifier to detect if the client is livenerf, and selectively not nerf those requests.

Open models are the endgame.

➕ show 1 reply
Razengan • today at 11:33 PM

Theory (Conjecture? Hypothesis?): What we notice as "model nerfing" is the company diverting compute to training/running new unreleased models..

Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced

colordrops • today at 11:33 PM

This repo already has too much visibility now. Anthropic will soon benchmaxx it.

gigatexal • today at 11:02 PM

This is genius. I’m so worried opus 5.5 will get nerfed cuz sonnet 5 was such trash I can’t go back.

solenoid0937 • today at 11:19 PM

Hot take, none of the models are getting "nerfed", people are just getting used to the new level of intelligence.

➕ show 6 replies