logoalt Hacker News

johnfn • yesterday at 11:36 PM • 8 replies • view on HN

"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.

I made a graphic to explain why people feel like the models get nerfed:

https://x.com/thesilenceturns/status/2103551351825543610

The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.


Replies

prodigycorp • yesterday at 11:53 PM

Incorrect.

Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.

Your chart is wrong.

➕ show 3 replies
gobdovan • yesterday at 11:52 PM

There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs, e.g.: https://www.anthropic.com/engineering/april-23-postmortem

Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.

➕ show 1 reply
HawtAds • yesterday at 11:58 PM

It's very much real but not necessarily malicious. We track upstream providers pretty closely. Sometimes it's a just matter of a single GPU runtime layer bug/update to break inference outputs. The model weights don't necessarily change/get quantized.

➕ show 1 reply
eek2121 • today at 12:08 AM

Admittedly, I didn't click your link, however, based on what you've stated, there is some inaccuracy. All these big companies take your requests and the context, and route it based on the content, cost, etc.

What Anthropic presents as Opus 5.5 isn't actually a single model...it's Anthropic's ecosystem as a whole. If you are lucky, you get the top model handling your issues all the time, however, that never happens. What really happens is that your request and content are graded along with your subscription (example: API? subscription, if so, what tier? how much has the user used it? Do we trust the user? how much? how much are they paying? are they asking something we think is dangerous?) and your request and context are routed accordingly.

Anthropic isn't alone in this behavior, Open AI does it as well, just look at the respective subreddits on reddit for both if you need some examples, or just play around with the various models from both companies.

There are a few folks who've done some analysis on this (their findings were posted on reddit and X), and a bigger multi-national study is apparently coming, though I admittedly don't know their findings.

I guess the tl;dr is that Anthropic and Open AI are actually selling you "best-effort" routers, so you may or may not get the best in class model, and only they get to determine if you do or do not. No guarantees.

physicallyIllfr • today at 12:02 AM

Its because they're addicts and addicts always grow numb and immune to their fix, needing more dopamine. They want to feel what it felt like the first time.

By the way, dont for a second think LLM hourly limits are all about revenue, they're playing into this psychology. They hire literal gambling UX designers, they want to turn you all into addicts. They want to make you reliant.

Want to run your llm like a slot machine? They'll let you do that spin the generation on a multiple, get 6x results, pick your favorite. Feel that high.

Just know you can get that same hit of dopamine by fostering your own intelligence and creating something with it. Token dealers are just selling you the shortcut, straight to the reward, short circuiting the the natural process.

Bad times ahead for many. This shit isnt good for your brain. And you all know the truth, you just wont admit it. Its doing damage, making you lazier, less intelligent.. Making you an addict.

hbn • yesterday at 11:39 PM

I have not been doing increasingly complex things since Opus 4.6 when models got really good.

My work at my job has stayed the same. But the model quality has varied.

They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.

➕ show 3 replies
Grimblewald • yesterday at 11:49 PM

Nerf is real, i think we initially get full precision models and later quants. My own logs show it clearly for opus 4.5 to 5, consistently a few months post launch, models start making quant based mistakes, like slipping in inappropriate tokens (e.g. chinese ones in english text) which doesnt happen at all in the first few months and regularly later. Additionally frontier problems previously done well start being done poorly, until later model variants where performance mostly holds, likely due to them training on your data reguardless of what boxes you tick.

My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.

How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?

486sx33 • yesterday at 11:39 PM

[dead]