logoalt Hacker News

talon8635today at 5:49 PM16 repliesview on HN

Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

For an industry that’s stagnant in progress yet relies on new frequent releases to survive (non-progress being an existential risk), this could make sense.

I have no idea if that’s what’s happened, I completely pulled it out of my butt. And I have no idea is the actual frontier is stagnating.


Replies

AmazingTurtletoday at 5:51 PM

> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.

Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.

show 5 replies
Aurornistoday at 6:45 PM

> to create a perceived improvement when in reality there isn’t really one?

This wouldn't explain progress on benchmarks (including closed sets), or the fact that newer models are providing solutions to major math problems that older models cannot.

show 6 replies
nxc18today at 6:05 PM

There must be some benefit if all the providers are doing it independently.

GPT5.6-Sol on Max thinking just became regarded as of a few days ago.

The boosters will tell me it’s my fault for using such an old, cheap out-of-date low quality near useless wish.com model (that was SOTA and better than human coders one month ago).

The cycle repeats.

show 3 replies
progvaltoday at 6:05 PM

This sounds similar to rumors about how SSD companies work. First they would design a new drive with better performance that everyone uses to benchmark against other models; then slowly change its parts to worse ones, either because they are cheaper, the originals are no longer available, or whatever reason

show 1 reply
fnordpiglettoday at 6:02 PM

The Opus 4-6,4-8,5 arc is exactly this. As one person commented in here, opus 5 is a terrorist. This is undeniable. Opus 4-6 was awesome. 4-8 was worse behaviorally but produced better code.

Fable seems to be following the same enshittification arc of other Anthropic models.

Generally OpenAI seems to be taking the opposite approach with an increasing improvement over time. As sad as I feel to say this, open ai seems to have the right strategy. Making your product worse over time rarely plays well with customers. At this point it feel often hard to justify using Anthropic for anything. I generally like Anthropic better as a company and they really had the initiative and advantage and customer good will, then proceeded to squander it faster than a cigarette company or the Sacklers could have.

show 2 replies
whatever1today at 6:36 PM

You can serve Fable from a cloud vendor (like AWS, Azure). They have frozen versions of the models, so likely this should not be an issue?

I would do a test to verify my suspicions.

show 2 replies
jaredklewistoday at 8:12 PM

This would only provides a benefit if we're approaching some sort of theoretical limit of how good LLMs can be with the current approaches and data.

Otherwise, even if one company did something like this, everyone would notice because the other companies would be pulling ahead. Are all the AI developers coordinating a "dumbing down" of models? i.e. Are Open AI, Anthropic, Google, Meta, DeepSeek, Mistral, xAI, and so on, all working together?

So we might be approaching some limit (the "there's only so round a sphere can get" argument). But I very much doubt there is some massive conspiracy between all the AI developers.

rfgplktoday at 6:22 PM

Yep, that's what they've been doing for a long while now. Also the amount of tokens you get per sub varies drastically from month to month. Needs to be regulated.

kleiba2today at 6:08 PM

> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Not unless your competitors do the same, or else you will only be perceived as falling behind others.

show 1 reply
Groxxtoday at 7:52 PM

The Shepard tone of "progress"

briffletoday at 6:01 PM

I have not been attributing it so much to malice, just that all the major cloud vendors seem to be running at full capacity, and can't build new datacenters fast enough. I just kind of assumed that as they got busy training newer models, that they allocated less resources to handle the existing systems, because they aren't able to get more capacity right now.

show 1 reply
wmftoday at 6:15 PM

...releasing a new model that’s marginally if at all better than the original...

This isn't what we see in benchmarks.

tsunamifurytoday at 5:52 PM

Every SOTA model I've used at launch uses deeper, longer inference then gradually turns down over time, until the next model comes out which seems to be trained on some new data, but mostly performance due to deeper longer inference for another period.

cyanydeeztoday at 7:12 PM

you mean like a Shepards Tone (https://en.wikipedia.org/wiki/Shepard_tone); i wouldn't doubt they slowly tweak quants to try to eke out.

there's also probably load balancers that downgrade models during high use.

holoduketoday at 6:55 PM

Nah its because they cache and preprocess requests by dumb models and send them too often to another dumb models instead of the top tier model.

ekjhgkejhgktoday at 6:35 PM

> For an industry that’s stagnant in progress

Yes, the AI technology is known primarily for how stagant it is.

show 1 reply