logoalt Hacker News

causalyesterday at 3:58 PM40 repliesview on HN

Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models? Trying to think of explanations:

1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months?

2) Distillation - also implausible for the reason above.

3) Benchmark hacking. AI companies have ways they can dial up performance artificially, and will reach for that to maintain the appearance of parity.

Other reasons?

Edit: Most replies are ignoring timing. It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.


Replies

ozgungyesterday at 9:04 PM

I think they dumb down their public models to be only slightly better than the competition. And the real competition is China, so the current state of the Chinese models would define the baseline.

I think one evidence is that the US has more than 5x the compute of China. With that difference in training speed, it should be impossible for Chinese models to close the gap that easily. It's also very unlikely that they sell the same public models to their private customers (military etc). We also know they talk about "unpublished internal models" for things like the last HuggingFace hacking incident. So it's not a bad theory.

https://epoch.ai/publications/trends-in-ai-supercomputers

show 2 replies
legucyyesterday at 5:40 PM

There is a widespread belief that the nature of intelligence is scalar, like how a person can have 100x more wealth than another person. If this were true, then we’d probably see breakaway RSI from a single lab.

But I think we’re discovering that intelligence is about universality, not magnitude. This is analogous to how building a universal Turing machine wasn’t merely a matter of building a calculator that could multiply higher numbers. The difference is that with calculators we consciously theorized about what universal computation would require, then we built one as a step change. Despite it having low memory and slow speeds, the first one built was as theoretically universal as any computer we have today, in terms of the surface of computations it can perform.

With intelligence, it’s turned out to be less discontinuous, which I believe has convinced people that intelligence is a never ending exponential rather than an S curve approaching a horizontal asymptote. I suspect the LLMs we have today are the same kind of thing we will have in 5-10 years, but in 5-10 years we’ll consider them to be fully universal. At that point we’ll still have improvements in tokens per second and volume of context window, but not in capability per token.

show 3 replies
logancbrownyesterday at 4:00 PM

Its possible no AI lab has any unique edge, and success is a combination of (a) having access to GPUs (b) having access to large amounts of data (c) know about the handful of techniques to build an LLM, of which nearly all are likely open source and documented in papers. So the cycle of growth is (a) and (b), get more GPUs and get more data and you have a better model.

show 2 replies
Planktonneyesterday at 8:43 PM

The simplest explanation is that 'Fable-level' doesn't mean anything; it's just hype, and there's not much difference in capability.

All you need to have Fable-level AI is to announce it, and have enough fans shift from insisting that model Y is the best now, way better than model X.

show 2 replies
chippiewillyesterday at 8:15 PM

> benchmark hacking

I think this is the main one. The benchmarks from this are heavily cherry-picked, and they also widely publicised their performance for 4.5 while downplaying the fact the benchmarks were "accidentally" in their training set

glimsheyesterday at 4:00 PM

4) There's nothing terribly special about Anthropic. No moat.

show 2 replies
redox99yesterday at 7:12 PM

I'm sure Grok 4.6 is not Fable level. Benchmarks are almost useless.

Having said that, Grok 4.6 (1.5T params) is without a doubt way smaller than Fable, maybe a Fable sized Grok would be Fable level?

show 1 reply
nikcubyesterday at 8:32 PM

It was said at the time that xAI acquiring Cursor was very smart because it would give them access to years of agent coding traces from millions of users.

$60B in SpaceX stock for Cursor was a bargain

Data + compute + being competent and smart enough to ship.

fwiw I don't think these are yet Fable level - the difference tends to get discovered in the long tail of tasks - but they're close enough, they're cheap, and the length of the frontier exclusive window is narrowing

show 1 reply
PeterStueryesterday at 8:00 PM

Why would you release a model if you are the current frontrunner? Only when a competitor pulls ahead, or comes close enough to actually get traffic, you prepare a new release.

yodsanklaiyesterday at 11:00 PM

Could it be that there's no magic formula, everybody uses the same known ideas, the same computation power, the same training data? if that's the case, we can imagine that models will be commoditized.

f311ayesterday at 8:31 PM

It's just model size and heavy RL, sometimes they overfit on specific tasks. RL can get you very far, prior models did not have such a focus on RL for agentic setups.

Look at deepseek, they improved it just by doing a lot of RL and you can see it from how it behaves. You provide very little information about a task, but since they are trained on similar tasks, they come up with a lot of assumptions and details on their own, because they were trained with such an info during RL.

moominyesterday at 4:05 PM

Yeah, I’m not convinced that there are any models as smart as Fable. Opus 5 definitely isn’t for all it has great benchmark scores. Fable displays judgement in a way I haven’t seen from any other model.

show 3 replies
jerfyesterday at 4:02 PM

Possibility: They're all hitting the same plateau of what LLMs can do with their current architectures.

I'm not stating this as a fact, but it's a hypothesis I'm keeping in my mix.

show 1 reply
extryesterday at 4:01 PM

It's because Fable is just synthetic RL tasks + scale. The secret has been out for awhile now.

show 2 replies
adastra22yesterday at 10:56 PM

Your implicit assumption seems to be that they didn’t start on this model until after Fable was released. They never stopped training though.

butifnot0701yesterday at 10:21 PM

I think it also shows that breakthroughs are not driven by innovative and research but mostly by scaling.

If this is the case, makes sense that frontier labs with similar access to compute driven by funding on same order of scale can produce improvement largely on similar pace

jannyferyesterday at 5:21 PM

Timing doesn't seem odd to me. It just seems like https://en.wikipedia.org/wiki/Multiple_discovery which I've noticed happen in many areas.

QuadmasterXLIIyesterday at 9:49 PM

I suspect that because each RLVR episode injects ~1 bit into the models capabilities, and training on a reasoning trace injects ~megabyte into a models capabilities, distillation is powerful enough right now that they’re all basically the same model

modelessyesterday at 8:39 PM

Researchers moving between companies (and other ways that techniques get leaked) is the largest cause of this IMO. It's happening continuously, so I don't see why the timing makes it implausible. A really underrated strength of Silicon Valley is California's ban on non-competes that allows this to happen and ensures robust competition between model providers both for talent (increasing salaries for workers) and in the marketplace (reducing prices for consumers). If OpenAI had been located in New York instead then Anthropic could never have succeeded, for example.

But I think the other reason you didn't mention is the timing of new compute coming online. Compute is the major factor limiting the training of these models and new datacenter investments are bearing fruit at around the same time.

tptacekyesterday at 7:05 PM

That's exactly what Anthropic said was going to happen!

Their big bet is that models are going to keep getting sharply better, not that they're going to quickly reach a plateau of quality that they can then defend.

show 1 reply
lanthissayesterday at 4:41 PM

what we're going through is the same thing as smartphones, the limiter is compute.

it used to be snapdragon came out HTC rushed out a janky phone everyone went omg htc is goat, then in the next few weeks and months others would impliment better versions and people would not notice those as much, finally sony would release a polished phone right as the next snapdragon cycle came.

eventually compute gains leveled off and apple won on taste.

nvidia/tpu is the new snapdragon. Anthropic and google both peaked on the first training run on a new tpu cycle.

you should expect amazing things within a few months of each other from everyone with access to chips and willingness to use them on a training run.

We haven't seen willingness from google to do that. So its currently xai,oai,anthropic, and probably soon meta.

Jcampuzano2yesterday at 4:03 PM

I'm pretty sure both Anthropic and OpenAI haven't necessarily been secretive that they have internal models that are much more capable than commercially available ones.

It's probably a mix of all of that plus simply always keeping one in the chamber to 1up everyone else when the time is right.

show 1 reply
ayewoyesterday at 4:10 PM

> 1) AI researchers talk and change companies often, so techniques circulate. This feels implausible because training and shipping a new model ought to take longer than 2 months?

The assumed timeline (2 months) is slightly wrong because Fable (Latin) is essentially the same as Mythos (Greek) albeit with protections against cyber and biological misuse.

Mythos (Preview) was publicly announced in April 2026 [1] which means other labs have had 4 months to catch up, not 2 months.

Assuming everyone had access to Mythos from the start, your expression, similar to other folks would have been "Mythos-level intelligence" and not "Fable-level intelligence".

1: https://news.ycombinator.com/item?id=47679258

show 1 reply
jcimsyesterday at 10:28 PM

Maybe the thing that has changed is the meaning of 2 months.

zahlmanyesterday at 7:12 PM

> It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.

When everyone's improvement (or at least, everyone's rate of increase in parameter count) is so rapid, "within 2 months" shouldn't be seen as "near-concurrent".

show 1 reply
inerteyesterday at 4:21 PM

No, it has happened to almost every other "sota" model before. There used to be a meme with a circular arrow going through Anthropic, OpenAI, Google as a hype circle. Now we can drop Google and add a couple of Chinese companies.

It's not an explanation of why it happens, I am just pointing Fable is not an exception, it has happened with almost every other model release by all these companies over the last 2-3 years.

HarHarVeryFunnyyesterday at 7:55 PM

I'm sure the SF AI scene leaks like a sieve, and companies have a pretty good idea what each other is working on.

show 1 reply
becquerelyesterday at 4:02 PM

More compute is coming online at all times.

user43928yesterday at 4:12 PM

I understand Mythos became internally available on the 24th of February.

Other labs catching up in half a year seems about right.

enraged_camelyesterday at 4:10 PM

I'm solidly in the "they are benchmaxxing" camp. This became very apparent with GPT 5.6 Sol. It, too, was widely hailed to have near-Fable level intelligence. But I used it non-stop for a week and realized that they had mostly just dialed up the relentlessness meter to eleven, most likely via heavy RLHF.

Last week I gave it a small-sized auth ticket to work on, then stepped away. I came back later that afternoon and found that it had worked for 3+ hours and written 25,000+ lines of code. I skimmed over the code and it looked like a small fix followed by a massive number of additional checks around it, including static analysis tooling.

I gave it to another GPT 5.6 and said "check this code and see if it addresses the ticket". It looked at it and said that 98% of it was garbage and should be thrown away (its own words). I then gave it to Fable, which said it was massively over-engineered. Fable's theory was that the agent implemented the fix first, but then compacted and lost crucial context, forgot what the original task was about, and kept going. After many compaction cycles it was completely lost.

Some people complain that Opus 5 stops before finishing a task. But to me, that behavior is vastly preferable to what GPT 5.6 Sol does.

show 1 reply
mikert89yesterday at 10:09 PM

benchamaxxing is easy, fable still seems organically more intelligent

bottlepalmyesterday at 4:02 PM

I think model level is more a function of the state of hardware. Once it exists and is available (and if a lab can afford it), then they can train their own 1T, 5T, coming up next 10T model.

show 2 replies
ReptileManyesterday at 7:46 PM

Sometimes you just need to know that something is possible, not exactly how it is done.

behnamohyesterday at 8:30 PM

> 2) Distillation - also implausible for the reason above.

DeepSeek V4 Flash 0731 is a distilled version of Fable into the original V4 Flash (announced before Fable), to the point that it also says load bearing and what not.

re-thcyesterday at 4:19 PM

> It's the near-concurrent release of the same jump in capability that I find suspicious; not the fact that labs can catch up eventually.

What are suspicious of? If the timing is similar maybe just everyone already are of similar capabilities and got there at a similar time?

> Anyone else find it weird how within 2 months of Fable releasing all the major labs suddenly had Fable-level models?

It means Anthropic had no real moat and no real lead. Is that weird to you?

show 1 reply
Traubenfuchsyesterday at 4:02 PM

> other reasons

Maybe research is sufficiently public and simple to reproduce or the next steps of how to improve things are sufficiently obvious to the smart people working on frontier AI.

dominotwyesterday at 9:25 PM

not suspicious at all. They are all doing the same scaling of test time, training data so getting similar results.

anyone with access to capital can produce frotier model. hell you can just ask chatgpt how to create a fontier model. recipe is not a secret despite what these 'labs' pretend

show 1 reply
inferniacyesterday at 7:37 PM

Maybe compute is the real moat (chinese possibly skip around it with distillation), xai is buildouts have been insanely fast (colossus 1 - 100,000 H100 GPUs brought online in 122 days lol) so maybe that explains them catching up

asked grok to give a compute estimate for each: - SpaceX / xAI: ~1.4 GW (owned Colossus clusters) - OpenAI: ~2–3 GW (mostly rented/cloud) - Anthropic: ~1.5–2.5 GW (multi-cloud + xAI lease)

chatgpt estimates a lower: - OpenAI: ~1.5M H100-eq ± ~0.8M - Anthropic: ~1.4M H100-eq ± ~0.7M - SpaceX/xAI: ~0.6M H100-eq ± ~0.3M

but it felt obligated to mention that "for single tightly interconnected NVIDIA training clusters, SpaceX/xAI has been unusually strong."

show 2 replies