logoalt Hacker News

“We have information that Moonshot distilled Fable for the development of K3”

173 pointsby softwaredougtoday at 2:42 PM409 commentsview on HN

https://xcancel.com/mkratsios47/status/2079933645888880708


Comments

himata4113today at 5:10 PM

Does this matter? Distillation is not illegal by every definition of the word.

There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.

Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre-cursor acquisition.

And lastly, kimi architecture is vastly different than that of fable as it uses mechanisms developed by... kimi themselves. US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.

Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.

edit: (moved this to bottom) The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens.

show 18 replies
throwa356262today at 4:29 PM

Kimi K3 was released July 16, Fable ban was lifted on July 1 but access was still limited.

How did Moonshot "distil" a huge model in such short time and still had time to run the benchmarks and do the usual release thingies?

I think Anthropic is desperate to stop foreign competition and the administration is happy to help because they too are heavily invested in these companies

show 9 replies
madducitoday at 5:08 PM

So what is the issue here? Distilling is still fair, on the same level like Anthropic scraped copyright protected material for their training.

So here robbers are blaming robbers?

These claims are just pointless, everytime

show 10 replies
linkregistertoday at 7:21 PM

Commenters are overlooking the significance of this information and posting emotional reactions based on perceptions of fairness or feelings of schadenfreude.

The economic viability of Anthropic and OpenAI rely on their being able to charge more for model access than their R&D and inference costs. If the market price for SOTA model access drops below that level, then these businesses will have to decide whether to continue to lose money or to reduce spending on R&D.

Moonshot's papers [1] claim that their training load was primarily from synthetic data and model self-teaching rather than RLHF and therefore keep their costs low. If Moonshot genuinely does not rely on human-led training, they will surpass US closed-source model providers. The United States government considers US supremacy in "AI" as a national security consideration.

This announcement is noteworthy because it implies that Moonshot's success is in fact due to distillation. It's in the interest of US frontier labs to place barriers to this if they find themselves in the position of subsidizing rival labs' research.

1. Kimi K2, https://arxiv.org/html/2507.20534v1

show 3 replies
JumpCrisscrosstoday at 8:42 PM

"Samuel Slater (June 9, 1768 – April 21, 1835) was an early English-American industrialist known as the 'Father of the American Industrial Revolution', a phrase coined by Andrew Jackson, and the 'Father of the American Factory System'. In the United Kingdom, he was called 'Slater the Traitor' and 'Sam the Slate' because he brought British textile technology to the United States, modifying it for American use. He memorized the textile factory machinery designs as an apprentice to a pioneer in the British industry before migrating to the U.S. at the age of 21."

https://en.wikipedia.org/wiki/Samuel_Slater

show 1 reply
sent-hiltoday at 6:39 PM

Reminds of the quote by Bill Gates.

> "Well, Steve [Jobs]… I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."

Source: https://www.goodreads.com/quotes/824084-well-steve-jobs-i-th...

show 1 reply
bradfatoday at 4:51 PM

I can understand that the AI labs might care about other labs distilling their models as it can eat into their competitive advantage, but do consumers care at all? Aren't consumers benefiting from this practice by getting better cheaper models as a result?

show 4 replies
ekelsentoday at 7:50 PM

Reminds me of this classic line from the 1973 movie The Sting:

"What was I supposed to do? Call him for cheating better than me in front of the others?!"

Said in response to being out-cheated at a high-stakes poker game.

Except in this case, it sounds like that's exactly the path they have chosen.

https://getyarn.io/yarn-clip/7612c4ce-1077-479f-a7bf-617dbc6...

skeledrewtoday at 5:15 PM

Super interesting. So Fable was really made available... a couple weeks ago? And K3 a few days ago? That's a really impressive feat to distill enough data AND train AND review to get a release that works really well in that time period. Mad props to the Moonshot team :flame:.

kjs3today at 8:48 PM

People who built a business model around stealing other peoples stuff are vewy, vewy upset that someone has built a business model around stealing their stuff.

Oh no! Anyway.....

overgardtoday at 8:54 PM

It's both acceptable and inevitable. Play stupid games, win stupid prizes. This is what they deserve for the awful way they treat creators and creative people's intellectual property and livelihood.

teravortoday at 6:23 PM

the distillation everyone talks about in respect to LLM's isn't nearly as easy as most think.

none of the frontier labs provide probability distributions over the tokens which is the actual method of distillation you use to train a smaller model based on a larger one. they don't even provide all the tokens.

therefore this so-called distillation the frontier labs whine about is just a set of clever methods to work the existing LLM into the training process for a new model. methods like having the existing model grade the output of the new model and work those grades into the RL method. give the new models structured tasks and use the existing model as a source of truth for those tasks and a myriad of other hacks.

efficiency scales with the gap between the models and generally allows an efficient bootstrap process. the implication that distillation wouldn't allow further advancement is false however, you can then start doing the same thing the frontier labs have been doing: dumping cash on humans to provide the signals or burning tokens on exploratory paths and grading the results.

what openai and anthropic don't like is that fact that all the cash they burned can be used to benefit everyone and not just them. and that no matter how much more cash they burn to build up the gap it will closed at a small fraction of the price.

laweijfmvotoday at 8:52 PM

if its that simple, why doesn’t Anthropic just distill its own models and release Fable 5.1, 5.2, ...?

HeavyStormtoday at 5:04 PM

Poor AI labs... All they hard earned training, done via scraping a lot of people works for free, now being scraped through payed subscriptions...

exabrialtoday at 8:49 PM

Oh the irony... LLMs go and read a billion pirated book, but now are crying when their models get distilled from a million queries.

MiguelVieiratoday at 5:34 PM

Here's a site that asks the same questions to 22 models and compares how similar their responses are.

https://typebulb.com/u/lab/you-re-relatively-right/full

According to these results GLM 5.2 is very similar to Google Gemini and Kimi K3 is very similar to Fable 5.

The American frontier labs are not similar to each other.

show 2 replies
thih9today at 7:08 PM

One of the replies:

> @MehdiKarech

> I don't remember letting Anthropic or Open Ai scrapping my GitHub, my research gate and all my online writings L O L

https://xcancel.com/MehdiKarech/status/2080000779859939678#m

show 1 reply
Geeetoday at 4:57 PM

You wouldn't distill a car.

show 5 replies
goldylochnesstoday at 8:23 PM

it was obvious from day 1 this is what happened

the chinese labs haven't been improving at training from scratch, they've been getting better and better at distilling models, so naturally they continue to follow this path

firstly, it shows the vulnerability of exposing a model to users. a highly capable group can quite literally suck the functionality out of your model and take it for themselves, so there's no sandboxing it

secondly, it's actually interesting to know that distillation is so powerful. in the science-fiction scenario of meeting some other form of intelligent life, they might have their own models and this would be a way of siphoning intelligence off of them

storustoday at 5:56 PM

I doubt they did any distillation as Hinton defined it (requiring logit access). They most likely ran a bunch of prompts/conversations and captured the results. Those conversations already missed thinking tokens, replaced by some confusing quasi-summaries. Then they took those and ran basic SFT or maybe DPO if they had competing responses. As there is no copyright on the output of AI, I am not sure where is the "covert industrial distillation" part of the problem.

nchmytoday at 6:31 PM

We have information that Claude distilled billions of copyrighted, and otherwise-created-by-others, materials for the development of their entire business.

feverzsjtoday at 4:49 PM

If web scraping is legal, so is distilling.

show 1 reply
throwa356262today at 4:37 PM

In the meantime, reddit is making fun of Opus for "distilling" Qwen:

https://www.reddit.com/r/ClaudeCode/comments/1tqaist/opus_48...

(don't take this too seriously)

softwaredougtoday at 7:31 PM

OpenAI and Anthropic should enter into distillation agreements with other US labs. Turn a threat into a profit center.

Other US labs cannot directly distill from OpenAI/Anthropic as it’s a violation of the terms of service. It holds other US labs back. Leading them to build second tier models And in the end OpenAI/Anthropic may be unable to prevent distillation.

Why fight it when there’s clear money to make here?

mrandishtoday at 5:46 PM

How was K3 trained on data distilled from Fable when Fable was only publicly available in the last two weeks before K3 was released? The timing just doesn't work.

grim_iotoday at 5:29 PM

So, if it's that easy and fast to "copy" Fable, is it really worth that much in the first place?

Sounds like the opposite of the conversation Anthropic would want to have.

jmward01today at 5:11 PM

If 'distillation' means training on outputs then what is the legal concept of ownership of outputs? And, more broadly, is this something that could be skirted by doing it in different countries that have different legal structures? Basically, are they saying they own those outputs, not the companies that paid for the tokens, and only they can train on them? I suspect a lot of companies are saving their token histories and using them to fine tune internal models.

show 1 reply
blkstoday at 8:31 PM

Considering that LLM are created using stolen content or against licensing, it’s only fair to distill and open source them.

InsideOutSantatoday at 8:30 PM

We already know K3 is really good; you don't need to glaze it even more by telling us it's like Fable.

strictneintoday at 6:57 PM

The level of discourse here anytime distillment is mentioned is so mundane. Do we need 40 people saying the same thing about how they don't feel bad and it serves them right and all that surface level stuff on every single one of these? This is the level of insight one receives anytime you mention chocolate and dogs "Oh it's poisonous for dogs!". Yes, we've all heard that 100 times. Thanks for adding nothing to the conversation.

A more interesting part of this discussion is that consistently these Chinese models are held up as a great achievement, and that they're "catching up" when in reality they're just using the work of Anthropic and OpenAI to try and keep up with them. This isn't even to say it's not a valid tactic, but it definitely colors these announcements and proclamations about foreign companies catching up to American ones.

If I get a 1600 on the SAT and you copied my answers and got a 1540, your achievement isn't that significant.

show 3 replies
NetOpWibbytoday at 6:49 PM

GOOD

I love using Claude but Fable's unusable wrt useful work like cryptography, biology, &c.

Kneecapping my productivity when I pay $100/month is annoying af.

jerrythegerbiltoday at 5:24 PM

“However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.”

What’s actually happening behind the scenes is that certain inference providers will classify a prompt and it’s re-routed transparently to Anthropic and that’s used for distillation training, only distilling the complicated traces they need, originating from real user prompts and traces. These inference providers are explicitly blocked in the claude cli if you reverse engineer it.

The real picture is that these Chinese labs have figured out how to get exactly what they need, at a high quality, directly from distinct and unique real user prompts.

It’s only “covert” because Anthropic doesn’t like it, while simultaneously being perfectly fine to do.

blaufasttoday at 7:17 PM

The frontier labs' work is more akin to discovery than artistic expression. An art piece is valued for its uniqueness and individuality, but AI is valued for verifiable correctness. Discovery cannot be unseen and is easily replicable. I think the AI labs are in a tough situation because their work is more similar to fundamental scientific discovery than say, a unique painting or song.

Mendel doesn't get a cut every time somebody uses the principles of heritability he discovered, and Einstein's family aren't getting royalties if you compute relative speeds. I think the frontier labs should expect to be treated more like scientists than artists in this regard.

wnmurphytoday at 6:33 PM

It's funny to me that these models were created by effectively "distilling" all available content including the proprietary works of many other people, but now it's a problem that someone is doing the same to them.

You're using available information (copyrighted works, or the output of another model) to train a model to encode the information in a new form. Why is the former not theft, but the latter is theft?

seydortoday at 8:22 PM

Good pretext for bannign chinese APIs in the USA and their vassal states.

If distillation is so good, why aren't US companies distilling each other?

pandinustoday at 5:29 PM

As with many others among these threads I don't see how the timing works out for K3 to have trained on distilled Fable usage. There should be at least a tacit academic acknowledgment of Kimi's own design efforts.

Distillation itself, however, is still clearly valuable - else competitors wouldn't pay so much to their rival on distillation campaigns or try to circumvent anti-distillation defenses.

As for the morality of it, if you paid for the tokens they're yours. It is already understood that you own the output. Seems to me like a variation of ordinary business arbitrage. Providers might object to certain use-cases or intention and try to craft terms around that, but that's hard to enforce at scale.

show 1 reply
kamranjontoday at 6:54 PM

So here is an important question I think.

If LLM outputs aren't copywriteable and you create your own synthetic training set using Fable and share it publicly on huggingface, and someone else uses that training set to fine-tune a model, would this be considered illegal?

I ask because this happens all the time, synthetic datasets have basically become a key aspect of training a model at this point. I even generated a synthetic set from DeepSeek v4 to aid in fine-tuning a classifier just a few weeks ago.

So I just wonder on what grounds any of this makes sense, I wouldn't be surprised if some of these American labs were using open models on their own self hosted infrastructure to generate training data, but by nature of them being open nobody has to know.

I'll make a prediction: I don't think we will ever see any of the evidence of this "distillation" before they end up implementing some type of ban.

goldenarmtoday at 8:14 PM

I have information that Anthropic distilled the internet for Fable

econtoday at 6:56 PM

I have an idea! If they are so hungry for citable content they should start a cheap or free blogging platform with images and video and a blogroll and verified credentials and resumes, with your own html css etc and domain name and a git server and a mail client and their own advertisement platform and aggregator and a chat platform, scientific journals too obviously, tools for writing and publishing books and documents. API available everywhere to avoid training on its own output.

Because there is no way in hell I'm going to make an effort creating quality content for existing platforms. The website should be entirely my own without moderation subject only to my local legal system.

Can just insert this comment as a prompt and vibe code everything in a few days⸮

cmiles8today at 5:47 PM

But wasn’t fable distilled from knowledge taken from others? I get why Anthropic is angry here, but it would appear they’re not really in a position to complain about this.

sscaryterrytoday at 8:06 PM

How much credible, provable evidence? None really.

show 1 reply
4chandailytoday at 6:18 PM

Seems to me like Moonshot is a paying customer, and if their business isn't worth the money Anthropic is charging, perhaps they should raise the price per token charged for it. Otherwise, I don't see why this is a story. "AI Company pays another AI Company for training data" just isn't that interesting.

wseqyrkutoday at 7:49 PM

Every model is distilled internet. This is just the natural progression of that.

alastairrtoday at 6:01 PM

Presumably the frontier labs themselves can do their own distillation far better than the chinese labs can. Why can't they just beat them at their own game and release / host low cost intelligence and own the whole game. There will always be a market for the more expensive frontier intelligence.

cregytoday at 8:40 PM

I think the major take away is, Chinese labs are very good at stealing others AI work and offering it did a fraction of the cost. This will be the reason the AI stock market bubble bursts. Unless they add in protection similar to patents, which stops these copied models being used by business.

mrhottakestoday at 5:56 PM

Good. If Fable is really so smart, it wouldn't let itself be distilled.

muldvarptoday at 7:01 PM

Okay? We have information that Anthropic sucked up all of the internet for the development of Fable.

NichoPaoluccitoday at 5:13 PM

I wonder if this points at a “shared” future (or at least things will eventually converge there whether companies like it or not). Ultimately, if you’re going to release these models that are fundamentally built on shared data - it’s pretty wishful to assume you’ll be able to harbor that model and the data, forever, and profit from it.

It also leads me to think about things like the original release of Fable 5, people were complaining that it was safeguarded too much - if you lock the models down too much they cease to be useful. So it’s going to be increasingly difficult to protect a model from competition while ALSO keeping it useful.

show 1 reply
jchwtoday at 7:10 PM

Is this person trustworthy? I struggle to believe that in the relatively short time Fable was available it has already been distilled so effectively. If this really is actually true, very impressive work.

eyevztoday at 7:58 PM

K3 frequently refers to itself as Claude in reasoning when instructed to play a role.

show 1 reply

🔗 View 50 more comments