logoalt Hacker News

bluegattytoday at 12:07 AM1 replyview on HN

I see your 'fine point' but I don't think it holds - 'distillation' is a perfectly reasonable term to describe the process of creating outputs from one model to that expose key training element, to use in another model.

I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues.

Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces.

Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing.

It's hard to draw the line.

But the Chinese models are absolutely distilling - and would not be competitive without this distillation.

At the same time, there's a lot of real innovation and regular building going on at the same time over there.


Replies

HarHarVeryFunnytoday at 12:29 AM

No - you can't distill if what you are given doesn't have the thing in it that you want to distill out of it.

I don't know why it's so important to you to use the word "distillation", but it's the wrong word to use.

BTW OpenAI on twitter also said that Kimi 3 "cannot be explained away by distillation or anything like that". The timeline of how long it takes to train a model and when Fable was released don't even line up. This is just Anthropic as usual trying to manipulate the US government into helping them shut down competition.

show 1 reply