logoalt Hacker News

colingauvintoday at 4:11 PM9 repliesview on HN

Where is the actual evidence of distillation? I keep seeing this repeated ad nauseam but I must have somehow missed the evidence.


Replies

voxic11today at 4:16 PM

Distillation a pretty well documented technique that actually pre-dates LLMs https://arxiv.org/pdf/1503.02531

Here is a project that guides you through it if you want to prove to yourself that it works https://github.com/arcee-ai/DistillKit

show 3 replies
hadlocktoday at 8:35 PM

It turns out you can train a 1b model at almost 1000 tokens/s on a m5 max laptop. As a personal experiment, I've been asking Sol for synthetic training data and synthetic agentic training data (model distillation in it's purest form), plus modified opencode, codex transcripts etc for training data, and nobody's even paying me to do it. If I'm doing it has a hobby, you can bet industrial users are doing it.

endymi0ntoday at 8:51 PM

Been using a lot of Kimi K3 lately and the answers have been… „load-bearing“ to the point of hilariousness. It‘s obvious from where they distilled, even if sceptics rightly point out it can‘t have been the only source of their secret sauce, as it‘s been better than the current Opus 4.x at the time of release.

show 1 reply
flexagoontoday at 7:32 PM

Why does Kimi insist its name is Claude?

show 1 reply
v64today at 7:55 PM

The evidence is Anthropic's own reporting [1]. You may doubt that they're telling the truth, but that's what they're reporting.

[1] https://www.anthropic.com/news/detecting-and-preventing-dist...

InsideOutSantatoday at 4:23 PM

Musk confirmed in federal court that xAI does it: https://techcrunch.com/2026/04/30/elon-musk-testifies-that-x...

It's also how providers build their smaller models out of their larger ones; they publicly talk about the process.

JacobAsmuthtoday at 4:45 PM

[flagged]

retinarostoday at 4:27 PM

there is no evidence. it shortcuts post training by a huge margin this is true. but that is all.

show 1 reply