logoalt Hacker News

WhitneyLandtoday at 3:02 PM4 repliesview on HN

False dichotomy right?

Are Chinese labs impressively innovating? Clearly.

However this doesn’t rule out possible gains due to distillation.

I don’t know the degree of the latter but both things could certainly be true.


Replies

SirHackalottoday at 3:34 PM

Didn't Anthropic train on our collective data just to sell it back to us for $100/month? On top of that, Apple is suing them over alleged IP and trade secret theft by ex-Apple employees. Hard to feel too sympathetic, and I’m not an Anthropic hater in particular…

show 2 replies
fnord123today at 3:58 PM

Also possibly true: Anthropic is running Kimi locally in their hardware and "distilling" it.

cmatoday at 3:55 PM

If they can distill fable into a full model post training run in ~15 days without the real thinking traces, yet we know Claude chats degraded with the thinking traces removed (chat resume bug from earlier in the year they reported stripping thinking to shed load as being the cause of degradation), how big can this degree be?

apitoday at 3:14 PM

“Distillation” is just indirectly pirating the largely pirated training data used to train the original model.

“You stole my warez!”

show 1 reply