logoalt Hacker News

qriosyesterday at 4:42 PM0 repliesview on HN

The interesting part is what they try to not to say: More indications for K3 is based on distillation from Claude and GPT.

From [1]:

> As you might guess, this suggests that distilling reasoning traces may have been possible for a long time without ever breaking the cryptography.

> An anecdote: we find that prefilling Kimi-K3 reasoning with a few tokens of Opus reasoning measurably shifts its response toward Opus’s

> A small memorization analysis showed that specific Claude and GPT reasoning spans are up to ~6 orders of magnitude easier to extract from Kimi-K3 than from the next-closest model.

[1] https://x.com/kotekjedi_ml/status/2087147042888114428?s=42