logoalt Hacker News

tristanjlast Saturday at 11:21 PM2 repliesview on HN

There's little doubt that Kimi K3 was distilled off Claude.

Anthropic stated in February that Moonshot AI (the creator of Kimi) distilled ~3.4 million exchanges from Claude models, as explained in their press release https://www.anthropic.com/news/detecting-and-preventing-dist...


Replies

kamranjonyesterday at 3:11 AM

It’s so funny to me that Anthropic can make claims like this one with zero evidence provided.

DeepSeek and others like Minimax are publishing deep research on Multi-Head Latent Attention and Mixture of Experts, Multi-Token Prediction, novel Sparse Attention approaches, I mean they trained long context models on a fraction of the resources and gave everyone the recipe.

Chinese labs might not have the funding of labs like Anthropic, but at least they provide the receipts.

show 1 reply
overfeedyesterday at 12:09 AM

While it sounds like a lot, do you suppose 3.4 million sessions come even close to being sufficient to train a frontier model?

Assuming each session was 10,000 words each, that's 34 billion words; lets call it 50 billion tokens (0.05 trillion) unfairly pilfered from Claude. That left Moonshot needing to scrounge for the other 14.950 trillion training tokens required for a baseline frontier model.

show 2 replies