There's little doubt that Kimi K3 was distilled off Claude.
Anthropic stated in February that Moonshot AI (the creator of Kimi) distilled ~3.4 million exchanges from Claude models, as explained in their press release https://www.anthropic.com/news/detecting-and-preventing-dist...
While it sounds like a lot, do you suppose 3.4 million sessions come even close to being sufficient to train a frontier model?
Assuming each session was 10,000 words each, that's 34 billion words; lets call it 50 billion tokens (0.05 trillion) unfairly pilfered from Claude. That left Moonshot needing to scrounge for the other 14.950 trillion training tokens required for a baseline frontier model.
It’s so funny to me that Anthropic can make claims like this one with zero evidence provided.
DeepSeek and others like Minimax are publishing deep research on Multi-Head Latent Attention and Mixture of Experts, Multi-Token Prediction, novel Sparse Attention approaches, I mean they trained long context models on a fraction of the resources and gave everyone the recipe.
Chinese labs might not have the funding of labs like Anthropic, but at least they provide the receipts.