Raw stolen data that is in no way related to AI vs a very very expensive transformative compilation of that raw stolen data that results in usable AI. The frustration is completely understandable to me, since distillation skips much of that "very very expensive" part.
And yes, I understand the stolen data was expensive to make, so I understand the owners of it are also frustrated, but that's partly a problem with current law. Would the authors of the world be rich if OpenAI bought a single copy of their book to legally scan? For best sellers, that's somewhere around pennies, so no. Should the authors get a share in OpenAI? Current laws says, unambiguously, "no".
Frustration all around is reasonable.
Not really, in this context the word "distillation" is really being abused, or at least used in a different sense than when it was originally introduced in the Hinton et al "Distilling the Knowledge in a Neural Network" paper, where it was essentially referring to knowledge compression.
The way Anthropic are using "distillation" is just in a very broad vague sense to claim that some data generated by their model was used to help train another one. They are not talking about something like internal logits, expensive to derive, that would be useful to train a smaller model, but rather about any output from their model, even outputs with redacted reasoning (i.e. incomplete outputs that do NOT reflect the underlying knowledge of the source model).
Given the way Anthropic are using the word, IMO it's better just to think of this as cheap training data, and indeed very similar to the way Anthropic themselves got cheap training data just by taking it (even in cases where that was illegal - copyright). The alternative for Moonshot would be to pay for human generated reasoning data, just as the alternative for Anthropic would have been to pay human developers for coding data etc, not just take it from wherever they could lay their hands on it.
So, I guess Moonshot may have violated Anthropic's TOS, in using Anthropic output to compete against Anthropic, but unlike Anthropic they at least didn't break copyright law since model output is not copyright, and technically they may not have even violated Anthropic's Terms of Service unless they owned the accounts used to access Anthropic's models (perhaps not - they may have used one of the anonymizing Chinese token resellers).
So yeah - pot calls kettle black.