Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.
While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. I think this is much less true for distillation (which is kind of the whole point).
ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point was just that training an LLM takes more resources and expertise than distilling from an existing LLM so I don't think the equivalence between training and distilling is entirely justified.
No - distillation is not data inputs.
Raw materials vs. Value add.
They are different things, like ore and metal.
Distillation is a new thing we need to understand, it's probably closer to IP than not.
This is why I don't give a shit that this is happening. It's actually kind of funny to me.
I disagree that LLM models are the product of enormous quantities of copyright infringement.
The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then applying that knowledge in a new way. If that's a violation of copyright, then a human doing the exact same thing would be a copyright violation too. But it isn't.
No, they settled that yesterday, so all is forgotten. Press releases were queued for today so just in the nick of time.