"> but they are different.
How, and why?"
How are they even remotely the same?
They're not even used the same way.
One is raw data input, the other is training content - designed to train LLMs.
One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.