> The vast, vast, VAST majority of the original datasets were from pirated books and the like
And there's been significant legal consequences as a result
> Also, arguably a robots.txt is the exact mechanism to follow to do the mass GET-ing
You're free to argue this of course, but the courts have largely rejected it already pre LLMs. See for example hiQ Labs v. LinkedIn
Yes, anthropic and openai have really been brought to their knees and ipos cancelled because of the legal consequences of obtaining their training data.