OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models. There is just not enough non-copyrighted data out there.
What tosh. It's copywrited material, they pay to access it like everyone else.
You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.