logoalt Hacker News

ungutyesterday at 4:32 PM2 repliesview on HN

OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models. There is just not enough non-copyrighted data out there.


Replies

yorwbayesterday at 4:50 PM

You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.

show 3 replies
deadbunnyyesterday at 5:08 PM

What tosh. It's copywrited material, they pay to access it like everyone else.