logoalt Hacker News

yorwbayesterday at 4:50 PM3 repliesview on HN

You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.


Replies

tygonyesterday at 8:27 PM

I wonder if anyone has run the numbers on what the actual cost, both in cash and logistical headache, contacting so many copyright holders would be. That seems like quite the feat to calculate.

show 2 replies
ungutyesterday at 11:25 PM

Pretty easy to assertain that they don't acquire them legally due to the plethora of evidence and court cases against them. No copyright holder would be sueing them if they knew they sold the works in the first place.

I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days.

spwa4yesterday at 5:11 PM

I'm pretty sure we would know if they did that. And we don't.

Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!

show 1 reply