logoalt Hacker News

greenlimeteatoday at 3:10 PM3 repliesview on HN

This is potentially very bad.

Like, Library of Alexandria or Council of Nicaea bad.

We may never be able to recover the information if, say, one of these AI companies copied or translated it wrong then destroyed the source material.

Maybe it's from bad OCR, or maybe from a bad actor - but there are a lot of ways history and information could change in this game-of-telephone like transfer of knowledge.

What is the point of destroying the source material? I don't buy the copyright thing.


Replies

TeMPOraLtoday at 4:10 PM

> What is the point of destroying the source material? I don't buy the copyright thing.

It is the copyright thing.

Despite what people say about scanning, the fact is, non-destructive scanning machines have been built and perfected long time ago. This was preferred in the past, back before some major kerfuffle with the publishers during COVID, but that incidentally happened to be before LLMs became a thing, so AI companies never had that option available.

show 1 reply
1970-01-01today at 3:22 PM

They're destroyed for 3 reasons:

It is the cheapest way to get them scanned.

It is the fastest way to get them scanned.

It doesn't need to be safely archived for another century until it is resold to someone that has not yet been born.

show 1 reply
ianm218today at 3:23 PM

> What is the point of destroying the source material? I don't buy the copyright thing.

The main reason is the machines they use to scan books at scale destroy the books in the process.

show 1 reply