logoalt Hacker News

ACCount37today at 1:03 PM5 repliesview on HN

Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.

The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.


Replies

voakbasdatoday at 1:34 PM

Think about that last point for a moment. Our “rights to read” are diminished significantly with digital works as compared to printed works. Right of resale. Right to lend.

In the end, digital publishing just isn’t right and will lead to massive gap in our historical records. They require active curation and cannot be preserved simply by resting on a dusty shelf.

show 2 replies
edoloughlintoday at 2:51 PM

> What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

Strictly speaking, no one needs the Sistine Chapel or the Pietà etc. It would be a shame if they were mulched and recycled, though.

show 1 reply
TeMPOraLtoday at 3:45 PM

Machines for non-destructively scanning books were developed and perfected long ago. The destructive scanning is neither technological limitation nor an issue of expedience. It's an issue of copyright law and fair use.

show 1 reply
dragonwritertoday at 1:36 PM

> if copyright wasn't a thing, there would be much less need to scan any physical media.

Because there’d be much less content created in any media to capture in the first place.

show 1 reply
classifiedtoday at 1:27 PM

It's the law, logic doesn't enter into it.