logoalt Hacker News

RIMRtoday at 5:14 PM0 repliesview on HN

I feel outraged, but I also worry this article isn't necessarily responding to what's actually happening.

It's a weird practice to destroy a book when you digitize it, but it's at least an understandable legal strategy to ensure that the digital copy "replaces" the physical one. However, this is only going to apply to books that have active copyrights.

This author suggests that a rare 18th-century botanical text could fall victim to the same fate, but I am somehow doubtful that this is the case. Non-destructive scanning is trivial, and these kinds of books are likely being processed in a quantity that would allow for it without backing up the pipeline.

I really like 404 media, but it doesn't really seem like the evidence points to the conclusion here. Yes, AI companies are shredding books that they digitize, and yes, AI companies are digitizing old, rare books. But the rationale for the book shredding doesn't exist for the old, rare books, so I would need more evidence than just "putting two-and-two together".

At the end of the day, old books with no resale value, including rare books, end up destroyed with some regularity by libraries and bookstores. While this may be an excessively generous take, at least this way the books are getting digitized before they become pulp. The real tragedy will be if the old, rare books that were digitized are never shared with the rest of us, because they were ONLY digitized to train AI, and not to actually preserve anything.