logoalt Hacker News

awakeasleeptoday at 2:56 PM6 repliesview on HN

Anyone who is into books eventually finds out that we're permanently losing them all the time. Like they're thrown out, and lost forever. University libraries throwing out huge collections of out-of-print material to make room for new books, study spaces, even cafés. Municipal libraries turning over their collections. Books that never made it to libraries going out of print, tossed in the trash after yard sales.

The AI companies digesting this stuff is a net win for humanity. And I'm not a fanboy! Ideally they'd upload them to Anna's archive too, but even if they keep it private forever, at least these books live on in some way in the model weights. Thats better than a landfill.


Replies

ygjbtoday at 3:07 PM

It would be really great if the companies who are doing this would commit to placing the scanned files into a public trust that would coordinate with organizations like the Gutenberg project to ensure that the scanned materials enter the public domain on schedule. Publishing encrypted archives with the keys in escrow would be a good first step.

IMO that would go a long way to resolve any concerns about losing books. I still don't like the idea of extremely hard to find or last prints being actually destroyed for this, but it certainly makes it more palatable.

show 2 replies
bitmasher9today at 3:06 PM

I would feel much better about this process if they were uploaded and if it were framed as a knowledge preservation project. This would only slightly increase the cost of the project, but have a huge impact on its perception and its net positive impact.

Of course, actually benefiting humanity is only a minor, indirect concern for investors.

show 2 replies
Aurornistoday at 3:17 PM

> Ideally they'd upload them to Anna's archive too

Not only can they not do that, they must scan physical copies because they are forbidden from using digital pirated copies from sources like this.

Anthropic had a big settlement because they were caught using downloaded digital copies. As a response they’ve ramped up their book scanning and others have followed.

show 1 reply
alightsoultoday at 3:00 PM

If or when they go bankrupt or reach agi, they will just delete them. I hope Anna's archive already has them anyways. Apparently openai's newest unreleased model, gpt 6, is capable of continuous training at Inference time, like a person is. That might be enough to delete them

show 1 reply
red_green_yelltoday at 3:22 PM

This is the right answer but the reason they can’t upload the scans is copyright law as demonstrated by Google having to settle with the publishers and allow them to remove their books and limit free access to 20% of text. The AI companies are essentially compressing the information in a huge swath of books that would otherwise be headed to landfill and making them 1000x more accessible. This is unquestionably one of those instances where capitalism is taking money from rich investors and benefiting the 99%.

Finnucanetoday at 3:08 PM

That is the problem: they are digesting it. They are not creating a new kind of library, where you could say, show me the text of "How to Fix Your Ice Cream Problems". (An actual book I own) It is not being done for our future reference.

show 1 reply