I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.
I support Anna's Archive, by the way. Information wants to be free.
Physical books and digital content is special in that you can mostly archive their content almost permanently for cheap. Buildings, paintings, idols, living things, natural features of the environment ... not so much.
So the solution is:
- mandatory copyright registration and renewal with links to where the work can be acquired
- a blanket carve out for any non-commercial trust-style org so that they can scan books etc and keep the data on their servers. They should be able to issue digital membership cards for a fee so that patrons can access the archives. Any work that is "live" based on the registration database will be locked. All "dead" material can be shared with members.
In this way, a hundred digital preservation societies can bloom.
“Rare books” usually refers to rare editions of books. Any books out there where there are only a few extent copies of the text itself, are probably not of very much interest or social value, since almost no one is able to read them, by definition.
If you think there is priceless knowledge locked up in books so rare that it is on the verge of being lost forever, then AI labs are not really the problem!
Pretty funny that they just took Anna’s archive and ingested it.
As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.
The very first paragraph is fascinating: "Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022."
Is the corpus of human knowledge useful for high quality AI training now essentially frozen in time? Also, how useful old books really are for AI training besides helping AI acquire knowledge about history?
I highly doubt they destroy digital copies of the books after scanning. They will want to train their future models on the same content. So what prevents them from making these digital copies available to the public? Copyright!
These stories are weird, because actual professional specialized book dealers pulp books by the millions. People keep pointing out, and it doesn't seem to sink in, that model trainers only have use for a single copy of a book. Even if they were literally burning these books to spite you, they'd be destroying an infinitesimal fraction of the books the book trade already destroys.
It is not natural in the industry to preserve books!
The copyright thing I get, sort of (I mean, it's galling, because it's such a total special pleading argument from a cohort of people who otherwise have absolute contempt for copyright on anything other than code). The model trainers are getting away with something other people haven't gotten away with. OK, sure.
But this seems like the AI water use story, where the reality is that existing industries do whatever the bad thing is at scales cosmically larger than AI ever could, and we're zeroing in on this weird little slice of it that AI does. Like, let me know when we stop growing pecans in the California desert, and then we can talk?
As much as I hate piracy in a sector in financial crisis like book publishing (because Anna’s project is piracy), I hate even more what these large AI companies are doing: privatizing human knowledge.
On one side, there’s copyright law, which exists to support the work of creative people. “Information wants to be free” is bullshit spread by people who have never spent a minute in their lives trying to create something themselves. Artists need some form of reward.
On the other side, buying and destroying copies of rare books is quite scary. We would lose access to those books if they weren’t digitized. They are creating walls around knowledge that they acquired because there are no laws in place to protect authors.
This is scary, and it reminds me of Fahrenheit 451.
Do not believe Anna’s claims, since physical book sales are plummeting — the main source of income for writers — and shadow libraries are killing the incentive to write. But even more importantly, do not believe AI companies will help you discover and access knowledge.
We might end up with all of humanity’s books digitized and accessible for free, and LLMs capable of writing entire books for us. But there would be no human writers left.
In a world like that, what motivation would we still have to read?
The AI companies should work with the Internet Archive to release the digitized copies once the copyright expires.
Unrelated: So with this one copy BS are you not allowed to have backups of the data?
To clarify: Are they scanning and destroying a single copy of Book X or are they buying up all copies of book X, scanning it once, then destroying all copies of book X they can get their hand on?
I am baffled at these practices and somewhere confused on what's the end game here? monopoly on information? altering data? exclusive subscription based knowledge? Feels like we have welcomed the AI era with open hands hoping( at-least assuming) that data democracy will be there, yet feels like its a long road!
I wholeheartedly believe the AI controversy on destroying books is being stirred up by the companies themselves.
Copyright law requires you destroy a book, if you format shift it. If you digitise, you need to ensure its not a "copy" but that your one license went with the book.
So... If enough people complain, they get to pressure for copyright changes. Which will just so happen to have massive carveouts to let them do whatever they want.
What evidence do we have that they are "destroying" books?
I'm not saying this in their defense, but as someone who has worked at companies who has scanned books at scale, and generally speaking, I wasn't on site there, but I knew we/they were pretty delicate with the books. And while the kneejerk reaction might be "hey, why would they go through the effort?" -- my guess is that they are following or even hiring people that have done this process in the past (out of laziness) and just follow what works easiest. The literal machinery is not designed to destroy the books for various practical reasons. Books that are bound are easier to be kept in order and work with. Getting a flat scan is done with specialized tools, you don't need to put it on a plate (it would be too slow that way anyway)
All of the above is just to justify my question: Who knows that the books are being destroyed? (I also agree with the general sentiment that there's a good chance these books are just cheap and bulk, they aren't pulling one of a kind rare books.)
I'm sure the AI companies will retain scans of the books for training on newer models
The question I have is, do these companies keep copies of the scans after they have finished training on them? If so, then it isn't the worst outcome. Not great but at least the information is not completely destroyed forever just the original physical being of it.
Deeper thought however, eventually this will all be lost to time and I suspect that about 99% of all printed materials probably would never be read again simply due to the huge volume of it and sheer obscurity. Ernest Becker and his work 'The Denial of Death' might have some thoughts on this.
go to any second hand book store and just pick out something at random from the 1950's for instance, something about pottery or bird watching or whatever. The history of Bisbee Arizona, I don't know. Look up the author, see if they even left a trace of their work and the vast majority of the time they have already been forgotten to the great void of the universe. In the end, it all goes away. Clinging only creates pain.
I'm not saying that we should let them just do this, I am just saying that long term it is a tough battle to fight only to lose the war.
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).
This whole situation is such a disgusting consequence of copyright law. The most frustrating part is that its so artificial. It is 100% the consequence of stupid laws.
Being purchased and juiced for model weights is about as noble of an end as any book could hope for.
Someone should build the digital equivalent of a fire department. Train a model on the books, then if the originals get destroyed you still have the smoke.
>It’s outrageous is that it’s legally permissible
No its not.
>but ethically, it’s an extremely serious crime against humanity.
Its only a crime if they dont also upload the scans to the internet.
>After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers.
This Law on the other hand is a crime against humanity.
>Anna’s Archive needs a plan to combat the destruction of physical books by AI companies.
No it doesnt.
>If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth.
This however is an unvarnished good.
Look, piracy is the only realistic media archive we have.
We should be inviting, and working to eliminate opposition to, AI companies to assist in piracy.
This US v Them mentality is weird. If Anthropic has 10 million books scanned, get a copy. Thank them for the copy. Spread the copy.
"Whoever destroys a book destroys a link in the chain of human knowledge"
-- Thos. Jefferson
It's giving Vishnu, but the world cannot exist without Shiva.
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
Aren't AI companies all about the rare book auctions now?
that's ironic, the url annas-archive.gl is blocked by my local DNS category for AI Threat Detection.
Can someone name a rare book that was destroyed as part of AI scanning? I want to know what kind of thing we're losing.
Google Books was a great resource until the lawyers got involved. I was able to find and download (one screenshot at a time) a rare family history. The author died 100 years ago. The published disappeared 80 years ago. But now Google has locked it behind a limited preview.
Google probably has the best collection of high quality scans, followed by the Hathi Trust. None of which are useable by anyone outside of those systems.
I can imagine 100y from now, most if not all books and knowledge are in electronic format or even just as part of an AI, then a wild solar flare wipes out all electronics in a minute..
It's not a bad idea but we need a multi-pronged approach, with at least one other prong being "destroy the companies that are doing this".
The hysteria around AI and data centers has hit a precipice. It's actually a bit embarrassing now. I am pretty sure there are foreign adversaries that are trying to stop the US, but I also really blame the AI companies for doing the most horrendous job imaginable in pitching AI to the public. Not a shock that people are against something that tech bros have claimed will destroy everyone's lives in the next 5 years. These books were probably going into a landfill without AI companies getting them, regardless. Tons and tons of books go into the garbage every day.
[dead]
[dead]
People keep repeating the "rare books" without providing any evidence that they are rare. Anyone who has collected books knows there are massive volumes of old books that can be bought by the pound.