AI Companies Are Buying Books Up to a Million at a Time, Shredding Them, and Calling It Preservation. A Judge Already Ruled It Fair Use.
TL;DR
On July 21, 404 Media reported that a quiet supply chain now exists to feed print books into AI training pipelines, and that ISBNdb, keeper of the self-described world's largest book database, sits in the middle of it: bulk orders of 1,000 to 1,000,000 books, a strict NDA on every engagement, buyer names never disclosed. The books have their spines sliced off so the pages feed through high-speed scanners, and the originals are shredded or recycled. Books printed before 2022 sell at a premium because they predate the LLM flood, the closest thing text has to a no-slop guarantee. The legal groundwork is already poured: a federal judge ruled last year that buying a print book, scanning it, and destroying the original is fair use, since only one copy exists at any moment. And used booksellers fear rare, nearly extinct titles are being swept into the same pipeline as everything else; the buyers are anonymous by design, so nobody can rule it out.
The sales pitch writes its own headline
ISBNdb spent years selling book metadata. Its new pitch, surfaced by 404 Media's Emanuel Maiberg, targets AI labs directly: "The world's best AI training data is sitting on a shelf." Books, the copy continues, are "curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative."
The logistics are industrial. The service brokers orders from 1,000 books up to 1,000,000 per engagement. And discretion is not an afterthought, it is a line item: a "strict NDA on every engagement," with buyer names "never disclosed."
The company knows exactly how this looks, because it wrote the headline itself: "The optics problem is real. 'AI company destroys two million books' is not a headline that generates sympathy." The recommended fix is not to stop destroying books. It is to describe the work as "digitally preserving" them, which is preservation in roughly the sense that taxidermy is wildlife conservation.
How a book becomes tokens
The pipeline exists because of a simple cost asymmetry. Cut the binding off a book and it becomes a stack of loose pages that feeds through a commercial document scanner like printer paper. Keep the book intact and a human, or a delicate page-turning machine, has to handle every spread. Google Books did it the slow way two decades ago, borrowing library books and handing them back. The new pipeline does not hand anything back.
At the exit of that pipeline the digital copy joins a training corpus and the paper goes to a shredder or a pulper. Nothing about the machine knows or cares whether it is eating copy 40,000 of an airport thriller or one of the last surviving copies of anything.
Why pre-2022 print is premium
The date cutoff is the cleverest part of the pitch. Since ChatGPT shipped in November 2022, AI-generated text has been soaking into the web, and training new models on old models' output measurably degrades them, a failure mode documented in Nature as model collapse. ISBNdb says the quiet part in its marketing: "Print books from the pre-LLM era are structurally guaranteed to be free of this contamination."
Strictly, "guaranteed" is sales copy: GPT-3 text was reaching print by 2020, when Pharmako-AI became the first book co-written with the model. But machine text in pre-ChatGPT print is a rounding error, which is why the 2022 line holds as a heuristic, and why the market prices it like a guarantee.
There is a perfect precedent for this, and it is not flattering: low-background steel. After 1945, atmospheric nuclear tests faintly contaminated all newly smelted steel, so instrument makers needing radiation-clean metal salvaged it from pre-war shipwrecks. Pre-2022 print is the web's sunken battleship: the only text left that provably predates the fallout.
The law already blessed the shredder
The legal question got its first real answer in Bartz v. Anthropic in June 2025. Judge William Alsup ruled on summary judgment that training LLMs on books was fair use on the facts before him, calling the use "spectacularly" transformative. And he ruled on the shredder specifically: because Anthropic bought its print books legally, destroyed each one after scanning, and kept the digital copy internal, the destructive digitization was a format change rather than a new copy. One copy went in. One copy came out. Fair use.
To be precise about what that is: one district court's ruling on one company's facts, not a nationwide rule that every book-shredding pipeline is lawful. But it is the only on-point precedent so far, it was never tested on appeal because the case settled, and the supply chain that has grown up since is exactly what an industry acting on a green light looks like.
The same order went the other way on the seven million-plus books Anthropic had downloaded from pirate libraries. The infringing act was building and keeping that pirated internal library, whether or not a given book was ever used to train anything. It produced the $1.5 billion settlement, roughly $3,000 a work across the approximately 500,000 pirated works the settlement covered, which got final approval on July 20, the day before the 404 Media story ran.
Put the two outcomes side by side and the incentive is unmissable. Piracy ended at about $3,000 for each work the settlement covered. Used books cost a few dollars each, and destroying them is precisely what makes the copyright math come out clean. The ruling did not merely permit the shredder. It priced every alternative higher.
Project Panama and "all the books in the world"
Anthropic ran this play at scale, and it ran it quietly. In February 2024 it hired Tom Turvey, the former head of partnerships for Google's book-scanning project, and tasked him with obtaining "all the books in the world," per the court's summary judgment order. In January, the Washington Post, working through more than 4,000 pages of court documents, put a name on the program: Project Panama. "Project Panama is our effort to destructively scan all the books in the world," an internal planning document read. The next line: "We don't want it to be known that we are working on this."
Millions of print books were bought, cut apart, scanned by commercial vendors, and recycled to help train Claude. No physical archive was kept.
The part that cannot be undone
404 Media's sourcing suggests the pipeline is not selective. One used-book seller said orders jumped from about 20 books a week to hundreds a week starting in April, picked in ways that only make sense if software is walking down ISBN lists, and that rare and out-of-print titles ship out with everything else. The seller suspects AI buyers but cannot prove it; the NDAs exist precisely so that nobody can. Rare-book dealers in the Netherlands report the same pattern of unexplained bulk buys with the same unconfirmable destination. And nothing anywhere in the pipeline checks how many copies of a title survive before the spine comes off.
Ingram, the largest book distributor in the US, has begun warning publishers about bulk orders bound for scanning and offering ways to opt out. That is the current state of the safeguards: an opt-out, for the books that still have someone to speak for them.
My Take
This got to me. A bookseller told 404 Media that rare books with almost no surviving copies are disappearing into these anonymous bulk orders. Books that survived wars, fires, and centuries of handling, heading into a pipeline whose endpoint is a shredder, so an AI can learn to write a better marketing email.
ISBNdb's website literally says "'AI company destroys two million books' is not a headline that generates sympathy," and they still built an entire business around making it happen quietly. They offer NDAs as a feature. They coach clients to call it "digital preservation."
I've covered AI companies scraping the internet, torrenting libraries, and stealing music. This is worse, because it's irreversible. You can re-upload a website. You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal, so it's going to accelerate.
"We shred rare books and offer NDAs so nobody finds out" is a legitimate business model in 2026. What a timeline.
Key Takeaways
- ISBNdb brokers bulk print books for AI training: orders from 1,000 to 1,000,000 books, a strict NDA on every engagement, buyer names never disclosed (404 Media, July 21).
- The scanning is destructive by design: spines are sliced off so pages feed high-speed scanners, then the originals are shredded or recycled. It is cheaper than careful, and at bulk scale nobody pays for careful.
- Pre-2022 print is premium because it predates mass LLM output: the textual equivalent of low-background steel, and a hedge against model collapse.
- Bartz v. Anthropic gave the pipeline its legal cover in 2025: a district court held that buy, scan, destroy, keep the file internal was fair use on Anthropic's facts. Piracy cost $1.5 billion at about $3,000 a book, so the ruling made the shredder the cheap, court-blessed path.
- Anthropic's Project Panama destructively scanned millions of books under an explicit instruction to keep the effort unknown, with the buying led by the former Google Books partnerships chief.
- Nothing in the pipeline checks scarcity. Booksellers say rare and out-of-print titles are shipping into anonymous bulk orders they suspect are AI buyers, the NDAs make confirmation impossible, and the main safeguard so far is an Ingram opt-out for publishers.
Sources: 404 Media, "AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop", Futurism, The Next Web, Washington Post, "Anthropic 'destructively' scanned millions of books", Akin Gump on Bartz v. Anthropic, Authors Guild on the settlement, Nature, "AI models collapse when trained on recursively generated data"