Why AI Companies Are Buying Up Old Books by the Million — Then Destroying Them
As AI models increasingly train on internet text that's itself AI-generated, the industry has found an unlikely solution to keep its output from degrading: raid used bookstores, buy printed books by the truckload, and then destroy them after scanning. According to a 404 Media investigation, a company called ISBNdb—once a simple tool for libraries and booksellers—has quietly become a broker connecting AI labs with bulk suppliers of secondhand books, purchasing anywhere from 1,000 to a million books per order.
The reasoning is blunt and, in its own way, logical: books published before AI chatbots existed are guaranteed to be free of "AI slop"—the increasingly common problem of AI models being trained on text that was itself written by other AI models, a cycle researchers compare to inbreeding.
Quick Summary & Key Takeaways
- Massive Bulk Purchases: ISBNdb can arrange buys of up to a million books at a time from libraries, retailers, and estate sales for AI training data.
- Destructive Scanning: Books are typically put through a machine that cuts off the spine, scans the pages, then discards the paper—faster and cheaper than non-destructive digitization.
- Anthropic's "Project Panama": Court filings revealed Anthropic spent tens of millions of dollars acquiring millions of books, using a hydraulic cutting machine to prepare them for scanning.
- Legal Cover Exists: A June 2025 court ruling found that Anthropic's use of lawfully purchased, scanned books for AI training was "transformative" under copyright law—as long as the books were bought, not pirated.
- Companies Know It Looks Bad: ISBNdb explicitly tells clients "the optics problem is real," warning that headlines about destroyed books don't generate public sympathy.
Why Old Books, Specifically?
| Reason | Why It Matters for AI Training |
|---|---|
| Pre-2022 Publication | Guaranteed to contain no AI-generated text, since chatbots weren't widely available yet |
| Edited & Curated | Passed through professional editing and publishing, unlike raw web crawl data |
| Domain-Specific Depth | Offers structured, authoritative knowledge on specific subjects that's hard to replicate from web scraping |
| Legally Cleaner | Purchased physical copies offer stronger legal footing than scraped or pirated digital text |
What Happened? Inside the Buy-Scan-Destroy Pipeline
The mechanics of this practice, as reported by 404 Media, are straightforward but unsettling in scale. ISBNdb markets itself directly to AI companies with a simple pitch: the world's best training data is sitting on a shelf, offering curated, peer-reviewed, domain-specific human knowledge structured in a way no web crawl can replicate. To access it, though, the books generally have to be destroyed—their spines cut off, pages fed through industrial scanning equipment, and the physical remains discarded.
This isn't a hypothetical process. Anthropic's own book-acquisition effort, internally codenamed "Project Panama," became public through court filings during copyright litigation. According to the Washington Post's earlier reporting, Anthropic spent tens of millions of dollars acquiring millions of books and used a hydraulic cutting machine to remove bindings before scanning pages with industrial-grade imaging equipment. Internal documents unsealed during the litigation showed Anthropic was acutely aware of how badly the practice would look if it became public.
Small booksellers have started noticing the pattern too. One dealer told 404 Media that his weekly sales jumped from roughly 20 books to hundreds starting in April, and that he's confident the buyers are AI labs—based on the randomness of the selections and the fact that every purchased book has an ISBN, the standard identifier AI training pipelines rely on to catalog acquisitions. He noted his inventory includes rare, out-of-print titles, raising the uncomfortable possibility that some of the last surviving physical copies of certain books are being destroyed in the process.
The Legal Backdrop: Why This Is (Mostly) Allowed
This entire industry exists in the shadow of a specific legal precedent. In June 2025, U.S. District Judge William Alsup ruled in Bartz v. Anthropic that using lawfully purchased books to build a searchable digital library and train AI models was "transformative" under copyright law—a legally significant distinction from Anthropic's separate, unauthorized use of pirated digital copies, which was not protected by that same reasoning. Anthropic ultimately settled a related claim over pirated books for $1.5 billion.
Part of the judge's logic centered on physical destruction itself: because scanning and discarding a purchased book effectively replaces one legitimate copy with another (digital) copy, rather than creating unauthorized duplicates, the practice was treated as legally comparable to simply owning and using the book. In effect, the ruling gave AI companies a court-blessed roadmap—buy books outright, then digitize and dispose of them, rather than relying on pirated text.
Why It Matters: The AI Slop Problem Is Getting Real
This story connects to a growing technical concern across the AI industry:
- Model Collapse Is a Genuine Risk: As more of the internet becomes populated with AI-generated text, models trained on that data risk gradually degrading in quality—a phenomenon researchers describe as being similar to inbreeding, where each generation compounds the flaws of the last.
- Clean Data Is Becoming Scarce and Valuable: The scale of these purchases—up to a million books per order—signals just how urgently AI labs are hunting for verified, human-written text as a countermeasure.
- Cultural Preservation Concerns: Critics point out that destructive scanning, unlike the non-destructive digitization practiced by libraries and the Internet Archive, permanently eliminates physical copies—including potentially rare or out-of-print titles that can't easily be replaced.
- An Industry Aware of Its Own Optics: ISBNdb's own marketing material acknowledging that "the optics problem is real" suggests companies are treating secrecy as a deliberate strategy, not an oversight.
💡 AI Tech Safar Insight
There's a strange irony sitting at the center of this story: an industry built on digitizing and generating infinite text at scale is now paying to physically destroy printed books because they can't trust their own digital ecosystem anymore. It's a quiet admission that the "unlimited data" promise of the internet has a shelf life—once a large enough share of online text becomes AI-generated, the well genuinely starts to run dry. Whether buying and pulping millions of used books is a sustainable long-term solution or just a temporary stopgap will likely become clearer as AI labs exhaust the supply of pre-2022 printed material that hasn't already been scanned by someone else.
Frequently Asked Questions (FAQs)
Q1: What is ISBNdb, and what does it do for AI companies?
ISBNdb is a company that helps AI labs bulk-purchase anywhere from 1,000 to a million printed books at a time from libraries, retailers, and estate sales, sourcing training data that's guaranteed to be free of AI-generated text.
Q2: What is Anthropic's "Project Panama"?
It's the internal codename for Anthropic's large-scale book acquisition and digitization effort, revealed through court filings, in which the company spent tens of millions of dollars buying and scanning millions of physical books before discarding them.
Q3: Is it legal for AI companies to buy and destroy books for training data?
Largely yes, based on a June 2025 court ruling that found using lawfully purchased books for AI training was "transformative" under copyright law—distinct from using pirated digital copies, which carries separate legal risk.
Q4: Why are old books considered better training data than the internet?
Books published before AI chatbots became widespread are guaranteed to be free of AI-generated text, helping labs avoid "model collapse"—a quality degradation risk from training AI on text that was itself written by AI.
What Do You Think?
Does destroying physical books for AI training data cross an ethical line, or is it a reasonable trade-off given the legal purchase and digital preservation involved? Share your thoughts in the comments below!
Related Link:
Frontier AI Security 101: Sandboxes, Breaches & Risks
Source: Reporting based on 404 Media, Tom's Hardware, and Futurism.

Comments
Post a Comment