Key takeaways
- ISBNdb removed its AI training service landing page after backlash over rare books being destroyed for LLM data sourcing.
- Court documents from January revealed Anthropic's Project Panama program purchased and destroyed books at scale; the company settled with authors for $1.5 billion.
- Bulk book orders targeting rare and specialized publications disrupted educational institutions already struggling with resource constraints.
- AI firms turned to printed material because internet-sourced data quality has degraded, but destroying books during cheap scanning reflects cost-cutting over necessity.
The database service ISBNdb moved to contain a public relations disaster this week, scrubbing its website of all references to an AI training initiative following widespread criticism over the systematic destruction of rare books. What began as reports of suspicious bulk book orders has snowballed into a larger reckoning over how artificial intelligence companies source training data, with court records exposing a controversial program at Anthropic that purchased and destroyed vast quantities of printed material.
ISBNdb Backs Away From AI Service
ISBNdb issued a damage-control statement acknowledging the mounting backlash. “We’ve seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised,” the company wrote. “We don’t train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We’ve taken the page down.”
The now-deleted landing page had promoted a strikingly blunt vision of corporate data harvesting. According to the cached content, ISBNdb framed the opportunity with language designed to appeal to AI developers: “The world’s best AI training data is sitting on a shelf. Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.” The messaging positioned rare and specialized publications as superior training material compared to internet-sourced data.
The Supply Chain Under Scrutiny
Unusual Bulk Orders Raise Red Flags
The controversy crystallized around an unexpected trend documented by 404 Media: book suppliers and secondhand dealers reported receiving abnormally large orders for inventory they had not actively marketed. The pattern suggested coordinated purchasing by parties with specific demand for bulk quantities of books, historically uncommon in the antiquarian and used-book markets.
The timing of these orders coincided with mounting pressure within the AI industry to locate higher-quality training material. As companies struggled with low-quality web-sourced data and the contamination risk of feeding machine learning systems on AI-generated text, rare and professionally edited books emerged as an attractive alternative—assuming the sources could be acquired cheaply and without legal entanglement.
Impact on Educational Institutions
Schools and public libraries, already operating under resource constraints, found themselves unable to compete with bulk purchasing. The shift in available inventory created additional financial pressure on institutions already cutting back on acquisitions. One anonymous book dealer quoted by 404 Media’s Samantha Cole articulated the moral conflict: “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.”
The Physical Destruction Question
Central to the controversy was what happened to the books after purchase. The destruction process appeared driven by economics rather than espionage or knowledge hoarding. Books do not need to be destroyed during scanning and archival—the source material can be preserved. However, cost-cutting approaches to digitization prioritize speed over preservation. Efficient high-volume scanning often involves breaking book spines to feed pages through automated systems, rendering the original volumes useless afterward. The crumpled remains then enter the waste stream.

Project Panama and the Anthropic Settlement
The current uproar gained historical context in January when court documents exposed “Project Panama,” an initiative within Anthropic designed specifically to acquire and destroy books at scale. The unsealed legal filings revealed that Anthropic had implemented a deliberate program targeting this acquisition and disposal cycle.
Rather than face sustained legal challenge, Anthropic settled disputes with authors for $1.5 billion. The settlement amount underscored the severity of the practices involved. However, the legal outcome carried an uncomfortable implication: the courts determined that the underlying practices, while generating terrible publicity, were not necessarily illegal. Anthropic’s massive financial payout functioned as a legal and reputational circuit-breaker rather than a declaration that the conduct violated law.
The newly public record demonstrated that AI firms had already calculated the cost-benefit mathematics: expensive books, divided by the volume needed, subtracted from improved model performance, compared against legal settlement risk—and determined that the equation penciled out in their favor financially. The only miscalculation was underestimating public reaction.
The Broader Crisis of Training Data Quality
Why AI Firms Turned to Books
The pivot toward printed material reflected a genuine technical problem. The open internet, long considered the primary training resource, has degraded as a data source. Low-quality web content, spam, and artificially generated text all contaminate datasets. When machine learning systems train on output from other AI systems, they risk entering negative feedback loops where quality steadily declines with each iteration.
Books offer something the internet at scale does not: editorial gatekeeping, professional fact-checking, domain expertise, and structured presentation of specialized knowledge. A technical manual, academic text, or professionally edited magazine represents concentrated, vetted information. To AI developers, it appeared obviously valuable. The only problem was acquiring it without attracting scrutiny.
The A24 and Google Partnership
The industry’s struggle for cleaner training material extended beyond textual sources. Earlier in the summer, A24 announced a formal partnership with Google to assist in training DeepMind systems using media-related content. The collaboration aimed to help generative systems trained on entertainment material escape the visual degradation that characterizes many AI-generated images—the characteristic muddy, over-processed aesthetic that has become a industry joke.
The Publicity Spiral and Backpedaling
What ISBNdb failed to anticipate was the speed at which the story would spread and the intensity of public response. The landing page promoting AI training data sourcing faced immediate criticism across media and social platforms. The reputational damage accumulated so quickly that the company moved to erase all evidence of the initiative within days.
ISBNdb’s claim that the service was merely “a test of market interest” and “never brought to life” cannot be fully verified given the deletion of materials. However, the scale of the underlying book orders suggests that regardless of ISBNdb’s direct involvement in any formal service, the infrastructure facilitating bulk purchasing already existed and was being actively used by AI-focused buyers.
What Comes Next
The incident leaves multiple unresolved questions. AI companies will continue needing training data superior to what the general internet provides. Rare book suppliers and dealers now face reputational and possibly legal pressure if they participate in bulk sales suspected of ending in destruction. Libraries and educational institutions must contend with reduced access to inventory as purchasing patterns shift.
ISBNdb’s retreat from the AI training market does not solve the underlying problem—it merely removes one visible intermediary from the transaction chain. The business incentives that drove bulk book orders remain unchanged. Until AI firms identify alternative sources of high-quality training material or develop methods that do not require massive quantities of copyrighted text, the pressure on rare books, archives, and specialized publications will likely continue, conducted through less conspicuous channels.
Frequently Asked Questions
What was ISBNdb's AI training service?
ISBNdb operated a marketing landing page promoting the sale of books as training data for AI models, describing books as "the world's best AI training data" because they contain curated, peer-reviewed, domain-specific knowledge. The company claims the service was never launched and has since removed the page.
What is Project Panama?
Project Panama was an Anthropic program exposed in January court documents that systematically acquired and destroyed books at large scale for AI training purposes. Anthropic settled with authors for $1.5 billion following the revelation.
Why are AI companies destroying books instead of preserving them?
Books do not need to be destroyed during scanning and digitization, but cost-cutting approaches prioritize speed over preservation. Efficient high-volume scanning often involves breaking book spines to feed pages through automated systems, after which the original volumes are discarded as waste.