BreakingNews
Inside the AI Industry's Book-Shredding Pipeline After Bartz v. Anthropic

Image: Flickr / Wikimedia Commons / Unsplash

Inside the AI Industry's Book-Shredding Pipeline After Bartz v. Anthropic

Courts have ruled that buying a book, scanning it, and destroying the original is fair use. That ruling is now the legal foundation for an industry-wide practice.

July 29, 20267 min read

This article was produced by the AETW editorial team.

Bartz v. Anthropic established that buying a book, scanning it, and destroying the original is fair use under US copyright law. Court filings and bookseller reports show the buy-scan-destroy pipeline has since scaled into an industry-wide sourcing practice.

The mechanics of Project Panama

Court filings unsealed in Bartz v. Anthropic describe an internal Anthropic effort known as Project Panama, which began in early 2024 to build a permanent library of digitized books for training Claude. Anthropic hired Tom Turvey, the former head of partnerships for Google Books, to run the acquisition side of the operation.

The company purchased used and out-of-print volumes in bulk from sellers including Better World Books and World of Books, then routed them through contractors such as Datamation for digitization. One vendor proposal referenced converting between 500,000 and two million books over a six-month window. An internal planning document described the goal directly: Project Panama was, in its own words, an effort to destructively scan all the books in the world, and the same document instructed employees not to discuss it outside the company.

The physical process is straightforward and irreversible. Workers or machines cut the spine off each book with a hydraulic-powered cutter, feed the loose pages through a high-speed industrial scanner capable of 80 to 120 pages a minute, and then discard or recycle what remains. Non-destructive alternatives, like overhead or flatbed scanners, exist and are slower and more expensive. Destructive scanning wins on cost and throughput.

Sources for this section

What Judge Alsup actually ruled in Bartz v. Anthropic

On June 23, 2025, Senior US District Judge William Alsup of the Northern District of California issued a summary judgment order that split the case into two very different outcomes. He held that training an LLM on copyrighted books is exceedingly transformative and therefore fair use. Separately, he held that destructively scanning legally purchased print books to build a searchable digital library is also fair use, reasoning that the digital file simply replaced a print copy the company already owned rather than creating an unauthorized additional copy.

The third part of the ruling cut the other way. Anthropic had also downloaded more than seven million pirated books from shadow libraries like LibGen and PiLiMi to build the same central library, and Alsup found that piracy was not protected by fair use regardless of what the books were later used for. That distinction, buy-and-destroy is legal, pirate-and-keep is not, became the basis for a certified class action covering US copyright holders whose books came from those pirate sources.

The piracy claims settled for $1.5 billion, the largest copyright settlement on record, with final court approval on July 20, 2026. The deal pays roughly $3,000 per book across more than 400,000 registered works, and court filings noted that more than 91% of eligible rightsholders filed claims. Federal judges in separate suits involving OpenAI and Meta later reached similar conclusions on the narrower question of whether training itself is transformative, though the piracy and sourcing questions in those cases remain unresolved.

Sources for this section

ISBNdb and the rest of the sourcing market

The buy-scan-destroy pipeline is no longer a single-company project. ISBNdb, which bills itself as operator of the world's largest book database, has repositioned part of its business toward brokering high-volume book acquisitions for AI developers, handling orders of up to a million volumes while keeping buyers anonymous and covered by a strict non-disclosure agreement on every engagement, according to its own site.

Booksellers describe similar patterns showing up on ordinary marketplaces. Sellers on platforms like Alibris and Biblio have reported abnormal-volume, subject-agnostic orders with no sensitivity to price, which dealers say does not match typical collector or resale demand. A German seller noted an overnight bulk order from the Canadian firm Zoom Books for unrelated academic titles, and booksellers in Europe and Australia have flagged unusual bulk purchases of rare, specialist, and out-of-print titles that in some cases exist in only a handful of surviving copies.

Anthropic's book lawsuit and settlement are far from the only active fronts. Meta faces its own author-led copyright suits, and nearly 400 newspapers are separately suing OpenAI and Microsoft over news content used in training. The book-sourcing practice sits inside a much larger legal fight over what AI companies are allowed to acquire and how.

Sources for this section

Why pre-2022 books command a premium

ISBNdb markets pre-2022 print books to AI labs specifically because they predate widespread generative AI, meaning the text inside is verifiably human-written rather than diluted by AI-generated filler that has since spread across the open web.

There is a second motive: data poisoning defense. Researchers have shown that a corpus can be manipulated with a small number of adversarial documents planted to create hidden behavior in a trained model, and ISBNdb's own marketing cites Anthropic-affiliated research suggesting a range as small as 250 to 500 crafted documents could be enough to plant a backdoor across a training set of trillions of tokens. Physical books printed before large language models existed are, by definition, immune to that kind of contamination, which labs frame as a clean chain of custody worth paying for.

Sources for this section

The preservation argument critics are making

The Authors Guild has publicly disagreed with Alsup's fair use finding on digitization, arguing the ruling overlooks precedent, such as Capitol Records v. ReDigi and Hachette v. Internet Archive, holding that format-shifting a copyrighted work without permission is not automatically transformative and can still cause real market harm. The Guild has said it expects the digitization holding to face appeal.

The narrower, more concrete concern is preservation rather than copyright theory. A mass-market paperback destroyed after scanning is not culturally significant on its own. A rare or out-of-print edition with only a few surviving physical copies is a different case: once every surviving copy has been fed through a cutter and a scanner, there is no way to recover the original object, regardless of how good the digital copy is. Booksellers, not AI companies, have been the ones raising the alarm about specialist and rare titles entering this pipeline.

Anthropic and other labs have not disputed that the scanning is destructive. Their position, backed by the court's reasoning, is that a legally purchased book is the owner's property to convert, and that a searchable digital replacement is a reasonable substitute for a print copy the company already paid for.

Sources for this section

What operators and builders should track next

For any US team sourcing training data, the practical takeaway from Bartz v. Anthropic is narrow but firm: buying a book and destructively scanning it is settled, court-approved practice, while acquiring the same content through piracy is not, and the financial exposure for the latter is now measured in billions of dollars, not a rounding error. Teams evaluating data acquisition vendors should treat lawful chain of custody, receipts, invoices, and proof of purchase, as a hard requirement rather than a nice-to-have.

Ingram, the largest book distributor in the United States, has already warned publishers about this trend and started offering an opt-out mechanism. In the European Union, the Digital Single Market Directive lets rights holders opt out of text and data mining entirely, which means a sourcing strategy that is fully compliant in the US may still create exposure abroad.

The bigger open question is what happens once the current wave of litigation against Meta, OpenAI, and Microsoft plays out. Alsup's reasoning on transformative training has held up in parallel cases so far, but every one of those cases still has to resolve how the underlying content was acquired. That is the fight that determines the actual cost of building a frontier model, not the training run itself.

Sources

Brian Weerasinghe

AI & Technology Researcher

Brian Weerasinghe is the founder and editor of AI Eating The World, where he covers artificial intelligence, tech companies, layoffs, startups, and the future of work. His reporting focuses on how AI is transforming businesses, products, and the global workforce. He writes about major developments across the AI industry, from enterprise adoption and funding trends to the real-world impact of automation and emerging technologies.

Trusted AI LeaderTrusted AI LeaderTrusted AI LeaderTrusted AI Leader
Trusted by 10,000+ builders

The AI brief for people adapting to changes in work

Join readers tracking AI news, workflow shifts, and practical tools they can use to adapt faster.

Free, no spam, unsubscribe anytime.