Hook
Most people think AI training data comes from scraping the open web or licensing digital archives. They’re wrong. The real play is in physical books—bought, scanned, and systematically destroyed. Anthropic just spent millions on millions of dead trees. They’re not building a library. They’re building a data moat. And the market hasn’t priced this yet.
Data doesn’t lie; emotions do. The numbers are stark: millions of physical books, millions of dollars, zero cloud scrapers, zero ethical gray zones—just pure legal arbitrage. This is the kind of structural inefficiency I live for. Let’s break it down.
Context
In 2025, a U.S. court ruled that converting a legally purchased physical book into a non-distributed digital copy, then destroying the original, counts as “fair use” under copyright law. The logic: you’re not creating an extra copy—you’re just changing the format. One-for-one replacement. It’s clean. It’s legal. And it’s a goldmine for AI companies desperate for high-quality, human-generated text uncontaminated by the noise of modern AI-generated garbage.
Enter ISBNdb, a company that offers exactly this: buy any book by ISBN, subject, or year, scan it in a destructive manner (cut the binding, shred the pages), then provide the digital copy to the client with a legally binding NDA and verifiable destruction certificate. They even market it as “pure human text from before the AI era.” Anthropic is their flagship client.
Core
Efficiency eats sentiment for breakfast. Let’s look at the numbers.
Anthropic paid several million dollars for millions of books. That’s about $2-5 per book, including scanning and destruction. Compare that to licensing a single digital book from a publisher: often $10-50 per title, plus restrictions on use. More importantly, the data from physical books is “clean”—no AI-generated paragraphs, no modern data poisoning, no SEO spam. Every sentence was written by a human, edited by a human, and published before the internet went crazy.
But the real alpha is in the scarcity. Physical books are finite. Once you buy and destroy a book, no one else can ever scan that exact copy again. If Anthropic targets rare or out-of-print titles, they create a permanent data moat. Competitors like OpenAI or Google can’t replicate that dataset. This isn’t just data acquisition—it’s a physical-world bottleneck.
The legal framework is the key. The court’s “one-for-one replacement” reasoning gives a green light only if the original is destroyed. That means the digital copy is unique in the sense that no other physical instance of that exact book can be legally converted. ISBNdb’s business model is pure legal arbitrage: they take a commodity (a common book) and transform it into a unique, exclusive digital asset.
Spread the truth, not the panic. But the truth here is that this model creates a massive asymmetric advantage for the first mover. Anthropic has already committed millions. Others will follow. The question is: how many rare books are left?
Contrarian
Most people see this as a cultural tragedy—destroying books for AI. I see it as a market inefficiency being exploited by rational actors. The ethical hand-wringing misses the point: the court already decided that the expression (the text) is preserved; the physical medium is irrelevant. If the goal is to train a better model, and the legal path is clear, why wouldn’t you do it?
The real blind spot is the risk of regulatory backlash. The same court could overturn this precedent. Or Congress could pass a law banning “destructive scanning” for AI training. That would wipe out the value of those digital copies overnight. But in a bear market for data quality, where every scrap of clean text is precious, taking that risk might be rational.
Another angle: the NFT comparison is lazy. When Banksy burned his art to create a digital NFT, he was creating scarcity—only one digital copy existed. In the book case, the digital copy can be infinitely replicated. The “uniqueness” is only in the legal claim that no other physical copy was digitized in the same way. But once it’s copied once, the genie is out of the bottle. The court’s logic relies on the assumption that the data remains non-distributed—but what happens when a model trained on that data is public? That’s the next legal battle.
Takeaway
This is a data quantity game dressed as a quality play. The real takeaway is that AI companies will pay a premium for scarcity, even if that scarcity is legally constructed. The book destruction model is a proof of concept: if it works, expect similar arbitrage in other physical media—maps, manuscripts, even museum pieces.
Code is law; liquidity is life. But in AI, data is the only scarce resource. And these guys are destroying the supply to lock it up.