The global AI industry is facing a data liquidity crisis. Clean, human-generated text is becoming a scarce resource. Web data is poisoned by AI-generated content. Copyright lawsuits loom. The solution? Burn physical books. Destroy them. Turn paper into ash. And claim the digital copies as fair use.
This is not dystopian fiction. It is happening now. Anthropic, the AI company behind Claude, spent millions of dollars purchasing millions of physical books. They hired a service provider to rip off bindings, cut pages, scan them, and then discard the original paper. The result: a clean, exclusive dataset for training. Legally defensible, according to a 2025 court ruling. The ruling said that converting a legally owned physical book into a non-distributable digital copy, then destroying the original, qualifies as fair use. One-for-one replacement.
Context: The Data Scarcity Paradox
We live in an age of information abundance. The internet contains trillions of tokens. Yet AI companies are desperate for data. Why? Because the web is polluted. AI-generated text floods forums, blogs, and social media. Data poisoning attacks are real. And copyright holders are suing. The result is that the highest-quality, lowest-noise human text is increasingly found in physical books published before 2022. These books have not been contaminated by AI output. They are pure. They are also finite.
A new market has emerged. ISBNdb, a company that provides book data and scanning services, now offers "destructive scanning" as a product. They will buy books by ISBN, filter by topic or publication year, scan them with high-resolution cameras, and then shred the originals. They advertise legally binding NDAs and verifiable destruction. Their marketing copy explicitly states that pre-2022 physical books are attractive because they have less exposure to AI-generated text and data poisoning. This is a data sourcing strategy built on legal arbitrage.
Core: Data Sourcing as Macro Asset Play
Let me frame this through the lens I use for all crypto macro analysis: liquidity first, incentives second, centralization is the inevitable entropy of scale.
This book-burning strategy is not about data quality. It is about exclusive access to a finite resource. Physical books are a non-renewable asset class. Once destroyed, they are gone. The AI company that buys them gains a permanent competitive advantage in data quality. This is analogous to a token burn mechanism in crypto: by removing supply from the market, you create scarcity. Here, the scarcity is in clean, copyright-safe training data.
From my work auditing ICO liquidity in 2017, I saw how unsustainable tokenomics could create short-term advantages that later collapsed. The same applies here. Anthropic is spending millions. But is this scalable? The total stock of relevant physical books is limited. There are only so many pre-2022 non-fiction, high-quality books in circulation. And every company wants them. This creates a bidding war. Prices will rise. The marginal cost per token will increase.
I also draw on my 2020 analysis of DeFi yield fragility. At that time, I predicted that unsustainable yield farming incentives would lead to a 70% drop in APYs. The same logic applies here: the current legal safe harbor for destructive scanning is a temporary yield. Courts can reverse. Legislatures can act. Public opinion can turn. The moment that happens, the entire value proposition collapses. The books are already burned. The digital copies may become orphaned.
Centralization is the inevitable entropy of scale. This data sourcing strategy centralizes control in the hands of a few well-capitalized AI companies. They can afford to buy and destroy. They can afford the legal teams. They can afford the reputational risk. Smaller players cannot. This creates a moat. But moats built on burning cultural artifacts are not defensible—they are inflammatory.
Contrarian: The Decoupling Fallacy
The prevailing narrative is that destructive scanning is a brilliant workaround for a systemic problem. It decouples AI training from the polluted web. It provides clean data. It respects copyright through the one-for-one ruling. This narrative is seductive.
But it is wrong. The real driver is not data quality. It is legal arbitrage. The one-for-one logic is a fragile fiction. Digital copies are not physically unique. They can be copied infinitely. The court's reasoning relies on a strict factual scenario: the original is destroyed, and only one digital copy is kept. But in practice, multiple copies exist during scanning, processing, and backup. The chain of custody is impossible to guarantee. Future courts may reject this reasoning. And even if they don't, the cultural damage is irreversible.
During the 2022 Terra/Luna collapse, I mapped contagion risk across centralized exchanges. I saw how a single point of failure could trigger systemic collapse. The same applies here. The entire market for destructive scanning rests on a single legal precedent. If that precedent falls, the data supply chain for these AI models is severed retroactively. The digital copies may be deemed infringing. Companies like Anthropic could face billions in damages.
Takeaway: The Cycle Is Early—Position for the Rotation
Right now, the market is in the euphoria phase of this data sourcing cycle. The court ruling is fresh. The books are burning. The press is fascinated. But the smart money is already looking at alternatives: long-term licensing agreements with publishers, synthetic data generation, and cooperative data pools. These are the stablecoins of the AI data world—boring, but sustainable.
I predict that within 18 months, destructive scanning will face a regulatory crackdown or a high-profile lawsuit. The reputational damage will be severe. The companies that pivoted early to sustainable data partnerships will outperform. The rest will be left holding ash.
Centralization is the inevitable entropy of scale. Centralization is the inevitable entropy of scale. Centralization is the inevitable entropy of scale.
The signal is clear: burn books, burn bridges. Build networks, build trust.