MPC-lab

Market Prices

Coin Price 24h
BTC Bitcoin
$64,628.7 +2.01%
ETH Ethereum
$1,918.92 +2.23%
SOL Solana
$74.02 +1.11%
BNB BNB Chain
$572.8 +1.17%
XRP XRP Ledger
$1.09 +3.16%
DOGE Dogecoin
$0.0707 +0.86%
ADA Cardano
$0.1638 +4.26%
AVAX Avalanche
$6.42 -0.50%
DOT Polkadot
$0.7644 +0.17%
LINK Chainlink
$8.45 +1.71%

Fear & Greed

29

Fear

Market Sentiment

Event Calendar

{{年份}}
30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

28
03
unlock Arbitrum Token Unlock

92 million ARB released

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

18
03
unlock Sui Token Unlock

Team and early investor shares released

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

12
05
halving BCH Halving

Block reward halving event

Altseason Index

43

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

Market Cap

All →
1
Bitcoin
BTC
$64,628.7
1
Ethereum
ETH
$1,918.92
1
Solana
SOL
$74.02
1
BNB Chain
BNB
$572.8
1
XRP Ledger
XRP
$1.09
1
Dogecoin
DOGE
$0.0707
1
Cardano
ADA
$0.1638
1
Avalanche
AVAX
$6.42
1
Polkadot
DOT
$0.7644
1
Chainlink
LINK
$8.45

🐋 Whale Tracker

🔴
0x4e61...91f7
12m ago
Out
1,224,247 USDT
🔵
0x8716...889b
3h ago
Stake
4,288 ETH
🔵
0x4d61...96a0
1d ago
Stake
1,152,748 USDT

💡 Smart Money

0x4619...e2e9
Early Investor
+$1.2M
94%
0xe251...34a5
Arbitrage Bot
+$2.5M
92%
0x951e...2756
Market Maker
+$1.8M
93%

🧮 Tools

All →
Stablecoins

The Great Book Burn: How AI Companies Are Destroying Physical Libraries for Training Data

Ansemtoshi

We didn't anticipate that the most controversial data source in 2026 would not be scraped from the dark web or generated by rogue models, but purchased from used bookstores and then fed into industrial shredders. Over the past year, millions of physical books have been systematically destroyed not for pulping or recycling, but for their text — carefully scanned, OCR'd, and then discarded. The reason? AI training data that is legally pristine, untouched by AI-generated noise, and protected by a recent court ruling that makes this destructive act a form of 'fair use.'

I’ve spent years inside DAO governance, where we debate the provenance of every vote and every transaction. Trust is built on verifiable history. But what happens when the source of truth for an entire generation of AI models is a pile of ashes? This isn't a thought experiment. Anthropic has already spent millions acquiring and destroying hundreds of thousands of physical books through a service called ISBNdb. The model: buy the book, cut off the binding, scan every page, then discard the paper copy. The legal cover comes from a 2025 U.S. court decision that says as long as the digital copy replaces the physical one one-for-one and is not distributed, it qualifies as fair use.

The Context: Data Hunger Meets a Legal Loophole

AI companies have an insatiable appetite for high-quality, human-generated text. Web-crawled data is increasingly polluted by AI-written content, poisoned datasets, and copyright lawsuits. Physical books, especially those published before the AI era (roughly 2022), offer a clean signal. They are written by humans, edited by humans, and contain no digital watermarking. The problem is obtaining them legally without triggering copyright infringement.

Enter ISBNdb, a company that positions itself as a data compliance broker. They purchase books in bulk — from publishers' remaindered stock, library deaccessions, and used bookstores — then perform what they call "destructive scanning." The books are transformed into digital files under strict non-disclosure agreements, and the physical copies are verifiably destroyed. The result is a dataset that an AI company can claim is both legally obtained (they bought the book) and non-infringing (the original is gone, so the digital copy is a replacement, not a copy). The court's reasoning in the 2025 case, which involved Anthropic, was that this one-to-one conversion does not harm the market for the original work because the digital version is not made available to the public.

The Core: Technical Analysis of a High-Stakes Data Pipeline

From a technical standpoint, the process is brutally efficient. High-speed scanners can process thousands of pages per hour. The resulting files are stored in cloud object storage — likely Amazon S3 or Google Cloud Storage — with each book generating between 50 and 200 MB of PDF or TIFF images. For millions of books, that means petabytes of storage. The data is then run through OCR pipelines to extract machine-readable text, often with manual quality checks for layout errors or damaged pages. The destroyed physical books are either shredded and sent to landfills or incinerated, adding a carbon footprint many would prefer to ignore.

But the real innovation isn't in the scanning technology — it's in the economic and legal engineering. ISBNdb charges a premium for books that are rare, out of print, or from specific subject areas. They claim to offer "attestable destruction" with legal protections. For AI companies, this creates a moat: a unique dataset that no competitor can replicate because the physical source no longer exists. It's a form of data scarcity manufactured by destruction.

Yet the fragility of this pipeline is staggering. Based on my experience auditing decentralized protocols, I recognize that any system relying on a single legal ruling and a limited physical resource is inherently unstable. The court decision could be overturned on appeal. The supply of rare books is finite. And the storage costs for hundreds of petabytes of data are not trivial. Furthermore, the digital files are themselves at risk: a single cloud outage, a ransomware attack, or a metadata error could render years of work useless.

The Contrarian: This Strategy May Undermine the AI It's Supposed to Help

Here's the counter-intuitive angle: destroying physical books to obtain "clean" training data might actually harm the long-term capabilities of AI models. Why? Because books are not neutral. They contain biases of their time — outdated science, cultural stereotypes, narrow perspectives from primarily Western authors. A model trained exclusively on pre-2022 physical books will lack understanding of the digital-native world: social media dynamics, platform economies, modern slang. It will be a historian, not a contemporary thinker.

Moreover, the cultural loss is irreversible. While ISBNdb claims to avoid destroying "rare, unique, or near-extinct" books, there is no public list of what they have destroyed. The burden of proof lies on those who suspect a loss. But the very nature of destruction means we cannot know what we've lost until it's gone. I recall a conversation with a librarian friend who said, "Every book is a unique artifact — even a mass-market paperback has marginalia, a specific printing history, a journey through hands." The court's ruling only considered the "protected expression" (the text), ignoring the physical object's cultural and historical value.

There's also a reputational risk that many investors underestimate. Anthropic is betting billions on public trust. Being associated with book burning — even for the noble goal of advancing AI — is a public relations nightmare. Already, social media is calling it "the Data Holocaust." This could trigger regulatory backlash, especially in Europe where the AI Act demands transparency in training data sourcing. Destroying evidence of your data sources is the opposite of transparency.

Freedom isn't the absence of laws; it's the presence of consent. The authors of those books never consented to their work being used to train an AI. The one-to-one replacement argument treats the physical book as a fungible container, ignoring the intellectual property embedded within. This is a loophole, not a principle.

Takeaway: Are We Building Intelligence on a Foundation of Ashes?

The race for pristine training data has pushed AI companies to extremes. Destructive scanning is the logical endpoint of a market that values data purity over everything else — even over the physical artifacts of human culture. But the very act of destruction creates a new kind of liability: the model's knowledge becomes untraceable, its biases unaccountable, and its future limited by the narrow window of books we chose to destroy rather than preserve.

Liquidity isn't just capital; it's the ability to convert assets without destroying them. In this case, we are converting physical books into digital files by eliminating their physical existence. That's not liquidity — it's consumption. And once consumed, that source of knowledge can never be replenished. As we push forward into an AI-driven future, we must ask: What are we willing to burn today for the models of tomorrow? And who will hold the match?

The Great Book Burn: How AI Companies Are Destroying Physical Libraries for Training Data