MPC-lab

Market Prices

Coin Price 24h
BTC Bitcoin
$64,100.4 +0.95%
ETH Ethereum
$1,866.79 +0.62%
SOL Solana
$73.7 +0.70%
BNB BNB Chain
$598.9 +1.58%
XRP XRP Ledger
$1.07 -0.17%
DOGE Dogecoin
$0.0700 -0.10%
ADA Cardano
$0.1919 +0.10%
AVAX Avalanche
$6.66 +0.23%
DOT Polkadot
$0.8586 +3.78%
LINK Chainlink
$8.13 -0.29%

Fear & Greed

27

Fear

Market Sentiment

Event Calendar

{{年份}}
18
03
unlock Sui Token Unlock

Team and early investor shares released

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

12
05
halving BCH Halving

Block reward halving event

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

28
03
unlock Arbitrum Token Unlock

92 million ARB released

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

Altseason Index

43

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

Market Cap

All →
1
Bitcoin
BTC
$64,100.4
1
Ethereum
ETH
$1,866.79
1
Solana
SOL
$73.7
1
BNB Chain
BNB
$598.9
1
XRP Ledger
XRP
$1.07
1
Dogecoin
DOGE
$0.0700
1
Cardano
ADA
$0.1919
1
Avalanche
AVAX
$6.66
1
Polkadot
DOT
$0.8586
1
Chainlink
LINK
$8.13

🐋 Whale Tracker

🟢
0xc22d...dc68
1d ago
In
1,576.07 BTC
🔴
0x84c3...7ba5
6h ago
Out
21,398 SOL
🔵
0xfed2...5786
2m ago
Stake
18,868 BNB

💡 Smart Money

0x593c...8778
Experienced On-chain Trader
+$2.6M
82%
0xd99d...822d
Market Maker
+$4.2M
92%
0xe8f0...1bad
Institutional Custody
+$2.0M
86%

🧮 Tools

All →
Research

The Book Shredders: When AI’s Hunger for Clean Data Turns Libraries Into Fuel

0xKai

We don’t need more users; we need more stewards. But somewhere in a warehouse outside Nashville, a machine is slicing the spine out of a first edition and feeding its pages to a scanner. The scan is saved. The original is shredded. A few hundred miles away, a large language model is learning to sound human from the ghost of a library that no longer exists.

This is not a metaphor. It is the newest arbitrage in the AI supply chain. Anthropic has spent millions of dollars buying millions of physical books, not to read them and not to preserve them, but to destroy them after scanning. ISBNdb, a book data service, now sells “destructive scanning” to AI developers: it will source books by ISBN, subject, and publication year, scan them, and then legally dispose of the originals with a notarized paper trail and a binding NDA. The logic is that if the original book is destroyed, the digital copy is a one-to-one replacement, and therefore fair use. A 2025 court ruling gave that model a legal green light. The result is a data economy that treats physical culture as fuel.

I know how that story ends. In 2017, I spent months auditing the whitepaper of a project called OmniChain. The token distribution heavily favored early investors, and the egalitarian rhetoric was a layer of polish over a pre-committed exit. I wrote a five-thousand-word exposé, and the project rug-pulled a year later. The lesson I carried out of that mess is simple: a well-written legal section is not the same as a fair system. The same is true for a well-reasoned court opinion about one-to-one replacement. It may make the destruction of a library legal. It does not make it just.

Context: The Clean Data Crisis

The AI industry has a dirty-data problem. The web is filling with AI-generated text, and models trained on machine output drift into bland, repetitive, increasingly useless prose. Researchers call it model collapse; lawyers call it copyright litigation; data engineers call it a nightmare. So the industry is going hunting for text that is reliably human. Physical books fit the bill — especially books published before 2022, before the current wave of generative text poisoning. A hardcover bought from a warehouse has no SEO spam, no comment sections, no hallucinated gloss, no watermark from an earlier model. It is pure, dense, carefully edited human expression. The catch is that scanning a book and copying its contents into a training corpus has historically been illegal, and the big AI companies have been sued for less. Then the court changed the math.

The 2025 ruling is narrow, but it is powerful. If you lawfully buy a physical book, scan it, and destroy the original, you have not increased the number of copies in the world. You have merely changed the medium. No net harm to the copyright holder, no flooding of the market, no unfair competition. In that logic, the book is consumed, not copied. That is exactly the kind of clever, bounded reasoning that sounds reasonable in a courtroom and collapses in the real world, because digital copies are not physical objects. The court’s logic assumes a world without packet loss, without leaking hard drives, without a single backup that turns a one-to-one replacement into a one-to-a-million reproduction. And the people running the shredders know it.

Anthropic has made this strategy concrete. It hired the former lead of Google’s book scanning project, it has spent millions on millions of books, and it has a supplier in ISBNdb that openly markets the service. ISBNdb lets AI developers filter by ISBN, subject area, publication year, and more. It promises binding confidentiality agreements and verifiable destruction. It even acknowledges, in its own procurement posts, that headlines about AI companies destroying books raise reputational questions. The reputational questions are not a side effect. They are the business model.

Core: The Part of the Machine No One Wants to Audit

Let’s be precise about what is happening. This is not a copyright breakthrough. It is a gap between two regimes: law and engineering. The law sees a book as a bundle of protected expression; the engineer sees it as a file to be parsed. The one-to-one replacement doctrine is a governance rule, and governance rules are only as good as the data infrastructure underneath them. In a DAO, you would never approve a treasury expense that could not be traced to a transaction on a public ledger. Here, the entire supply chain operates behind NDAs, and the only public fact is the destruction certificate. That is not accountability. It is a receipt.

I have spent the last year auditing DAO treasuries and token distribution models, and I have learned to look for the line item that no one wants to discuss. In this case, the line item is everything that happens after the scan. A single book can produce fifty to two hundred megabytes of raw images. Millions of books mean petabytes of scans, and petabytes need storage, replication, and backup. Before a single token reaches a model, someone has to run OCR, correct misrecognized characters, split chapters, remove duplicates, filter pages with stains and marginalia, and decide whether the marginalia belongs to the book or to the person who once owned it. That is not a small operation. It is a factory, and the factory is the part of the story that the marketing material leaves out. ISBNdb says it will handle the approvals and destruction; it is less noisy about the cleaning, the quality checks, and the archive that quietly holds the files. That is where the cost lives, and it is also where the value is destroyed.

There is also a distribution bias hiding in the shelves. Books that are easy to buy in bulk are not a random sample of human culture. They are the leftovers: remainders, warehouse stock, out-of-print titles that did not sell, popular classics in cheap editions, and whatever a liquidator can source in large quantities. That is not a neutral representation of human thought. It is a representation of the secondary book market, which overweights certain countries, certain languages, certain genres, and certain authors. If you train a model on destroyed remainders, you will get a model that sounds like a remainder table. The court ruled on copyright; it did not rule on epistemology. But the epistemology is the problem.

The competition angle is even more uncomfortable. Physical books are finite. A rare edition, once destroyed, is gone forever. Every copy that is shredded for an AI training run is a copy that cannot be sold to a reader, archived by a library, or consulted by a future researcher. That makes destructive scanning a powerful moat. Anthropic can pay a premium to buy up entire categories of books, destroy them, and thereby prevent competitors from using the same source. It is a fuel war, except the fuel is the written record of human civilization. And because the court has blessed the one-to-one replacement logic, there is no legal penalty for being the most aggressive buyer. The only constraint is capital. In a market where Anthropic is valued in the hundreds of billions, millions of dollars for millions of books is a rounding error. The real cost is the cultural loss, and that cost is not on anyone’s balance sheet.

This is the same pattern I watched in crypto after the Bitcoin ETF approval. A technology designed for peer-to-peer exchange became a Wall Street custody instrument, and an asset class was born. Now the same institutional logic is treating literature as a strategic mineral. The buyer does not care about the author’s intent; it cares about the token yield of clean text. The seller does not care about the reader; it cares about the liquidation price. The legal system does not care about the cultural object; it cares about the copy count. And the public is left with a narrative problem: no one knows the titles of the rare books that have already been destroyed, because the records are sealed by confidentiality agreements. The absence of evidence is not evidence of absence. It is the whole point.

Core: The Gray Zone the Court Left Open

Even within the court’s own logic, there is an unresolved fault line. The ruling blesses the conversion of lawfully purchased books into non-distributed digital library copies, as long as the originals are destroyed. But it does not answer what happens next. A language model is not a static library. A language model is a distribution mechanism. When a model is trained on the scanned text and released to millions of users, the text is not locked in a vault; it is baked into the weights. The weights are distributed. And models memorize. Ask a large model about a well-known passage, and it can reproduce it almost verbatim. That means the digital copy is not contained. It is leaking through every API call, every chat response, every generated summary.

The court did not rule on that. It could not, because the case was about the act of scanning, not about the act of generation. But this is where the one-to-one replacement fiction becomes dangerous. The physical book is gone, so the copyright holder has one fewer copy in the world. Yet the model itself may have become thousands of copies, in the sense that matters to a publisher: a source that can be queried, reproduced, and paraphrased. The legal theory treats the scan as the final copy. The engineering reality treats the scan as the first copy. That gap will not remain empty forever.

There is also a claim still hanging over Anthropic’s data procurement practices. The court did not grant summary judgment on the part of the case involving copies allegedly taken from a central library’s pirated archive. That is not a footnote; it is a sword. If that claim survives, the entire supply chain is exposed, because the same laboratory processes that destroyed lawful books may also have drawn from unlawful sources. In my years auditing token distributions and treasury flows, the projects that failed were almost never the ones with a single obvious crime. They were the ones with a culture of cutting corners, where the first corner cut made the second one feel routine. The same pattern may be playing out at the intersection of AI and physical books, and no destruction certificate can prove otherwise.

Core: The Open Source Gap

The most underreported consequence is what this means for open source AI. Open models have always depended on public datasets like Common Crawl, Books3, and the Internet Archive. Those datasets are increasingly polluted with machine-generated content, and several have been pulled from distribution because of copyright challenges. Meanwhile, closed labs can pay to destroy physical books and obtain a clean, exclusive corpus that no open researcher can legally access. That is not a technical edge. It is an extractive edge. It means the gap between open and closed models will become a gap between open data and destroyed data.

We have seen this movie before in finance. After the ETF approval, Bitcoin became a toy for Wall Street custody desks, and the peer-to-peer cash vision that defined the early protocol was quietly retired. The same logic is now classifying books. A book is no longer a thing to read; it is a feedstock with a purity score. The buyer’s goal is not to preserve the text; it is to own a source of text that no competitor can access. That is the opposite of a public library. It is a private quarry.

The open source community will respond, as it always does, by building synthetic datasets and better filtering tools. But synthetic data cannot replace the dense, edited, human-shaped prose that lives in physical books. You cannot synthesize your way out of a hole you dug by destroying the source material. The only honest response is a governance response: a visible record of what exists, what has been scanned, and what has been lost.

Core: What a Ledger Would Change

I keep coming back to my own side of the industry because we spent the last decade building exactly the wrong answer to this problem. We built Web3 on a promise that data could be owned, traced, and governed by its creators. Most of that promise has not been kept. We turned Bitcoin into a Wall Street custody toy, we turned NFTs into speculative receipts, and we turned DAOs into group chats with a token. But the underlying instinct was correct: trust is the only protocol that cannot be coded, so we need the next best thing, which is verifiable, decentralized provenance.

The book-shredding model is a perfect proof case for why data provenance is not optional. Right now, a buyer can destroy millions of books and give you an NDA saying it happened. There is no public ledger of destroyed ISBNs. There is no registry of who scanned what, into whose hands the digital copies went, or how many backups exist. There is only a legal paper trail, and a legal paper trail is not the same as a fact.

A blockchain-based registry could change that. Before a book is destroyed, its ISBN, edition, condition, and provenance could be recorded. After the scan, a cryptographic hash of the digital file could be anchored. That would not prevent the destruction, but it would make the destruction visible. It would let journalists, historians, and regulators see what is being consumed. It would turn a secret supply chain into an accountable one. It would also create a market for preservation: if a rare book is listed on a public registry, a library, a university, or a community DAO could bid to keep it alive. Right now, there is no mechanism for that. The only bids that matter come from AI companies with the deepest pockets.

This is where privacy-preserving compliance becomes relevant. The answer is not to ban destructive scanning. The answer is to make it impossible to do in the dark. KYC requirements and regulatory compliance are not the enemies of decentralization; they are the conditions under which a market earns trust. The same is true for data provenance. We do not need to stop the sale of physical books to AI labs. We need to record the sale, the scan, and the destruction on a public ledger, with the exact same rigor we expect from a DAO treasury. If a protocol cannot explain where its assets came from, we call it a governance failure. If a model cannot explain where its training text came from, we should call it a data governance failure.

In 2022, after Terra collapsed, I retreated to a cabin in Yilan for three months and wrote about the soul of the ledger. The phrase sounded romantic then. It carries a different weight now. A ledger can record a hash, but it cannot record a binding, a bookplate, a marginal annotation from a stranger who read the same page sixty years ago. Those are exactly the things the one-to-one replacement logic discards. The text survives. The object does not. And the object is where memory lives.

When I helped founders structure DAOs in The Alignment Circle, the lesson was always the same: if you design a process that rewards the appearance of compliance, you will get people who shred the evidence. ISBNdb’s model is a governance process designed to reward destruction. It is not malicious. It is rational. It is the outcome of a legal system that has no unit of account for cultural memory. The only way to fix that is to create a unit of account: a public record that says a particular book existed, was scanned, and is now no longer available to the world. That record is not a sentiment. It is an infrastructure choice.

The Banksy Trap

Some will try to soften this story by comparing it to Banksy burning one of his own artworks and issuing an NFT. The comparison is tempting, but it is wrong. In the Banksy case, the physical object was destroyed and its value was transferred to a unique digital token. The scarcity continued. In the case of an AI training run, the digital copy is not unique. It is a file on a hard drive, with backups and replicas and model weights that spread it across the world. Destroying the book does not create digital scarcity; it creates a legal fiction. The rare thing is not the copy. The rare thing is the certificate that says the copy was made lawfully. That is not art. That is bookkeeping.

It is also not a cultural preservation strategy. A training run is not an archive. An archive preserves access; a training run consumes its source. The model may be able to generate text that resembles the book, but it cannot offer the book to a reader, cannot show the original pagination, cannot represent the physical object. The one-to-one replacement doctrine treats the text as the entire value of the book. That is a category error, and it is the category error on which an entire industry is being built.

Contrarian: The Object Was Already a Copy

I want to be careful here, because the easy take is to condemn the shredders and move on. But the contrarian truth is more uncomfortable: the problem is not that the books are destroyed. It is that the exchange is treated as equivalent. The court’s one-to-one replacement logic treats a book as a container of text. If the text survives, the book has survived. But anyone who has held an annotated copy of a poem knows that is not true. The binding, the paper, the marginalia, the library stamps, the inscription from a stranger who read the same page decades ago — that is the book. The text is only the score; the object is the performance. When a machine destroys the object, it destroys part of the performance. That is not a copyright harm. It is a cultural harm, and the courts have no word for it.

There is another uncomfortable layer. Most physical books, especially remainders and warehouse stock, are not unique. They are mass-produced copies, and the information they contain is already replicated in other libraries, other archives, other formats. Destroying a warehouse full of unsold paperbacks is not the same as destroying the last illuminated manuscript of a medieval poem. The law has no way to distinguish between those two acts because the law only regulates copies, not meaning. But the deeper the AI industry digs into physical archives, the more likely it becomes that the next purchases will target the rare titles. The confidentiality agreements mean we will not know until it is too late. That is why the provenance ledger must exist before the next round of buying, not after.

We often say in this community that we built not for the peak, but for the valley. The valley is where stewardship is tested. Right now, in the valley of the AI data crisis, we are being tested, and many of us are failing. The book shredders are not villains; they are symptoms of an economic system that has no accounting line for collective memory. We can build a better ledger, but a ledger cannot decide what is worth keeping. That is a human decision, and we keep outsourcing it to the highest bidder.

Takeaway: The Provenance Officer Is Coming

In five years, every serious AI company will have a data provenance officer. That person will be asked to certify where the training data came from, whether the source was authorized, and what the supply chain cost. The question is not whether that role will exist. The question is whether it will report to legal or to ethics. If it reports to legal, it will produce more NDAs, more shredding certificates, and more precisely worded loopholes. If it reports to ethics, it will build what this moment demands: public registries of destroyed books, auditable hashes of scanned files, and a clear rule that no private training run may erase the last copy of a cultural artifact.

The book-shredding story is not a story about copyright. It is a story about stewardship. We don’t need more users; we need more stewards. And if we want models that actually reflect humanity, we need to start by recording what we are willing to lose. Trust is the only protocol that cannot be coded — but transparency is the closest thing we can build.

This sounds like an ending, but it is really a beginning. The next few years will decide whether the written record of human culture becomes a public commons or a private fuel reserve. The answer will not be written in a court opinion. It will be written in the choices we make about visibility, ownership, and consent. The books are already on the conveyor belt. The only question is whether we keep the receipt.