Emerald Icon

Emerald Pages

AI training data concept

Photo: Wikipedia

It sounds like the plot of a dystopian novel: a tech company buys hundreds of thousands of used books, slices their spines off with industrial machinery, feeds the pages through high-speed scanners, and then shreds the paper remains. Yet, this is not fiction. It is the current reality of the artificial intelligence industry, a practice exposed through unsealed court documents and investigative reports. This frantic, expensive, and culturally destructive strategy is a testament to the crisis gripping the AI world: the internet has run out of clean data.

The revelations, broken by outlets like The Washington Post and 404 Media, detail a process known as "destructive scanning." Desperate to avoid the "AI slop" and legal landmines of the open web, companies like Anthropic have turned to physical books. They are using hydraulic cutters to destroy books from the last century to train the next generation of chatbots. While the strategy is legally defensible, it is a short-sighted gamble that raises profound questions about the future of information, culture, and the very nature of intelligence.

The primary driver for this destruction is a mathematical reality. The massive datasets that powered previous models—scraped from Reddit, Wikipedia, and millions of blogs—are now contaminated. Researchers call this "model collapse," a phenomenon where an AI trained on its own output degrades into gibberish. Physical books printed before 2022 are the last refuge of 100% human-edited, fact-checked, and logically structured text. But as many have pointed out, the math simply doesn't work. With only 130 million unique books ever published, yielding roughly 30 to 40 trillion tokens, this cache is a drop in the bucket compared to the 100 trillion tokens needed to build the next generation of models.

The Mechanics of a Data Heist

How exactly do you acquire and destroy millions of books without public outcry? The answer lies in a sophisticated, covert supply chain. AI companies are not walking into local bookstores. They are hiring logistics veterans, like Anthropic's hire of former Google Books executive Tom Turvey, to lead initiatives like "Project Panama." These executives use their insider knowledge to bypass traditional publishers, who initially refused to license their digital catalogs.

Instead, they target the secondary market. Tech firms are contracting directly with industrial-scale wholesalers like Better World Books, buying tens of thousands of donated library books, textbooks, and out-of-print titles by the pallet. To hide their tracks, they use third-party vendors to setup hundreds of buyer accounts on platforms like Alibris and Biblio, quietly siphoning rare and obscure texts. The ultimate shortcut is platforms like ISBNdb, which facilitates anonymous, million-book orders complete with strict NDAs to keep the process quiet.

  • The Data Wall: AI companies have already scraped the entire public internet. The remaining text is either low-quality or AI-generated.
  • The Legal Escape Hatch: A June 2025 federal ruling stated that scanning legally purchased physical media is "fair use" if the original is destroyed, making this the only legal way to get high-quality data.
  • Synthetic Data Bridge: The goal is to use books to perfect "seed" models that can then generate flawless synthetic data, theoretically ending the reliance on human-generated text forever.

An Idiotic Strategy or a Necessary Evil?

From a long-term perspective, destructive scanning is a dead-end. It doesn't scale. It is incredibly expensive. And it represents an irreversible act of cultural vandalism. Every time a rare, low-circulation historical text is put through the hydraulic cutter, a piece of physical history is gone forever—not to preserve the text in a public library, but to teach a chatbot how to write a better email. It is a strategy that appears, to many, to be utterly idiotic.

However, from the boardrooms of Silicon Valley, this is a rational, if desperate, response to a crisis. The industry is facing a massive data shortage. They cannot train on pirated "shadow libraries" without facing multi-billion dollar lawsuits (Anthropic settled one for $1.5 billion). The June 2025 court ruling made destructive scanning their only legal loophole. It is a way to buy time, to inject their models with a final dose of pure human logic to bridge the gap until synthetic data generation becomes viable.

Ultimately, the strategy is a reflection of an industry that has painted itself into a corner. The relentless pursuit of larger models has exhausted the planet's supply of human text. Now, companies are willing to shred the very libraries that preserve our history to sustain their growth for a few more months. It is a frantic, last-ditch effort to keep scaling laws alive, revealing that the "intelligence" we are building is, at its core, a cannibalistic engine consuming its own past.

No Ads. By Us. For Us.

This article was made possible by readers like you. We hope it inspired you to support Emerald Book, so we can continue producing content like this.

We will never show you ads, sell your data, or require a subscription to consume our content. Your gift helps us keep the truth accessible.

Click the Support button to give a gift of any amount today.

Thank you for making this work possible.

Emerald Pages is a publication of
Emerald Book, Inc.

Follow us
Share
Scroll to Top