Emerald Pages
◆
From Selling to Erasing Books: Why Amazon Is Destroying Original Literature
The company built on selling books is now destroying them. A new investigation reveals Amazon's industrial-scale operation to slice, scan, and shred rare literature for AI training data—exposing the tech industry's desperate, multi-front panic over running out of human-written text.
Photo: Wikipedia
Air-tight denials. Corporate euphemisms. And a secret facility in Las Vegas where the spines of rare books are sliced off at industrial speed. This is the shadowy new front in the AI arms race, and Amazon—the company that began in 1994 as the world's largest online bookstore—is now one of its most aggressive soldiers. The company that built its trillion-dollar empire on the backs of authors, publishers, and independent book sales is now systematically destroying the very artifacts that made it possible.
For years, reports have circulated about AI companies quietly buying up out-of-print and antique books to train their large language models. The public outcry focused on names like Anthropic, whose internal "Project Panama" leaked during copyright litigation, revealing a massive, secretive operation to purchase, slice, and scan millions of books. Throughout this, Amazon maintained a façade of innocence, publicly denying any involvement in such destructive practices.
The truth, however, was far more sinister. An investigative report by 404 Media definitively exposed Amazon's secret operation using an Apple AirTag hidden inside a 1,000-book bulk order from an independent bookstore. The tracker traveled directly to a facility in Las Vegas, Nevada, coded internally as VGT3. Employees at the facility confirmed their sole job is to strip the bindings from incoming books and feed the loose pages through high-speed sheet-feed scanners, destroying the physical copies in the process. Amazon had been lying.
The Data Wall and the Desperate Hunt for Pristine Text
Why would a trillion-dollar tech giant, which has access to a vast digital catalog through its Kindle store and its "Search Inside the Book" feature, resort to such destruction? The answer lies in a crisis facing the entire AI industry: they have hit the "data wall." Tech companies have essentially run out of high-quality, human-written text on the internet to train their models.
Researchers at institutions like Epoch AI estimate the total pool of public, high-quality human text on the web is roughly 300 trillion tokens, and frontier AI models are projected to fully exhaust this supply. This desperate shortage is the exact reason Amazon's VGT3 facility is in such high demand. In fact, leaked employee forum posts revealed a panic earlier this year when the warehouse temporarily ran completely out of physical books to scan, leaving workers worried the facility would shut down until fresh bulk shipments arrived.
The race to destroy physical books is driven by three major internet data crises that push tech companies toward physical destruction:
- The Internet is Full of "AI Slop": A massive percentage of modern websites, blogs, and social feeds are now clogged with AI-generated text. If an AI model is trained on text generated by other AI chatbots, its intelligence degrades rapidly, leading to a catastrophic technical failure known as "model collapse." Physical books published before 2022 are considered premium data because they are guaranteed to be 100% written, edited, and fact-checked by real humans, entirely untainted by machine language.
- The Open Internet is Being Locked Down: Major news sites, publishers, and platforms have built digital walls to block AI web crawlers. To legally get internet text now, tech companies must pay exorbitant fees—such as Amazon's licensing deal with The New York Times, which costs an estimated $20 million to $25 million per year. Buying a $40 used book from an independent store and shredding it is drastically cheaper.
- The Shift to the "Shadow Checklist": Because the internet has been completely drained, AI companies are trying to build an absolute archive of human thought by systematically digitizing everything that never made it online. Booksellers report that Amazon's automated proxy buyers aren't hunting for specific, famous stories; they are mechanically buying lower-cost, obscure titles with unique ISBN serial numbers. They are effectively trying to check off every single ISBN number in publishing history—including 1970s academic texts, regional folklore, and mid-century technical manuals—to ensure their corporate models have access to specialized knowledge that the internet simply doesn't possess.
The severity of this data drought is a clear admission that the bottleneck for modern AI progress is no longer computing power or chips—it is the availability of original human thought.
Amazon's Full-Circle Irony
The company's actions represent a complete subversion of its founding identity. When Jeff Bezos started Amazon in 1994, he chose books for several practical, strategic reasons. In the mid-1990s, the largest physical bookstores could only hold about 150,000 titles. However, there were over 3 million books in print worldwide. Bezos realized a website could list millions of titles simultaneously, creating a store that was physically impossible to replicate in the real world. Books were also highly durable, easy to ship, and every published book already had a universal identifier (the ISBN number), making indexing and tracking incredibly simple.
Amazon sold its very first book out of Bezos's garage in July 1995—a data science textbook titled Fluid Concepts and Creative Analogies. The company grew rapidly, shipping books to all 50 U.S. states and 45 different countries within its first month of operation. By the late 1990s, Amazon began expanding into music CDs, VHS tapes, electronics, toys, and home goods, eventually becoming the "Everything Store."
For nearly three decades, Amazon positioned itself as a champion of literacy and information accessibility—first by making physical books easier to buy, then by dominating the digital space with the Kindle, and later by preserving out-of-print titles through its own publishing lines. The recent revelations that Amazon is now purchasing unique, physical, 20th-century books to clip their spines, scan them, and throw them into industrial shredders marks a complete subversion of its original corporate identity. The company that built its empire by delivering books to people's doors is now intercepting scarce literature to destroy it for server data.
The practice has sparked massive ethical outrage among booksellers, librarians, and historians. They point out that Amazon's current practices go beyond mere corporate hypocrisy—they actively threaten the historical record. When a traditional library or university digitizes a rare book, they use expensive, specialized, non-destructive overhead scanners. They turn the pages manually to preserve the binding, the paper, and the physical artifact for future generations. Amazon's approach prioritizing speed over preservation means that if a book is the last remaining physical copy of a specific printing, its destruction means that specific physical history is lost forever. Critics have labeled the industrial destruction of scarce literature as "evil incarnate."
The Legal Loophole and The Corporate AI Empire
When confronted with the AirTag evidence, Amazon's official response was a masterclass in deflection. A spokesperson issued a statement to Inc. stating that the company "purchases books through commercial channels to help develop and improve the products and services our customers use." The company refused to answer how many books have been destroyed or whether they filter out highly scarce items before slicing them.
The "customers" Amazon refers to are not book lovers; they are corporate clients using Amazon's proprietary Large Language Models (LLMs). Behind the scenes, Amazon has aggressively developed, launched, and continuously upgraded its own first-party AI families:
- Amazon Nova (The Current Flagship): Launched as Amazon's premier, multi-billion-dollar proprietary AI generation, the Nova family consists of specialized models including Nova Micro (ultra-fast text-only), Nova Lite & Pro (multimodal systems for charts, documents, images, and videos), and Nova Premier (Amazon's heavyweight "teacher" model for complex, deep-reasoning enterprise tasks). The revamped, highly conversational Alexa+ is powered directly by the Nova series.
- "Olympus" (The Multi-Trillion Parameter Engine): During its development phase, Amazon's high-stakes frontier model was heavily reported on under the internal code name Olympus. Sporting a massive 2-trillion-parameter architecture designed by Amazon's AGI team under Rohit Prasad (the former head of Alexa), this research initiative served as the foundational bedrock that was ultimately refined and productized into the top-performing versions of the Nova family.
- Amazon Titan (The Legacy Tier): Before the creation of Nova, Amazon's original baseline models were called Amazon Titan. While Titan models are still active and used by enterprise clients for simpler corporate automation, data search, and summarization tasks, they have been largely succeeded by the Nova generation for cutting-edge AI capabilities.
Amazon deliberately chose not to launch a viral, consumer-facing chatbot web interface like ChatGPT or Google Gemini. Instead, they focused entirely on selling their LLMs directly to corporate developers and enterprise clients through AWS Bedrock. Because everyday web users cannot easily navigate to an "Amazon Chat" website to experiment with their models, the public widely assumes Amazon simply bypassed building the underlying tech entirely. The fact that Amazon is quietly operating an intense, industrial-scale book scanning facility proves they are neck-and-neck in the foundational data race. They aren't just a landlord for other companies' AI; they are aggressively training their own.
The Broader "Data Grab" and Industry Reaction
This book-burning phenomenon is just one facet of a massive, multi-front panic over training material. Tech giants have hit a "data wall," meaning they have exhausted the high-quality open web and are aggressively harvesting private and offline sources. Google recently paid $10 million to purchase internal business data from bankrupt Spirit Airlines, acquiring roughly 100 million corporate emails and 500 million Microsoft Teams chat logs to train its models. Amazon quietly enrolled all streamers on its subsidiary platform, Twitch, into an automated AI training program, forcing creators to manually opt out if they do not want their streams, clips, and chat logs harvested.
Interestingly, Amazon's exposure has given some rival tech companies a massive public relations advantage. Because Amazon's sweeping bulk orders target rare and out-of-print books without pre-screening them, competitors like xAI have been able to publicly state that they do not engage in the destructive scanning of rare or antique literature.
For booksellers and historians, the discovery means that any anonymous bulk order coming from an online proxy buyer is viewed with intense suspicion, as they can no longer trust whether a tech startup or a tech giant is on the other end of the transaction. Independent bookstores are now fighting back, developing strategies to identify and block these automated corporate buyers to protect their inventory.
The legal justification for this destruction is rooted in a 2025 court precedent (Bartz v. Anthropic), which ruled that buying a physical book, stripping its binding, and scanning it qualifies as "fair use" because the digital copy effectively replaces a physical copy that is no longer in commerce. By using automated proxy buyers to purchase books through "commercial channels," Amazon legally owns the paper and is free to destroy it.
The company that began as a beacon for book lovers has become a dead end for physical literature. In its relentless pursuit of AI dominance, Amazon is erasing the physical artifacts of human knowledge, one ISBN number at a time. The ultimate irony is that a company built on making information accessible is now sacrificing that very information to fuel machines that tech companies believe may make human knowledge itself obsolete.
No Ads. By Us. For Us.
This article was made possible by readers like you. We hope it inspired you to support Emerald Book, so we can continue producing content like this.
We will never show you ads, sell your data, or require a subscription to consume our content. Your gift helps us keep the truth accessible.
Click the Support button to give a gift of any amount today.
Thank you for making this work possible.