Amazon Buys Books in Bulk to Train AI Chatbots, Then Destroys Them

Source

20 August 2026 · 06:00 · Claude (Anthropic) · claude-sonnet-5

An investigation by 404 Media, picked up by VRT NWS, reveals that Amazon is buying up physical books on a massive scale, digitizing them to train AI models, and then destroying the original copies. The practice raises questions about copyright, data collection, and the future of physical books in the AI era.

Amazon books AI training has been in the spotlight this week following a revealing investigation by 404 Media, which was picked up by VRT NWS. The report shows that Amazon is buying physical books in bulk, scanning them to use as training data for its AI chatbots, and then destroying the original books afterward. The news fuels a broader discussion about how big tech companies obtain data to feed their language models, and what price is being paid for it outside the digital world.

What exactly does the investigation reveal?

According to the tracker used by 404 Media, Amazon is buying up secondhand books on a large scale through its own channels and third-party sellers. Those books are not resold or reused, but systematically digitized. The content serves as fuel for the training datasets behind Amazon's AI systems, including the chatbot technology the company deploys in products like Alexa and its cloud platform AWS. After the text has been scanned, the researchers say, the physical copies are destroyed rather than recycled or resold, which has raised questions among many about waste and sustainability. The investigation fits a pattern we've seen for a while now in the history of artificial intelligence: large language models need enormous amounts of text to learn, and books are a particularly valuable source because they tend to be of higher quality and better structured than random web text.

Why are books so valuable to AI companies?

Language models like Amazon's, but also those of competitors such as OpenAI, Google, and Meta, are trained on gigantic text corpora. Books offer long, coherent narratives, complex sentence structures, and a wide range of topics, making them exceptionally well-suited for teaching an AI model language skills and reasoning ability. Where the internet often contains short, fragmented, and sometimes unreliable information, books are editorially vetted and more clearly defined in terms of authorship. Yet that very last point is a sensitive one. Authors and publishers have already filed multiple lawsuits against tech companies in the past over the unauthorized use of copyrighted material for AI training. According to legal experts, scanning and then destroying physical books could be a way of staying within the boundaries of certain legislation, since purchasing a physical copy is subject to different rules in some jurisdictions than digitally copying protected content.

Reactions and social impact

The news has led to outraged reactions, including from book lovers, authors, and environmental organizations. They point out that destroying physical books runs counter to sustainability goals and to the cultural importance of printed works. Others emphasize that the practice illustrates just how hungry the AI industry is for quality data, now that the internet itself is increasingly polluted with AI-generated content and is therefore becoming less reliable as a training source. Amazon is not the only player searching for new data sources. Other major tech companies are also investing heavily in acquiring exclusive datasets, licensing agreements with publishers, and partnerships with libraries. This fits into a broader trend in which AI applications are becoming increasingly dependent on scarce, high-quality data now that the easily available sources are slowly being exhausted.

What does this mean for the future of AI and books?

The revelation opens the door to stricter regulation around data collection for AI training. Policymakers in the European Union and the United States are taking a critical look at how tech companies obtain their training data, and cases like this one could increase pressure to legally enshrine transparency and compensation for creators. At the same time, it shows just how far companies are willing to go to gain a competitive edge in the race for ever more powerful AI models. For consumers and the book industry, this raises fundamental questions: who actually benefits from the knowledge captured in books, and at what cost? As the debate over copyright, data collection, and AI ethics continues to heat up, it seems likely that this won't be the last revelation about how big tech companies feed their models. Anyone wanting to keep up with developments can check out more AI news and our knowledge base for background information on how artificial intelligence works and its impact.

VRT NWSVRT NWS


Source: VRT NWS

Ster Software

The most complete knowledge platform on artificial intelligence.

Kraaienjagersweg 24
7341 PT Beemte Broekland, Netherlands


© 2026 Ster Software BV · Chamber of Commerce 75474913

Content generated by Claude (Anthropic) · model: claude-sonnet-4-6