Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Amazon, once an online bookseller, is destroying rare books to train AI models

Amazon is buying large numbers of rare books, cutting off their spines, and scanning them for AI training, according to an investigation by 404 Media that tracked a book to an Amazon facility in Las Vegas. Amazon said it purchases books through commercial channels to improve the products and services customers use.

Published
亚马逊被曝大量收购稀有书籍、切脊扫描用于AI训练
Image source: techcrunch.com

Amazon, the company that began as an online bookseller, is now buying large quantities of rare books, cutting off their spines and scanning them to train AI models, according to a TechCrunch report based on an investigation by 404 Media.

404 Media said it placed a tracking device inside a rare book and traced it to an Amazon facility in Las Vegas known as VGT3, which identifies itself with a symbol of a dinosaur holding a book in its claws.

In a statement, Amazon told 404 Media that it "purchases books through commercial channels to improve the products and services customers use."

The report explains the motivation: companies like Amazon need unfathomably large amounts of text to train their LLMs, and models have already ingested much of what is available on the internet — Anthropic, for example, was accused of illegally pirating books.

Rare books, especially ones that are out of print or impossible to find on the internet, offer a new source of coveted training data. These texts are especially valuable since there's no chance that anything published before 2022 was written by an LLM.

When LLMs train on AI-generated text, they risk "model collapse," a degradation in output quality — which makes human-written rare books an increasingly attractive data source for model builders.

The revelations sharpen the ongoing fight over training data and copyright, a battle already playing out in lawsuits from publishers and authors. What remains to be seen is how Amazon responds to the tracking-device findings, whether the book industry pushes back, and whether regulators take a closer look at large-scale scanning of physical books.

Why it matters

AI labs desperate for fresh human-written text are now reaching into the physical rare-book market; if confirmed, Amazon's large-scale scanning could fuel a new wave of training-data copyright disputes.

AmazonAI Training DataCopyright
Back to realtime news

Nearby Updates

All

08/18, 00:53

Claude to start watermarking AI-generated text, The Guardian reports

Anthropic's Claude assistant is set to start watermarking AI-generated text, giving readers, platforms and publishers a way to tell machine-written content apart. But The Guardian's report asks whether the watermark will degrade Claude's output quality, making that the central open question of the rollout.

08/18, 00:15

Groq raises $350M to fuel its pivot from AI chips to neocloud

Groq has raised $350 million led by investment firm Disruptive, with planned participation from Nvidia, valuing the former AI chipmaker at $3.5 billion as it pivots to a neocloud business running Nvidia GPUs. The valuation is down from $6.9 billion last September, before Nvidia hired Groq's founder and top talent in a licensing deal.

08/17, 23:55

China's NBS confirms 81.8% surge in AI equipment spending as broader investment falls

China's National Bureau of Statistics confirmed Monday that internet companies' equipment spending surged 81.8% year-on-year in the first seven months of 2026, while overall fixed-asset investment fell 6.7%. The data validates the AI infrastructure buildout by Alibaba, Tencent, ByteDance and Baidu, with Tencent's second-quarter operating capex reaching 51.8 billion yuan, up 190% year-on-year.

08/17, 23:55

Alibaba launches HappyShrimp, an AI music model that generates complete songs from text prompts

Alibaba has launched HappyShrimp, a new AI music model that generates complete songs from text prompts, according to a report carried by finance.biggo.com. The release extends Alibaba's generative AI push from language and image tools into the audio domain.