sakutto
Generative AI

Amazon Rare Books Are Cut Up for AI Training Data

AmazonTraining dataCopyright
Amazon Rare Books Are Cut Up for AI Training Data

How the Amazon rare books shipment was traced

The story starts from a question about where AI companies source training data. The method 404 Media chose turned out to be the answer.

A tracking device went inside a rare book

404 Media picked a rare book it suspected an AI company would buy, hid a tracker in it, and watched the package move across the country. 404 Media says Amazon's book-buying operation had not been reported before this. Training-data procurement usually stays inside contracts and internal teams, which is exactly why so few examples of it are observable from outside at all.

View official source →
"A 404 Media investigation was able to reveal Amazon's book buying operation, which hasn't been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination."— from 404 Media's investigation

The trail ended at a Las Vegas facility called VGT3

The tracker came to rest at an Amazon warehouse in Las Vegas, Nevada. The Amazon team working there goes by VGT3. Its logo is a dinosaur, teeth bared, with a book in its hands. Whatever else is undisclosed about the operation, the branding is not being coy about what happens to the books.

View official source →
"That final destination was an Amazon warehouse in Las Vegas, Nevada." (the destination)/"The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands." (the VGT3 team)— from 404 Media's investigation

What happens to the books inside the facility

This is the part of the reporting with the sharpest edge on it.

Employees say the bindings get cut off before scanning

Workers at the site describe the job as receiving massive shipments of printed books and cutting the bindings off so the books scan faster. The stated reason is scanning speed, and the printed book does not come back from it. There is no storage step and no resale step. Given that the material being bought is rare books, what gets destroyed is not replaceable inventory.

View official source →
"Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process."— from 404 Media's investigation

Amazon confirmed buying the books and nothing else

Amazon told 404 Media it purchases books through commercial channels to improve the products and services customers use. The purchasing is acknowledged. The cutting and scanning are not addressed. The stated purpose stops at improving products and services, with no mention of which models the text trains, or how the scanned material is retained afterward.

View official source →
"The facility, known as VGT3, identifies itself with a symbol of a dinosaur holding a book in its claws. Amazon told 404 Media in a statement that it \"purchases books through commercial channels to improve the products and services customers use.\"" (Amazon's statement)— from TechCrunch

Why old print books are wanted as AI training data

The question left at the end is why books, specifically. The answer sits on the training-data side rather than the publishing side.

Text published before 2022 carries a guarantee

Training a large language model takes an amount of text that is hard to hold in your head, and the open internet has already been scraped for what it can give. That pushes demand toward rare books, especially ones out of print or absent from the web entirely. Anything published before 2022 comes with a guarantee no newer text can offer: no language model wrote it. Once generative AI went mainstream, that certainty stopped being available for free.

View official source →
"Companies like Amazon need unfathomably large amounts of text to train their LLMs, which have already ingested what they can from the internet (and, in Anthropic's case, illegally pirated books). Rare books, especially ones that are out of print or impossible to find on the internet, offer a new source of coveted training data."— from TechCrunch

Model collapse is what that guarantee protects against

Feed a model enough AI-generated text and output quality can degrade, a failure mode known as model collapse. The training supply gets contaminated by the industry's own output. Material that is provably human-written is the resource that holds that off, which is why text surviving only on paper has become worth the cost of acquiring and destroying it.

View official source →
"These texts are especially valuable since there's no chance that anything published before 2022 was written by an LLM. When LLMs train on AI-generated text, they risk \"model collapse,\" which can occur when the quality of an LLM's outputs degrade after ingesting too much AI-generated text."— from TechCrunch

Fights over how training data gets sourced are already underway. TechCrunch's piece notes Anthropic ingesting illegally pirated books as a contrasting case. Buying a book legitimately and then cutting it apart is, at minimum, a different acquisition route than piracy. Whether purchase also confers the right to train on the contents is a separate question, and one rights holders answer differently. What disclosure obligations land on AI companies is still being written, in the EU AI Act provisions that took effect in August and in California's AI transparency law.

View official source →
"Companies like Amazon need unfathomably large amounts of text to train their LLMs, which have already ingested what they can from the internet (and, in Anthropic's case, illegally pirated books)."— from TechCrunch

The investigation goes subscriber-only past a certain point, so the free portion is the part anyone can independently check — and it is what every quote here rests on. Converting that page to Markdown keeps its headings and lists intact, which is what you want when collating quotes against the original.

Free ToolURL to Markdown ConverterConvert any public web page URL to Markdown. Preserves headings, tables, lists, and links — perfect for LLM and RAG preprocessing, research notes, and archiving web articles.Try it now →

Summary

One tracker, one rare book, and a corner of the AI training-data supply chain became visible. The destination was VGT3 in Las Vegas, the work is cutting bindings off and scanning, and the printed book does not survive it. Amazon has confirmed the purchasing and said nothing about the rest. What drives all of it is the premium on text that provably predates language models, which puts the most pressure on writing that exists nowhere but on paper. How much of that supply chain companies will be required to disclose is still being decided in law rather than in policy statements.

FAQ

Q. How did this come to light?
404 Media planted a tracking device in a rare book it expected an AI company to buy, then followed the shipment across the country. It ended at an Amazon warehouse in Las Vegas. 404 Media says Amazon's book-buying operation had not been reported before.
404 Media — We Tracked a Shipment of Rare Books
A 404 Media investigation was able to reveal Amazon's book buying operation, which hasn't been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination. 404 Media — We Tracked a Shipment of Rare Books
Q. What happens to the books after they are scanned?
They are gone. Employees told 404 Media the job is cutting the bindings off so the books scan faster, and the printed book does not survive that. Nothing about the process assumes resale or return.
404 Media — We Tracked a Shipment of Rare Books
Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. 404 Media — We Tracked a Shipment of Rare Books
Q. Why is old printed text valuable as AI training data?
Because nothing published before 2022 could have been written by a language model. Training on AI-generated text risks model collapse, where output quality degrades, so text that is provably human-written is worth hunting for.
TechCrunch — Amazon, which started off selling books, is destroying rare texts to train AI
These texts are especially valuable since there's no chance that anything published before 2022 was written by an LLM. When LLMs train on AI-generated text, they risk "model collapse," which can occur when the quality of an LLM's outputs degrade after ingesting too much AI-generated text. TechCrunch — Amazon, which started off selling books, is destroying rare texts to train AI

Related Tools

Related Tool Categories

Articles