Tracked by AirTag: Inside Amazon's Las Vegas Facility That Scans and Destroys Rare Books to Train AI
A 404 Media investigation hid a tracking device in a shipment of roughly 1,000 rare books and followed it to Amazon's VGT3 warehouse in Las Vegas, where workers cut off spines, scan pages for AI training data, and destroy the physical copies.
There is a certain irony in how the story was uncovered. To find out which AI company was quietly buying up the world’s rare books, journalists at 404 Media hid an Apple AirTag — a consumer gadget designed to help you find your lost keys — inside a shipment of roughly 1,000 rare books, and simply watched it travel across the country. Its final destination, revealed on August 17, 2026, was an Amazon warehouse in Las Vegas, Nevada. Inside, according to the investigation, Amazon employees spend their days receiving massive shipments of printed books, cutting the bindings off so the pages can be scanned more quickly, and feeding the results into Amazon’s AI training data pipeline. The physical book is destroyed in the process.
Amazon, the company that started life in 1995 as an online bookstore with the slogan “Earth’s Biggest Bookstore,” is now dismantling the physical stock of the used and rare book trade to fuel its machine learning models. The facility, identified by the code VGT3, even has a mascot: a snarling dinosaur clutching a book in its claws.
How the investigation worked
According to 404 Media’s Emanuel Maiberg, the newsroom had suspected for some time that AI companies were behind a wave of bulk purchases sweeping the rare book market. Booksellers had noticed anonymous buyers snapping up inventory at prices that made little sense for resale — no haggling, no questions about condition, just volume. The prevailing theory, which 404 Media had explored in an earlier July investigation, was that printed books had become one of the last uncontaminated troves of human writing left on the planet.
So the reporters tested it empirically. They placed a tracking device in a rare book shipment they suspected would be acquired for AI training, and followed it around the country until it stopped moving. The trail ended at VGT3 in Las Vegas. Amazon employees who work at the location told 404 Media that all they do is receive enormous shipments of printed books and cut the spines off so scanners can process the pages faster. TechCrunch, which covered the findings the same day, noted that the operation had never been previously reported.
Amazon’s response to 404 Media was brief: the company “purchases books through commercial channels to improve the products and services customers use.” It did not dispute that books are scanned and destroyed.
Why old books are suddenly gold
The economics here are straightforward and a little bleak. Large language models have already ingested essentially everything readily available on the open web — and increasingly, that web is polluted with AI-generated text. Training on synthetic output risks “model collapse,” the documented phenomenon where model quality degrades after ingesting too much machine-written content.
Printed books solve two problems at once. First, they contain knowledge that was never digitized: out-of-print titles, obscure local histories, technical manuals, and first editions that exist in a few hundred copies scattered across physical shelves. Second, anything published before roughly 2022 carries a guarantee that no web corpus can offer: it was written by a human being. As 404 Media put it in its July investigation of the book-sourcing industry, “the world’s best AI training data is sitting on a shelf.”
That earlier reporting introduced ISBNdb, a company that maintains one of the largest book databases in the world and offers high-volume book acquisition services to AI companies. Its pitch to clients was blunt about the reputational stakes, reportedly warning that “the optics problem is real.” In the wake of this week’s reporting, a book-database middleman that had marketed printed books as AI training data reportedly withdrew its offer following the scrutiny.
Amazon is not alone in this race. Anthropic disclosed in court filings during its copyright litigation that it had purchased and scanned millions of printed books for training data, destroying the physical copies afterward — and, as TechCrunch pointedly noted, Anthropic separately admitted to training on pirated books downloaded from torrent sites. Against that backdrop, buying books legally and destroying them quietly is arguably the “responsible” version of data acquisition.
The copyright calculus — and the cultural one
Why destroy the books at all? The most likely explanation is legal hygiene. In several jurisdictions, making a copy of media you own is tolerated only as long as you retain the original — the same logic that once governed ripping CDs. Keeping a scanned corpus while reselling the physical books would invite a straightforward infringement claim. Shredding the evidence keeps the acquisition inside the first-sale doctrine’s gray zone, even though whether training a model on a work constitutes “copying” at all remains an unsettled question in the courts.
Commenters on the story pushed back on the framing, noting that the used book trade already destroys staggering quantities of books every year — hundreds of thousands of tons are pulped annually when no buyer materializes. From that angle, a bulk buyer that pays sellers fairly might look like a lifeline rather than a vandal.
But that argument misses what makes rare books different. A pulped mass-market paperback is a lost copy; a scanned-and-shredded rare edition can be a lost artifact. Many of the volumes flowing into facilities like VGT3 are precisely the ones that exist in small numbers and were never digitized by anyone else. Once the spine is cut and the scan is proprietary, the public’s access to that text now depends entirely on Amazon’s commercial training pipeline — and the physical object is gone forever.
What to watch
The VGT3 revelation lands amid a broader reckoning over AI training data: publishers are litigating, governments are drafting disclosure rules, and the pool of high-quality human text keeps shrinking relative to synthetic output. Expect three follow-ons. Rare booksellers, now aware of who their bulk buyers are, may start pricing — or refusing — accordingly. Regulators and courts will eventually have to decide whether scan-and-destroy is a genuine legal safe harbor or just an untested assumption. And competing labs, none of whom can afford to be left out of the last clean data source on Earth, will keep buying.
The dinosaur holding a book in its claws turns out to be a better mascot than Amazon probably intended. The question the investigation leaves hanging is which side of the joke the rest of us are on.
Sources
- [1] https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/
- [2] https://techcrunch.com/2026/08/17/amazon-once-an-online-bookseller-is-destroying-rare-books-to-train-ai-models/
- [3] https://www.404media.co/ai-companies-are-buying-tons-of-old-books-because-theyre-free-of-ai-slop/
- [4] https://news.slashdot.org/story/26/08/17/1644216/tracking-rare-books-leads-to-an-amazon-ai-training-facility
- [5] https://aiweekly.co/alerts/amazon-scans-and-destroys-rare-books-at-vegas-ai-facility