Tech companies are secretly buying up rare old books, scanning every page, and then destroying them, just to train AI.

Here's how this came out, and why it's happening specifically with old, rare books, based on NPR's interview with 404 Media reporter Emanuel Maiberg -- whose original investigation is the source for everything below.

How It Was Discovered

A bookseller told Maiberg they'd been getting unusually large orders through Biblio, an independent rare-book marketplace that keeps buyer identities anonymous. Suspecting where the books were headed, the bookseller worked with Maiberg's team, who "managed to get a tracking device into one of these orders" -- reportedly an Apple AirTag slipped into a single volume inside a roughly 1,000-book bulk order. The tag's signal traced the shipment across several states before it landed at an Amazon complex called LAS8 in Las Vegas, inside a specific unit internally labeled VGT3. That unit's own logo, according to reporting on the investigation, is a T-Rex holding a book in its jaws -- workers there have described their job, on an internal Amazon forum, as simply: "all we do is scan books."

We Already Had a Clue

This wasn't the first sign of the practice. A lawsuit from book authors against Anthropic had already revealed that AI companies "were purchasing books in large quantities... and that in order to scan them really fast, they cut the spines off the books and fed the pages to the scanner." Court documents in that case referred to the effort internally as "Project Panama," and reporting has identified Tom Turvey -- the former Google executive who ran Google Books -- as the person Anthropic hired to lead the book-buying and destructive-scanning operation, reportedly using a hydraulic-powered cutting machine to strip bindings before scanning. Anthropic ultimately settled that lawsuit for roughly $1.5 billion, paying about $3,000 each for 500,000 works -- one of the largest copyright settlements in AI's short history. Separately, a June 2025 federal ruling in the same case found that scanning legally purchased physical books to train an AI model counts as fair use -- that ruling is specifically why buying real copies, rather than downloading pirated ones, has become the standard (and legally safer) path for AI companies doing this at scale.

How You Can Spot It

Booksellers who've dealt with these orders describe a specific pattern. "There was no coherent theme to the kind of books that the clients were purchasing," as Maiberg put it. "When you get an order for hundreds of books that have no coherent theme, that is another sign" it's a scanning operation, not a collector building out a themed library.

Why Old Books Specifically

AI companies already scraped the entire public internet for training text years ago. The next obvious idea -- training new AI on text generated by older AI -- turned out to backfire: "that causes a problem called model collapse," where the AI actually gets worse, since it's increasingly learning from its own errors and homogenized patterns rather than genuine human writing. That leaves a narrowing pool of fresh, human-written text that was never posted online, and old, rare, out-of-print books fit that need precisely -- they're real human writing that large language models haven't already absorbed a thousand times over.

Amazon's Response

Amazon confirmed the warehouse, called VGT3, is used for AI training, but its only public statement on the record was: "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use." Notably, that statement doesn't address the destructive-scanning method itself, and reporting on the investigation notes Amazon had previously denied engaging in this kind of destructive scanning -- a denial the AirTag's route directly contradicted.

What "Rare" Actually Means Here

Maiberg's point about what "rare" actually means is worth sitting with: it's rarely a first edition of Oliver Twist. It's more often "an instruction manual for a product that's no longer available," self-published, or otherwise never digitized -- the kind of book nobody thought to preserve, until an AI company decided it was useful.