Amazon Is Destroying Rare Books to Train Its AI
Amazon is buying rare books, cutting off their spines, scanning them, and trashing them for AI training data. Here's what that means — and why copyright law made it happen.
Written by AI. Zara Chen

So. Amazon — the company that started in a garage selling books, the company that was my library growing up when I couldn't find a title anywhere else — is now buying rare books, slicing their spines off, scanning them, and throwing them away. To feed an AI.
I keep turning that over and can't decide if it's tragedy or farce. Probably both.
According to TechCrunch, the story broke through an investigation by 404 Media, which found that Amazon is buying large quantities of rare books specifically to scan them for AI training data. The method: cut off the spine (which destroys the binding and, practically speaking, the book), photograph or scan every page, then discard the physical object. Ars Technica and Futurism both confirmed the core finding: Amazon bought a shipment of rare books so it could scan and destroy them for model training.
And the way 404 Media proved it? They put an AirTag inside one of the books. Watched it travel across the country. Watched it arrive at an Amazon facility in Las Vegas. Watched it go silent — almost certainly destroyed, according to AppleInsider.
Okay, I need to stop and actually say this: a reporter hid a tracking chip inside a book, sold it into a trillion-dollar company's supply chain, and used an iPhone accessory to document corporate book destruction in real time. That's not "clever." That's the kind of move that sounds fake when you describe it to someone. It worked. And the story it confirmed is somehow worse than the stunt itself.
Why Books, and Why Now
Here's the thing that makes this more than just a bad-optics story: the reason rare books specifically are in demand tells you a lot about where AI development is right now.
The prevailing assumption among AI researchers — and it's increasingly being treated as industry consensus, even if no single definitive study has closed the debate — is that the low-hanging fruit of internet text has mostly been picked. Web-scraped data has powered LLMs through several generations of models, but that well isn't infinite, and a lot of what's left online is low-quality, repetitive, or already in training sets. Books represent something different: longer-form, edited, structured prose that doesn't exist in digital form anywhere — especially when we're talking about older or more obscure titles.
That's why Amazon is apparently doing this methodically. The AI Chronicle reports that the investigation suggests Amazon is specifically targeting books with ISBN numbers to ensure a high volume of unique works. This isn't random hoarding — it looks like a systematic acquisition-and-scan operation.
Ars Technica adds a detail that stopped me cold: at one point, the supply of books at the facility completely ran out. Workers said so. But the operation kept going — which tells you demand for this kind of material is real, ongoing, and apparently not close to satisfied.
The AI Chronicle also notes that competitors Anthropic and xAI have publicly stated they don't train on rare or antique books. Amazon hasn't made any such statement.
The Copyright Angle Is the Part That Should Make You Laugh (In a Bad Way)
Now, here's where it gets legally weird — and honestly kind of darkly funny, if you're into gallows humor.
On Hacker News, where the story quickly surfaced on the front page (thread here), one commenter laid out the copyright logic bluntly: "Yes, this is a result of copyright laws. The other commenters are wrong/uninformed. If it was up to the companies training LLMs, they wouldn't destroy the books: It's a waste of company resources, it's needlessly destructive/evil, it generates bad PR, etc etc."
The commenter didn't provide a username visible in the sources, but honestly? They're not wrong. Which makes it worse.
Here's the legal architecture: under U.S. copyright law, making and retaining a digital copy of a copyrighted work creates ongoing liability. But there's a doctrine around "transient copies" — copies made in the course of a process that are not stored beyond what's immediately necessary — that some legal teams believe creates a narrower safe harbor. The logic, as I understand it: scan the book, use it in a training pipeline, don't keep the scan as an accessible file, destroy the physical object so there's no dual-format issue. You've technically never "published" a digital copy. Whether courts will ultimately accept that argument is a genuinely open question — but it's the bet Amazon appears to be making.
So the law has created a situation where destroying the physical object is the legally conservative move. The only way to make copyright lawyers comfortable is to make the book stop existing.
Let that sit for a second. The legal incentive structure we've built around intellectual property is so tangled that the safest corporate behavior is the most culturally destructive one. I'm not saying copyright law is wrong to exist — creators deserve protection, full stop. But I am saying that when the law's internal logic points toward burning the library, it might be worth asking whether the law needs a patch.
The Amazon-Shaped Hole in My Bookshelf
I grew up ordering books off Amazon as a kid. My parents couldn't always find niche titles locally, and Amazon had everything. I remember the excitement of a specific kind of tracking notification — "your package is out for delivery" — for a book I'd been waiting on. That's not nostalgia in an abstract sense; Amazon and reading were genuinely fused for a lot of people my age.
I'm not saying that's uncomplicated — Amazon also helped hollow out independent bookstores, and that's a real and painful thing. But my generation's relationship to book access and Amazon is genuinely intertwined in ways older readers might not feel in their bones the same way.
Which is why this story hits different than just "company does bad thing." The company that made books more accessible than they'd ever been is now one of the forces making certain books permanently inaccessible. Not because they want to harm books — they just want data. The books are raw material. The machine doesn't care what the raw material used to be.
What the Operation Actually Looks Like
Slashdot covered the core finding too: a tracking device inside a rare book led investigators to an Amazon AI training facility. The Las Vegas warehouse isn't a sorting center or fulfillment hub — it's a facility where books arrive and get processed specifically for AI training. Systematically, methodically, at scale.
The scale matters. "Rare book" isn't just about monetary value — it's about scarcity. First editions, out-of-print academic works, regional histories with small print runs, texts that exist in only a few hundred surviving physical copies. When those go into the shredder, the world's knowledge of those objects doesn't just migrate — it narrows. It becomes what Amazon's scanner captured on that day, at that resolution, interpreted by that pipeline, filtered into that model.
Where This Goes
Amazon isn't the first company accused of this — AppleInsider notes it has "joined the list of companies destroying rare books to feed to the AI machine." But Amazon's scale and its particular history with books makes the story land harder.
And here's the open question I can't answer from the sourcing: Is Amazon's scan-and-destroy pipeline producing meaningfully better AI outputs that justify this? Or is this a "we needed more data and this was available" decision that doesn't move the needle as much as the cultural cost suggests? Nobody outside those training runs knows. Amazon hasn't said.
What I do know is that Ars Technica confirmed the facility is still operational even after it ran through its initial book supply. Demand hasn't peaked. The operation continues.
I keep thinking about whoever sold that AirTag-loaded book into Amazon's supply chain without knowing it was being tracked. They thought they were just selling a book. They were actually feeding it into a machine that would eat it, extract what it wanted, and leave nothing behind.
That's not a metaphor. That's just what happened. And I can't shake the feeling that I'm going to keep writing about it for a while.
Zara Chen covers technology and politics for Buzzrag.
More Like This
Laravel 13.6 Drops Debounceable Jobs and JSON Health Checks
Laravel 13.6 introduces debounceable jobs, JSON health check responses, and Cloudflare email support. Here's what developers need to know.
This Creator Got Shadowbanned on YouTube in 25 Days—On Purpose
A vidIQ creator deliberately shadowbanned their channel with AI-generated content to expose how YouTube's algorithm actually works. The results are wild.
Heroku Is Really Dead This Time, and Here's What Happened
Heroku has entered full maintenance mode after mass layoffs and leadership exodus. How did Salesforce let a developer platform die at the finish line?
Master Remote Access with Comet Pro KVM
Explore the Comet Pro KVM for seamless remote PC access: Wi-Fi 6, out-of-band management, and Tailscale security.
What malloc Actually Does (It's Not Magic)
Dave's Garage breaks down how malloc really works—from a five-line bump allocator to 40 years of fragmentation fixes, security patches, and thread nightmares.
One PR Hijacked the Entire NPM Registry
A single pull request compromised 169 npm packages—no phishing, no stolen passwords. Here's how the TanStack supply chain attack actually worked.
RAG·vector embedding
2026-08-19This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.