Refacto AI

Industry story

Amazon facility caught destructively scanning rare books for AI training

big-tech compliance copyright training-data

Amazon is running an industrial-scale book-destruction operation out of its LAS8 fulfillment facility in Las Vegas, confirmed by 404 Media tracking an AirTag hidden inside one of roughly 1,000 rare books bought anonymously through Biblio. Anthropic was reported doing the same thing in June 2025. The more interesting question is why: a thousand rare books is rounding error against a 10-trillion-token corpus, so the plausible purpose is legal cover, not capability lift. Buying and destroying a licensed physical copy litigates better than a LibGen mirror, and after the Bartz ruling, that paper trail is worth real money in a courtroom.

Full analysis

Amazon is running a dedicated Las Vegas facility, marked with a dinosaur-holding-a-book logo, that destructively scans rare books to feed AI training data. 404 Media proved it by hiding an AirTag in one of ~1,000 volumes bought through Biblio by an anonymous, price-insensitive buyer and watching it land at the VGT3 corner of LAS8. Anthropic was reported doing the same thing back in June 2025. The question for anyone building with models: does physical-book acquisition actually matter to your training pipeline, or is this legal cover with a good story attached?

Reversibility: This is a Type 1 posture for the labs (a standing industrial operation, hard to unwind) but a Type 2 read for you (whether to build a scanning pipeline is a cheap experiment). What's really being decided is whether pre-digital corpora are a capability edge worth chasing, or a compliance hedge dressed in spy-thriller copy. No forcing function beyond the litigation clock.


The Skeptic. A thousand rare books against a 10-trillion-token corpus is rounding error. Even scaled to millions of volumes, the marginal capability lift is noise. The plausible reason to buy physical copies and shred them is a paper trail: "we owned a licensed copy" litigates better than "we torrented LibGen." Anthropic doing the same thing in parallel tells you this is industry-wide risk hedging, not a data arms race. The dinosaur logo and the AirTag make it a great read. The token economics do not support the drama. For the PM at the next desk: buying and scanning the book is mostly about what a judge sees later, not what the model learns.

The Safety Lens. The books are not the story. The pattern is: anonymous bulk buying, dedicated dark facilities, plausible deniability routed through a marketplace, zero disclosure or consent trail. It is engineered to be legally defensible without being transparent. When the same labs publishing responsible-AI charters run covert acquisition ops, they burn the trust any oversight regime needs. Data provenance is heading toward being a first-class audit requirement, in the EU AI Act's training-data documentation obligations and in every active copyright suit. These facilities become the liability. For the PM: the exposure here is a subpoena and a headline with your logo on it, not a bad model.

The Researcher. The AirTag reporting is genuinely good field work, but it confirms a known pattern. Books3, LibGen, Z-Library mirrors already established that labs digitize everything they can. What's new is operational scale: a branded, sustained, industrialized program, not opportunistic scanning. The research question nobody's answering is what's in those ~1,000 titles that Common Crawl doesn't already have. Rare books imply out-of-print academic monographs, pre-digital technical manuals, primary historical sources. That's a real signal about where frontier capability gaps still live: the long tail of expert knowledge that never made it onto a webpage. For the PM: the open web is scraped dry, so labs are going to the stuff that was never online.

The Compute Pragmatist. Everyone anchors on GPU FLOPs and misses that this is an I/O and preprocessing problem. Overhead scanners, OCR at volume, dedup against your existing corpus, formatting normalization. Rare books scan dirty. Old typefaces, foxed pages, footnotes, tables. Cleaning a rare-book token can cost 10 to 100 times what a clean web token costs to prepare. A dedicated Amazon fulfillment node strongly implies AWS is the infra layer underneath, which is its own tell about who profits either way. The binding constraint isn't acquisition, it's the cleaning compute. For the PM: getting the book is trivial, turning a shredded 1940s monograph into usable training text is the expensive part.


Where they split. The Skeptic and the Researcher genuinely disagree on whether this is a capability play at all. Skeptic says it's litigation theater with a rounding-error payload. Researcher says the long-tail expert knowledge in out-of-print books is exactly the corpus the open web can't supply, so the marginal token is worth more than its count suggests. That's the real fork: is a scanned rare book worth 1x a web token or 100x?

The Safety Lens and the Skeptic agree on the mechanism but weigh it opposite. Both see the physical-copy purchase as a legal move. Skeptic thinks it works. Safety Lens thinks it's the exact behavior that makes provenance a mandatory audit and turns the facility into the smoking gun.

What it hinges on. Two beliefs. First, whether long-tail physical corpora produce a measurable eval lift on domain tasks, or vanish into the training noise. Second, whether "we bought a physical copy and destroyed it" actually holds up as a defense, or gets treated like the Books3 downloads did. The Anthropic verdict in the Bartz case last year already found scanning purchased physical books was fair use while pirated downloads were not, which is why everyone suddenly owns a book shredder. That precedent, not the capability, is what's driving the dinosaur logos.

If you build models, the cheap experiment is to fine-tune on a few thousand cleaned public-domain out-of-print monographs in your domain and measure the eval delta before you assume the physical-source moat is real. Don't build a scanning line on vibes.


Prediction: No US court will invalidate the "buy a physical copy, scan it destructively, train on it" fair-use path before the next major appellate ruling in an AI copyright case, and by 2027-02-20 at least one additional large model builder beyond Amazon and Anthropic will be publicly reported running the same physical-book-scanning playbook.

Confidence: Medium The legal incentive is settled and copyable; timing of the next report is the only real variable.

Why: The reason these dinosaur facilities exist is the June 2025 Bartz v. Anthropic ruling, which held that training on lawfully purchased and scanned books is fair use while training on pirated copies is not. That gives every lab a legal template: buy the book, destroy it, keep the receipt. Amazon and Anthropic are already reported doing it, and the acquisition cost is trivial while the litigation cover is enormous, so the behavior spreads to any lab with a legal team and a data-provenance problem. The opposite outcome, a court slamming this path shut, is unlikely on this timeline because the controlling precedent points the other way and appeals move slowly. The thing that spreads here is the legal hedge, not a capability breakthrough.

Revisit by 2027-02-20: We're right if a third named large model builder (OpenAI, Google, Meta, Mistral, xAI, or comparable) is credibly reported buying and destructively scanning physical books, and no court has struck down the purchased-copy fair-use path. We're wrong if a court invalidates that path, or if no additional lab is reported doing it and the practice stays confined to Amazon and Anthropic.

Comments