目錄
In August 2026, 404 Media published a striking investigation: a reporter worked with a bookseller to hide an AirTag inside one book from a shipment of roughly 1,000 volumes, then tracked where the batch ultimately went. The signal left California, passed through Wisconsin and Colorado, and eventually stopped at Amazon’s LAS8 facility in Las Vegas—not in the ordinary print-on-demand area, but in a separate operation on the north side identified as VGT3. Based on descriptions posted by Amazon employees in internal work forums, 404 Media reported that the site receives large volumes of physical books, removes their bindings, and scans them, destroying the original copies in the process. Ars Technica and other outlets followed up the next day, quickly turning “an AirTag tracked into an Amazon warehouse” into one of the week’s most discussed AI-data stories.
404 Media described VGT3 as an Amazon AI training facility and argued that the available evidence strongly points toward AI training use; Amazon itself has not confirmed that claim. In responses to 404 Media and GeekWire, the company only said that Amazon purchases books through commercial channels to develop and improve products and services used by customers, without explaining which models or datasets ultimately receive the scanned text.
So what can actually be confirmed is not that “Amazon admitted it used all 1,000 rare books to train a specific AI model,” but something more concrete: large numbers of physical books are being purchased and sent into an Amazon facility, employees describe a workflow that includes cutting and scanning books, and that process destroys the original physical copies. How the resulting scanned data is ultimately used remains undisclosed by Amazon.
How Did an AirTag Reveal Amazon’s Book-Scanning Facility? The Original Investigation Came from 404 Media
An Anonymous Order of Roughly 1,000 Books Ended Up at Amazon LAS8 in Las Vegas
The investigation did not begin with a reporter randomly mailing out an old book. 404 Media first received a lead from a bookseller who had been tracking unusual bulk purchases in the used-book market. One seller received an order for roughly 1,000 books through Biblio from a buyer whose identity was not transparent, and similar orders had begun appearing more frequently in recent months. The bookseller agreed to let 404 Media provide an AirTag, which was hidden inside one of the books so the final destination of the shipment could be observed.
The AirTag eventually reached Amazon’s LAS8 facility in Las Vegas. At first, that location confused the reporters because most of LAS8 is associated with Print-on-Demand, where books are printed after someone orders them—the exact opposite of a process involving large-scale book destruction. The clue began to make sense only when they examined the AirTag’s final position: not in the normal print-on-demand area, but in a separate operation on the north side of the warehouse identified as VGT3.
According to descriptions from Amazon employees posted in workplace forums, VGT3 receives large quantities of physical books, with some workers cutting books apart and others handling receiving and scanning. 404 Media’s investigation further reported that bindings are removed so the pages can be scanned more quickly, destroying the physical books in the process. VGT3 even has its own sign featuring a Tyrannosaurus rex showing its teeth while holding a book. The earlier wording that it was “about to swallow the book” added more drama than the available evidence supports; the existing photo and descriptions are better summarized as “a T. rex holding a book.”
The Shipment Can Be Called “Rare Books,” but They Were Not All Unique Copies or Museum-Grade Antiquarian Volumes
404 Media itself referred to the shipment as rare books, and the bookseller involved in the investigation works in the rare-book market. But translating “rare” too literally can easily create the impression that every volume was a museum-grade unique copy, first edition, or centuries-old antiquarian book. The reality was more mixed.
When GeekWire summarized 404 Media’s reporting, it specifically noted that the bookseller observed that these bulk buyers were not targeting the “rarest” category of books—very old titles that did not even have ISBNs. That led the bookseller to suspect the buyer might instead be systematically searching by ISBN and other publishing metadata for books that had not yet been digitized.
So the more accurate description is “rare, obscure, difficult-to-find, or older books,” rather than treating the entire shipment as globally unique cultural artifacts. Preservation concerns still exist, but the significance of losing a physical copy depends on the edition, surviving quantity, and physical characteristics of each individual book.
Why Are AI Companies Looking at Physical Books Again? The Main Attraction Is Clean Text That Has Never Been Digitized
Old Physical Books Are Being Marketed as Sources Free from AI-Generated Content
A month before the AirTag investigation, 404 Media had already reported on another part of the supply chain. Book-data company ISBNdb was pitching services to AI customers for acquiring large volumes of physical books, with one of the central selling points being that older books were produced before generative AI became widespread and therefore would not be contaminated by the rapidly increasing volume of AI-generated text. ISBNdb even described physical books as edited, structured sources containing domain knowledge.
So the earlier idea that “physical books are among the few remaining high-quality corpora that have not yet been drained” can remain, but it should not be presented as if the industry has reached an agreed-upon threshold where “internet data is running out.” What can currently be confirmed is that the AI data supply chain is actively looking for two things: text that has not yet been digitized, and human-authored material created before generative AI became widespread.
For model developers, these books also offer another practical advantage. An obscure local-history book, old technical manual, or out-of-print professional title may have little value in the ordinary second-hand market, but if its contents have never appeared in Common Crawl or a major ebook database, its value as training data may be calculated very differently from its resale price.
Cutting Off the Spine Before Scanning Was Not Invented by Amazon
Destructive scanning itself is not new. When large numbers of bound books need to be scanned, cutting off the spine allows loose pages to be fed through high-speed scanning equipment instead of turning every page manually. What has drawn renewed attention is that generative AI companies are buying physical books at scale and applying this long-standing digitization technique to model-data acquisition.
Anthropic is one of the clearest examples already documented in court records. The 2025 Bartz v. Anthropic ruling describes Anthropic buying physical books, removing their bindings, scanning them page by page, converting them into searchable digital copies, and then selecting material from its central database for training Claude. The court record also explicitly states that the physical source books were destroyed during the conversion process.
So Amazon is not the first AI-related company associated with destructive book scanning. The difference is how the practice became public. Anthropic’s approach emerged during Copyright Litigation Discovery, while the Amazon operation was traced through the physical supply chain using a bookseller and an AirTag.
Has Amazon Actually Admitted That These Books Are Being Used to Train AI? Not Yet
Amazon Confirmed Buying Books, but Not What the VGT3 Scans Are Used For
This is the uncertainty that needs to remain intact throughout the article. 404 Media’s investigation headline directly described the destination as an Amazon AI training facility, and there is substantial circumstantial evidence behind that inference, including the purchasing pattern, the VGT3 workflow, Amazon’s AI business, and similar practices already documented at other AI companies. But Amazon’s formal response remained much broader, saying only that it purchases books through commercial channels to develop and improve products and services used by customers. When GeekWire directly asked whether the books were being used to train models, Amazon still did not confirm it.
For that reason, the earlier sentence claiming that “Amazon had previously denied this kind of practice” should be removed entirely. There is currently no reliable record showing that Amazon previously denied VGT3 or a similar internal book-scanning program. The confusion likely comes from statements made by other companies in recent weeks: Zoom Books said it does not digitize and then destroy books, while Anthropic said its acquisition program would not purchase and destroy rare or antiquarian books. Those were not Amazon statements.
The more accurate question is therefore not “Why did Amazon lie?” but this: Amazon has acknowledged purchasing books to improve products and services, and 404 Media identified an Amazon operation where books are physically cut apart and scanned, but the company has not disclosed what system ultimately receives that resulting data.
If You Buy a Physical Book, Can You Just Scan It? Ownership and Copyright Are Two Different Things
The First Sale Doctrine Lets You Dispose of That Copy, but It Does Not Automatically Grant Reproduction Rights
The earlier explanation that “a legally purchased book belongs to the buyer, so destroying it is not automatically illegal” is directionally correct, but the next legal issue needs to be separated from it.
Section 109 of U.S. copyright law, commonly known as the First Sale Doctrine, deals with a lawfully acquired physical copy. The owner can generally resell, lend, or dispose of that particular copy, and the copyright holder cannot continue controlling the circulation of that same physical object forever.
But scanning a book creates a reproduction, which means “buying the book” does not automatically equal “obtaining the right to scan it.” Whether the content can be copied without the author’s permission must be analyzed under other parts of the Copyright Act, including Fair Use. That is why current AI-training disputes cannot be resolved simply by saying that the company paid for the physical copy.
Anthropic Won on Specific Issues in One Case, but That Does Not Mean U.S. Courts Have Declared All AI Training Legal
In the 2025 Bartz v. Anthropic federal district court ruling, the court held that using the books at issue to train Claude constituted Fair Use. It also separately found that converting lawfully purchased physical books into a digital library qualified as Fair Use under the specific facts of that case, with one relevant factor being that the original physical books were destroyed rather than leaving a second circulating copy in addition to the digital version.
The same ruling, however, refused to excuse Anthropic’s acquisition and permanent retention of books obtained from pirate websites, analyzing “lawfully purchased and scanned” material separately from books first acquired through unauthorized sources.
So as of 2026, the safer formulation is that federal district courts have found AI training to qualify as Fair Use under the particular factual records presented in Anthropic and Meta cases, but that does not mean every model, every data source, and every acquisition method has been universally declared lawful. The source of the data, market effects, model outputs, and the purpose of use can still affect the legal analysis.
Why Destroy the Books After Scanning? Legal Reasons and Preservation Reasons Should Not Be Mixed Together
Destructive Scanning Prioritizes Scanning Efficiency, Not Physical Preservation
What the public evidence supports is that removing the binding makes scanning faster. In the Anthropic case, the fact that the physical copy was replaced by a digital version even became part of the court’s analysis of whether that particular digitization qualified as Fair Use.
Amazon has not publicly explained why books processed at VGT3 are not rebound, donated to libraries, or preserved long-term. The earlier draft attributed the decision directly to storage, temperature and humidity control, insurance, and labor costs. Those considerations may be plausible in practice, but at this stage they remain inference rather than documented Amazon reasoning and should not be presented as the company’s actual rationale.
What can be said more confidently is that if the objective is to acquire text at scale, cutting, scanning, and moving on to the next book is much faster than restoring each physical volume to collectible condition. That efficiency is the reason destructive scanning exists, and it is also what preservation advocates find most troubling.
Digital Text Can Survive, but a Specific Physical Book Contains More Than Text
For an ordinary mass-market used book, the loss may amount to one fewer physical copy. But if a book has significance because of its edition, special binding, signature, ownership stamp, bookseller label, or handwritten reader annotations, scanning the text does not preserve all of that information.
This is why the phrase “rare books” needs to be handled carefully. The investigation did not prove that VGT3 is systematically destroying unique cultural artifacts, but if the acquisition system is primarily optimized around ISBNs and textual value, it may prioritize “Is this book useful as text?” over “Is this specific physical copy worth preserving?” The first question belongs to model-data infrastructure; the second is the concern of libraries, archives, and antiquarian book markets.
Amazon Is Not the First Company to Scan Books—the New Part Is That the Supply Chain Was Traced to a Specific Facility
Anthropic’s Project Panama has already been exposed through litigation records, which show that the company bought, dismantled, and scanned physical books at scale. In recent months, second-hand booksellers in the United Kingdom, Ireland, and Australia have also reported unusual bulk orders: large quantities of titles with little obvious thematic connection, buyers that seemed relatively insensitive to price, and destinations that were sometimes forwarding warehouses. Some booksellers have begun refusing certain orders because they fear the books will eventually be subjected to destructive scanning.
What the Amazon investigation adds is a physical supply chain that can be followed end to end. Roughly 1,000 books left the seller, the AirTag eventually entered Amazon LAS8/VGT3, and VGT3 employees separately described book-cutting and scanning work. Previously, the question was “Who is buying all these books?” This time, at least one shipment reached a concrete destination.
So the better conclusion is not that “the entire AI industry is secretly destroying rare books,” but that destructive scanning is no longer only a special case documented in Anthropic litigation. There is now at least one independent investigation tracing similar physical operations to an Amazon facility.
Flock’s 120,000 License Plate Readers Offer a Different Comparison About Scale
In the same week, Flock Safety faced extensive criticism over its nationwide automated license plate recognition network and announced shorter data-retention periods, mandatory use of anomalous-search auditing tools, and requirements that law-enforcement searches be tied to case numbers. Flock’s system now spans 49 U.S. states and includes more than 120,000 cameras.
This is not the same legal issue as Amazon scanning books, and it would be inappropriate to argue that “photographing one license plate and scanning one book are both legal, but multiplying them by 100,000 suddenly makes them illegal.” The constitutionality and data use of automated license plate recognition remain the subject of continuing litigation, and rules differ across states.
What makes the comparison useful is the effect of scale. A single ALPR recording one vehicle is very different from 120,000 cameras forming a searchable database of movement across regions. Likewise, a company scanning a handful of books for internal search is not the same data-acquisition capability as building a large-scale supply chain for purchasing, cutting, and scanning physical books.
That comparison better explains the current controversy: many of the underlying technical actions have existed for years. What changes the environment is automation, databases, and scale connecting those actions into systems capable of processing vastly more material at once.
What Does This Have to Do with Website Creators? Physical Books Are Just the More Expensive Data Pipeline
Physical books at least require someone to obtain a copy, ship it, dismantle it, and scan it. Public websites do not have those logistics costs. If a page can be publicly read, acquiring its text is usually much cheaper.
That means blogs, newsletters, public databases, and other websites face a different kind of data supply chain. robots.txtcan tell crawlers which paths a site does not want them to access, but the IETF’s Robots Exclusion Protocol explicitly warns that it is not an effective content-security mechanism and cannot replace real Access Control.
That does not mean robots.txt is useless. It remains a standardized way of expressing crawler policy and can influence search engines and AI crawlers that choose to comply. It simply should not be understood as a technical anti-theft lock.
For content operators, a layered approach is more practical. Public articles intended to appear in search should continue allowing normal Crawling. Member-only material that should not be freely obtained should use login, permissions, and server-side access controls. Content published through third-party platforms requires a separate review of platform terms and whether an AI training opt-out is available. A single robots.txt file cannot solve all three problems at once.
What Creators Should Do Now Is Not Hide Everything
First Check How Each Publishing Platform Handles AI Training Data
The same article published on a self-hosted WordPress site, a social platform, and an email newsletter service is governed by different rules. Whether a platform obtains permission to use content for training, whether users can opt out, and whether an opt-out applies only to future data or also to existing material all need to be checked separately.
This step is tedious, but it is more practical than writing “AI training prohibited” underneath a work and assuming the issue is solved. A statement can preserve the author’s position, but actual access and authorization still depend on website controls, platform terms, and applicable law.
Treat robots.txt as a Policy Layer on Your Own Website, Not a Security Layer
You can continue configuring robots.txt, especially when known AI crawlers publish identifiable User-Agents, because it provides a direct way to express site policy. But RFC 9309 is explicit that it is not a content-security mechanism. Data that genuinely must not be publicly obtainable needs Authentication, Authorization, or other server-side controls.
Public SEO articles and private content therefore do not need the same strategy. Articles that should be found through Google naturally need to remain accessible to search crawlers. Paid educational material, internal files, or content that genuinely cannot be allowed to leak should never rely on Disallow alone.
Keep First-Party Evidence and Verification Records in Your Own Systems
On-location photos, interview recordings, test records, raw spreadsheets, purchase receipts, and version history matter for more than simply being “harder for AI to copy.” They answer a question generated text often cannot: how was this information obtained?
A published article can be crawled, and its sentence structures can be imitated, but the original interview files and testing records remain under the control of the person who generated the evidence. For content sites trying to build long-term credibility, that provenance is much more useful than deliberately trying to make an article sound “less AI-like.”
Separate Confirmed Facts from Unconfirmed Claims in Public Articles
The Amazon case itself is a good example. We can confirm where the AirTag ended up, how employees described VGT3, and how Amazon responded. What remains unconfirmed is which model receives the scanned data, how many books the operation processes in total, and which Amazon products currently use the resulting dataset.
Separating those two layers makes it easier for readers to understand where the investigation ends and where reasonable inference begins. As AI-generated content becomes more common, there is no need to invent a special kind of “human tone” to stand out. Simply marking the evidence boundary clearly already makes a substantial difference.
The most interesting part of the AirTag investigation is not only the T. rex at VGT3. 404 Media began with a data supply chain that was difficult to observe directly: booksellers knew someone was placing large bulk orders, and the public already knew Anthropic had used destructive scanning, but no one knew which company was receiving a particular shipment. A tracker hidden in a physical book finally followed one batch all the way to Amazon.
But that is also where the evidence should stop. Amazon has not confirmed that the content scanned at VGT3 becomes the training set for a particular AI model, nor has it disclosed the scale of the operation. 404 Media’s investigation provides strong clues and confirms that physical book-cutting and scanning operations exist at the facility. The next step still requires a more complete explanation from Amazon rather than filling the remaining gaps with assumptions.
For content creators, this does not need to become a conclusion that “nothing should ever be published again.” The more realistic reminder is simply that once text enters a public distribution system, the ways it may later be acquired, transformed, and reused can no longer be understood only as “someone reading this article.” Platform settings, crawler policies, licensing terms, and preservation of original source material are all becoming part of content work.
FAQ
Not fully. 404 Media traced a shipment of roughly 1,000 books to Amazon LAS8/VGT3 and obtained employee descriptions of book-cutting and scanning work. Amazon only said that it purchases books through commercial channels to develop and improve products and services used by customers, without confirming which AI model, if any, uses the VGT3 scanning data.
That cannot be established. 404 Media referred to the investigated shipment as rare books, but the bookseller also said the bulk buyer did not focus on the very rarest books that lacked ISBNs altogether. A more accurate description is rare, obscure, older, or not-yet-widely-digitized physical publications.
Destroying or disposing of a lawfully purchased physical book and reproducing copyrighted content from that book are separate legal questions. The U.S. First Sale Doctrine generally allows the owner of a lawful copy to dispose of that physical copy, but scanning creates a reproduction that must be analyzed separately. In the specific 2025 Bartz v. Anthropic case, the court found Anthropic’s training use and certain digitization of lawfully purchased physical books to qualify as Fair Use, but that does not mean every AI training-data acquisition method has been universally declared lawful.
One explicit selling point from data suppliers is that books published before generative AI became widespread do not contain the large volume of AI-generated text appearing online today, while many obscure works have also never been fully digitized into internet datasets. 404 Media’s July investigation documented book-data companies marketing older physical books to AI customers for exactly this reason.
No. Litigation records involving Anthropic already document the company buying physical books at scale, removing their bindings, scanning them, and creating a digital library. What makes the Amazon case distinctive is that 404 Media used a physical tracker to follow one shipment directly to Amazon VGT3.
Not reliably. robots.txt is a standardized crawler-policy mechanism, but the IETF explicitly states that the Robots Exclusion Protocol is not a substitute for actual content-security controls. It is useful for expressing crawler preferences, while content that genuinely must not be publicly accessible still requires login, permissions, or server-side Access Control.