AI companies have spent years pulling training material from the internet. Now, as more of the web fills with AI-generated writing, some are looking for text in a place largely untouched by chatbots: used-book shelves.
This Tweet is currently unavailable. It might be loading or has been removed.
This Tweet is currently unavailable. It might be loading or has been removed.
Until July 28, ISBNdb, a company that operates an online book database, advertised its ability to source as many as 1 million physical books per order for AI developers. The since-deleted page offered books “tailored to your LLM training needs,” according to 404 Media, including older, specialized, rare, and out-of-print titles gathered from used bookstores and other catalogs.
The company also promised confidentiality. Its marketing materials advertised a “strict NDA on every engagement” and said clients’ identities, strategies, and acquisition targets would not be disclosed. That secrecy attracted attention because converting physical books into AI training data can involve destructive scanning: cutting off their bindings, feeding the loose pages through industrial scanners, and discarding or recycling the originals.
By July 28, ISBNdb had removed the sourcing page and its promise of an NDA. The company said the service had been part of “exploring demand” and that it had “chosen to pivot away from that direction,” in a news update.
Still, ISBNdb’s brief sales pitch offered a glimpse into a book-buying operation that one major AI developer has already carried out on a much larger scale. Federal court records show that Anthropic purchased and scanned millions of physical books, while used-book sellers in the United States and Europe have recently reported unusual bulk orders for obscure and out-of-print titles.
The reports have opened several practical questions: Why have older books become so valuable to AI developers, how widespread is destructive scanning, and what happens to the originals once their pages become training data?
Why AI companies are hunting for old books
Older books offer something that has become surprisingly difficult to guarantee online: writing produced entirely by humans.
This Tweet is currently unavailable. It might be loading or has been removed.
Large language models are trained on enormous quantities of text collected from websites, articles, books, code repositories, and other digital sources. Since generative AI tools became widely available, however, the internet has filled with machine-written summaries, product listings, social posts, and articles.
That creates a problem for developers assembling new training datasets. If a model is trained too heavily on material produced by earlier models, it can begin reinforcing their mistakes while losing some of the variety and less common information contained in the original human data. Researchers call the process “model collapse.”
A physical book printed before the current AI boom offers a relatively clean alternative. Unlike a webpage that may have been quietly generated or rewritten by a chatbot, an older book provides a more dependable record of human writing. Books are also edited, structured, and often contain specialized information that cannot easily be found elsewhere online.
ISBNdb leaned heavily on those qualities in its sales pitch.
“Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate,” the company wrote. “Dense, edited, authoritative.”
ISBNdb also emphasized that many potentially useful titles had never been fully digitized. Its marketing effectively presented the pre-chatbot bookshelf as a reserve of human-created training material that had yet to be mixed with AI output.
That demand did not appear out of nowhere. It follows a much longer effort to turn printed books into searchable digital information.
AI didn’t invent mass book scanning
This Tweet is currently unavailable. It might be loading or has been removed.
People have been converting books into digital files for decades. Project Gutenberg began creating electronic versions of public-domain works in 1971 and now offers more than 75,000 free ebooks.
The industrial-scale version arrived with Google Books in 2004. Working with libraries and publishers, Google scanned books and made their contents searchable online. By 2019, the company said it had assembled a collection of more than 40 million books in over 400 languages.
Those scans also helped form HathiTrust, a digital repository created by research libraries in 2008 to preserve their collections and make as much of the material publicly accessible as copyright law allowed.
The project prompted an earlier version of today’s copyright fight. Authors sued Google for scanning copyrighted books without permission, but a federal appeals court ruled in 2015 that the project qualified as fair use. The court found that creating a searchable index and displaying limited snippets gave the books a new purpose without providing readers with a replacement for the originals.
Other digitization programs have emphasized preservation and access. The Internet Archive, for example, says it has digitized more than 25 million books since 2006 using nondestructive scanning designed to keep the bound volumes intact.
This Tweet is currently unavailable. It might be loading or has been removed.
But converting a book into data does not necessarily make it publicly accessible. That distinction concerned Charlie D. Becker, a second-generation bookseller whose family runs Becker’s Books in Houston. His store recently received a single order for 70 obscure titles, including a 1995 guide to metropolitan Denver and manuals explaining how to use WordPerfect in 1991.
After investigating, Becker suspected the buyer was using an algorithm to find underpriced books and relist them on Amazon, rather than acquiring them for AI training. Most of the titles were either unavailable on Amazon or listed there for as much as 20 times his store’s price, and the shipments appeared to be going to Fulfillment by Amazon preparation companies.
This Tweet is currently unavailable. It might be loading or has been removed.
Still, Becker said an obscure book can be especially easy to lose because few people consider it worth preserving. Books that fail to sell through Amazon’s fulfillment system may eventually be liquidated, while some titles have little more than scattered listings across private databases to prove they existed.
Mashable Trend Report
“Everyone assumes the internet preserved everything,” Becker wrote in a longer account of the orders. “It didn’t.”
Anthropic has already scanned millions of books
The buyers behind the recent bookstore orders remain unclear. The destructive scanning process, however, is no longer hypothetical.
This Tweet is currently unavailable. It might be loading or has been removed.
In 2024, Anthropic hired Tom Turvey, a former Google executive who had worked on partnerships for the Google Books project. According to a June 2025 federal court ruling, Turvey was tasked with helping the company obtain “all the books in the world” for an internal research library.
Turvey initially contacted publishers about licensing their books, but those conversations did not continue. His team then approached major book distributors and retailers about purchasing print copies in bulk.
Anthropic ultimately spent millions of dollars acquiring millions of physical books, many of them used. Service providers removed the books from their bindings, cut their pages to the appropriate size, and fed them through scanners. The searchable PDF files went into Anthropic’s internal library. The paper originals were discarded.
Engineers could then select groups of those books for inclusion in datasets used to train the large language models behind Claude.
This Tweet is currently unavailable. It might be loading or has been removed.
That process became central to a legal battle over how Anthropic obtained its training material. In June 2025, U.S. District Judge William Alsup ruled that the company’s use of books to train Claude was transformative and qualified as fair use under the specific circumstances of the case.
Alsup also found that converting legally purchased print books into digital files could qualify as fair use because Anthropic destroyed each physical copy and replaced it with one internal digital copy. In other words, the company did not keep both versions.
The ruling did not give AI companies blanket permission to copy any book they could find. Alsup drew a sharp distinction between the print books Anthropic had legally purchased and the more than 7 million pirated books the company had downloaded and stored in a permanent digital library. Purchasing physical copies of some titles later did not erase the original piracy, he found.
That distinction eventually became expensive. On July 20, a federal judge approved Anthropic’s $1.5 billion settlement with authors and publishers, resolving claims involving approximately 482,000 pirated books. Eligible rights holders are expected to receive about $3,000 per title, according to Reuters.
For some authors, the payment does not resolve the larger disagreement. Charles Graeber, one of the case’s original plaintiffs, told NPR that he was proud authors had secured a substantial settlement, but said the case had cost him more than two years of time, travel, and professional opportunities.
Fellow plaintiff Andrea Bartz questioned a system that allows companies to train commercial models on legally purchased books without negotiating separate licenses with the people who wrote them.
“The algorithm is being used to essentially try to put us out of a job,” she told NPR.
Anthropic has maintained that training AI models on books is protected by fair use. The court’s ruling nevertheless helps explain why physical books may be especially attractive to developers: A lawfully purchased copy gives the company a much stronger legal position than a file downloaded from a pirate library.
Destroying the original may be a legally useful distinction. For many readers, it is also the most unsettling part of the story.
Who is buying all these books?
Anthropic’s operation is documented in court records. The source of the more recent bookstore orders is harder to pin down.
This Tweet is currently unavailable. It might be loading or has been removed.
One bookseller specializing in uncommon and low-circulation titles told 404 Media that his weekly sales jumped from roughly 20 books during a good week to several hundred after the orders began arriving in April.
The requested titles did not appear to share a subject, author, genre, or language. They did, however, all have International Standard Book Numbers, or ISBNs, leading the seller to suspect that they had been selected through a book database.
The surge was financially helpful and allowed him to clear inventory that might otherwise have remained unsold. He was less enthusiastic about where the books might be going.
“I don’t like the end-use, and I don’t like that uncommon books are being pulped,” he said.
Booksellers in Europe have reported similarly broad requests. An antiquarian bookseller in the Netherlands received a list of approximately 3,000 English-language books organized by ISBN, ranging from an academic study of Irish folklore to a technical book about laser shock peening.
There is no public confirmation that every unusual bulk order came from an AI company or that every book purchased through these orders was destroyed. There is also no evidence that developers are intentionally hunting for the final surviving copies of rare titles.
The uncertainty itself is part of the concern. A company purchasing from a massive ISBN list could sweep up uncommon or out-of-print editions without first checking how many physical copies remain.
Why the story struck a nerve
Once the reports reached social media, they were accompanied by an unsettling visual: a cutting blade moving inch by inch through a book’s spine. Comparisons to Fahrenheit 451 and the Library of Alexandria followed quickly.
This Tweet is currently unavailable. It might be loading or has been removed.
This Tweet is currently unavailable. It might be loading or has been removed.
This Tweet is currently unavailable. It might be loading or has been removed.
This Tweet is currently unavailable. It might be loading or has been removed.
Book destruction carries a particular weight because it has historically represented more than the loss of paper. The 1933 Nazi book burnings targeted works deemed “un-German,” including books by Jewish, pacifist, and left-wing writers. The act became an enduring symbol of censorship and the suppression of ideas, according to the United States Holocaust Memorial Museum.
That symbolism is central to Fahrenheit 451, Ray Bradbury’s 1953 novel about a society where books are outlawed and burned. Bradbury said his warning extended beyond government censorship to television reducing knowledge to digestible fragments and eroding interest in reading. Users have also invoked the Library of Alexandria, another enduring symbol of lost knowledge, although historians believe it declined gradually through political upheaval, reduced support, neglect, and repeated damage rather than disappearing in one catastrophic fire.
Elon Musk joined the conversation on July 27. He wrote on X that he had asked the SpaceXAI team to preserve rare books in a library and scan them “the hard way,” without removing their spines.
This Tweet is currently unavailable. It might be loading or has been removed.
As the discussion spread, users resurfaced a 2011 Cracked article about libraries, universities, and retailers destroying unwanted books years before the current AI boom.
That history has informed a less alarmed response. Some users argued that the books shown in warehouses appeared to be ordinary, unwanted inventory rather than irreplaceable artifacts. If a book would otherwise be recycled without being read again, they asked, could scanning it first preserve something that would have been lost?
This Tweet is currently unavailable. It might be loading or has been removed.
This Tweet is currently unavailable. It might be loading or has been removed.
AI adds a complication. The text may survive, but inside a private company’s research library rather than a public archive. A forgotten manual or travel guide may have little resale value, yet still contain exactly what an AI developer wants: edited human writing created before the flood of chatbot output.
That has led authors and publishers to argue that they should have more control over how their work is used, especially when it is helping companies build commercial products. AI developers, meanwhile, continue to argue that training a model is a transformative use of the material rather than a replacement for the original books.
ISBNdb has removed its sourcing page, but the demand behind it remains. The web’s AI problem has sent developers back to the bookshelf. The next chapter will depend on whether they can extract what they need without leaving those shelves any emptier.
Topics
Artificial Intelligence
Books
