When does a book's economic value separate from the story and ideas within it, reducing it to the words it contains?
For Michael Ströter, a second-hand bookseller in Bielefeld, Germany, this question took on concrete form in the spring of 2026. Under normal circumstances, Ströter sells a few books a day, but his system began receiving an unusual string of orders in the early morning hours. Even technical books that had gone unsold for years and held extremely limited commercial value were being purchased. The orders did not target popular works; they focused on single-copy, niche-area books that were often not easily accessible in digital form.
At first glance, this activity could be seen as a new wave of automation in the global second-hand book trade. Algorithms may be identifying low-priced books across different markets, with large sellers collecting them in central warehouses and offering them for resale at higher prices.
However, copyright lawsuits filed in recent years and company documents made public have brought another, more unsettling possibility to the fore: Books are no longer seen merely as cultural products to be resold, but also as strategic raw data materials that can be used to train large AI models.
Generative AI systems do not need more text; they need text that is coherent, curated, and rich in domain knowledge. Although the internet has met this need on a large scale, the repetitive, low-quality, and mislabeled nature of digital content is making books valuable once again.
A book is not merely an object made of stacked paper. It is a structure of thought filtered through experts and editors, the accumulated knowledge of a specific era, and often a record of expertise and emotion that has no other equivalent on the internet.
Therefore, discussing the growing role of books in the AI industry solely through the lens of copyright constitutes an incomplete approach. This situation heralds a new data supply chain, a fresh licensing market, and a comprehensive transformation in the economic value of knowledge.
The traces of this transformation may be visible today in Europe's used-book shops. However, the story did not begin on physical bookshelves, but in libraries tucked away in the obscure corners of the internet.
The Past: From Pirate Archives to Physical Supply
Meta and LibGen: Risk Despite Everything
The role that books played in the development of large language models remained unclear for a long time in the technical reports published by technology companies. Training data was mostly described with broad categories such as "publicly available sources," "web data," or "licensed content."
Copyright lawsuits filed by authors gradually brought the processes behind these general statements to light.
The lawsuit filed by Richard Kadrey, Sarah Silverman, and other authors against Meta became one of the most significant examples of this process. The plaintiffs alleged that Meta used pirate book archives such as Library Genesis (LibGen) to train its Llama models.
Internal communications disclosed during the court proceedings showed that Meta employees were aware of the legal and reputational risks of LibGen. The documents revealed some employees expressing reservations about torrent usage, the removal of copyright information, and the public disclosure of the sources used. Records indicating that the usage decision was escalated to senior management also entered the case file.
These documents suggest that book data did not accidentally end up in the training corpus; it was deliberately evaluated within the framework of competitive model development goals.
Meta argued that using copyrighted works in model training fell within fair use. The company also claimed that authors had failed to demonstrate concrete substitution or economic harm in the market for their works.
The significance of the Meta case is not that a pirate archive was used. The lawsuit made visible the calculation AI companies make between legal risk and technological competition in their data decisions. Under the pressure to develop more powerful models, the legitimacy of data had become a risk item evaluated alongside data volume and model capacity.
The book was no longer cultural heritage; it was becoming a competitive advantage.
OpenAI: The Memory of the Dataset Matters as Much as the Data Source
While the Meta case focused on which sources were used, the proceedings against OpenAI revealed a different problem. When a training dataset is eliminated, how can its origin and usage history be audited?
According to court records, OpenAI deleted two datasets named Books1 and Books2 in 2022. The company initially stated that the datasets were deleted because they were no longer in use. The plaintiffs, however, requested documents regarding the reasons for the deletion decision, internal communications related to LibGen, and the legal or commercial nature of the decision.
In September 2025, the court requested additional clarification from OpenAI regarding its shifting positions on the reasons for the deletion. In November 2025, a magistrate judge ordered the production of certain internal communications and limited depositions of relevant in-house attorneys. However, this ruling was overturned by a district judge in February 2026. Consequently, the legal status of the communications in question and the obligation to produce documents have returned to a point more uncertain than the initial ruling suggested.
This distinction is important. Court records confirm that Books1 and Books2 were deleted; however, they do not definitively establish that the deletion constituted intentional spoliation of evidence. Similarly, allegations regarding the datasets' connection to LibGen and internal communications should be treated as part of the ongoing legal battle.
The OpenAI case demonstrates that data provenance is not merely a matter of the question, "Where did this text come from?" Other equally important questions exist:
- When was the data collected?
- In which model was it used?
- How was it transformed?
- When and why was it deleted?
- Who made the decision?
- How long are transaction records retained?
In financial systems, money movements, in supply chains, material flows, and in scientific research, data transformations are recorded, yet the training datasets shaping billion-parameter models lack similar oversight—a structural governance problem.
Anthropic: From Digital Piracy to Book Processing
The most concrete example of this transformation in book data sourcing was Anthropic's Project Panama operation.
The lawsuit filed by Andrea Bartz and other authors against Anthropic revealed that the company, on one hand, downloaded books from digital shadow libraries such as Books3, LibGen, and Pirate Library Mirror (PiLiMi), while on the other hand, it purchased physical books to build a separate data archive.
Judge William Alsup deemed the use of books to train AI models a transformative activity. In contrast, he treated the downloading of books from pirated sources to create a permanent central library as a separate legal matter.
This distinction created a critical economic incentive. While downloading a pirated digital copy of a book exposed the company to high copyright risk, purchasing a physical copy and digitizing it in-house appeared to be a more defensible method.
Project Panama was built precisely on this distinction.
The project, which gained momentum in 2024, was described in an internal document filed in court as an effort to "aggressively scan all books in the world." The documents also showed that the company did not want the project to become publicly known.
It emerged that within roughly a year, tens of millions of dollars were spent to purchase millions of physical books, have their bindings cut, and their pages run through high-speed scanners. Anthropic sourced the books in batches of tens of thousands from major second-hand book sellers such as Better World Books and World of Books.
One scanning supplier's proposal projected processing 500,000 to 2 million books within six months. The process involved separating book bindings with hydraulic cutting machines, running pages through scanners, and delivering the processed physical copies to recycling companies.
Although the final number of books and total cost were redacted in the documents, the scale of the operation shows that physical book sourcing was not an experimental endeavor but an industrial data production process.
This model divided the book into two distinct entities:
- The cultural and physical object
- The machine-processable data
From the company's perspective, economic value concentrated in the second entity; the binding, paper, printing features, margin notes, and the circulation history of the physical copy became waste.
The Evolution of the Data Supply Chain
The Meta case focuses on the use of pirated digital archives in model training; the OpenAI process on the origin, deletion, and record-keeping of training datasets; and the Anthropic case on the distinct legal statuses of pirated digital sources versus legally purchased physical books.
But read together, a common economic trend emerges.
In the initial phase, large book collections on the internet were used as a low-cost and fast data source. As legal risks, author lawsuits, and licensing demands increased, data collection processes became more complex. Companies moved not only toward acquiring more books but also toward building a provenance chain that could justify how the books were acquired.
This shift does not mean a complete abandonment of piracy in favor of a transparent and licensed data economy. On the contrary, legal uncertainty may have made data sourcing methods more expensive, more physical, but less visible.
Today: A Wave of Orders Seen in Used Bookstores
Project Panama was a documented operation from the past. The current picture in Europe is more ambiguous, but economically at least as striking.
Various used bookstores in Germany began receiving bulk orders in 2026 that fell outside their usual sales patterns. The requests were often transmitted automatically during the night, targeted niche works that had not sold for years rather than popular or collectible books, and typically purchased a single copy of a given title.
Michael Ströter, who normally sells at most a few books a day, reported that his first unusual order came through his system on April 30 at 2:53 AM. Among the purchased works was a technical book from the 1970s focused on a topic with extremely limited commercial demand, such as television education in daycare centers.
The fact that the orders were placed on behalf of Canada-based Zoom Books and later routed to a logistics hub in the Germany–Poland–Czech Republic border region heightened the booksellers' suspicions. Reports from taz and SRF indicated that similar automated ordering patterns were observed in different parts of Germany.
These observations do not definitively prove an AI data sourcing operation. However, they do indicate the emergence of an algorithmic purchasing model that diverges from the normal demand system in the second-hand book market.
Zoom Books: Categorical Denial, Undisclosed Customers
At the center of current debates is Canada-based Zoom Books.
In a statement to Publishers Lunch, the company described claims that it digitizes or destroys books to train AI models as "categorically false."
Zoom Books executive Reed Pannell also told taz that the company does not digitize or destroy books. However, when asked whether the company sells the books it purchases to other businesses, he responded that they cannot disclose information about their customers.
A company not disclosing its customers is standard behavior in terms of trade secrecy and does not, by itself, prove a suspicious relationship. In contrast, Zoom Books keeping its end-customer chain confidential makes it impossible to independently verify the economic purpose behind the purchases.
2077AI: A More Explicit but Still Incomplete Link
The case of Pieter de Vries in the Netherlands makes the AI connection in the current wave of orders more visible.
De Vries, a second-hand bookseller operating in Haarlem, received a direct email from an individual at the Singapore-based company 2077AI. Attached to the email was a list containing the ISBNs of approximately 3,000 English-language books.
The books on the list were not popular works aimed at a broad audience. Instead, technical and academic publications focusing on geomechanical modeling, surface treatment, folklore, and other specialized fields had been selected.
This example carries a stronger AI connection than the Zoom Books case. The buyer's company name, the method of communication, and the nature of the selected books increase the likelihood that the texts could be used for artificial intelligence or data processing purposes.
Although the 2077AI case is a strong signal, it does not constitute proof.
Data Mining or Book Arbitrage?
The hypothesis that bookseller orders are linked to AI training is plausible given Project Panama's documented history. Nevertheless, there are other commercially strong explanations for the current purchases.
The AI Data Supply Hypothesis
Indicators supporting this hypothesis include:
- Selection of niche books that are difficult to find in digital form
- Usually purchasing only a single copy of a book
- Targeting academic works with no commercial popularity
- Orders appearing algorithmic and automated
- Books from different countries being consolidated at central hubs
- AI-linked actors like 2077AI sending book lists
- Project Panama having previously used similar physical supply methods
These indicators form a pattern. However, a pattern does not substitute for a contract, payment records, or internal company documents.
The Book Arbitrage Hypothesis
An alternative explanation is a sophisticated second-hand book arbitrage model. In this model, large buyers automatically scan book markets in different countries to identify low-priced, scarce books. The books are collected in central warehouses, repriced, and listed for sale at higher prices on platforms like Amazon, AbeBooks, or other international marketplaces.
This strategy can be particularly profitable for:
- out-of-print books,
- specialist publications,
- limited academic works,
- single copies with no other sellers,
- technical publications needed by libraries or researchers.
Once a company becomes the sole or primary online seller of a specific book, it can gain pricing power even if demand is low.
Therefore, purchasing a book that hasn't sold for years does not necessarily mean its text will be used for AI training. The physical scarcity of the book can also hold economic value.
Two Uses of the Same Logistics Infrastructure
The most realistic possibility is that data supply and book arbitrage are not entirely separate worlds.
Both models require similar infrastructure:
- Algorithms scanning different markets,
- ISBN-based inventory matching,
- Automated price comparison,
- Global logistics hubs,
- High-volume bulk purchasing,
- Physical sorting,
- Storage and recycling.
This infrastructure could have been built for resale; it could later supply books to AI developers or digitization firms. Conversely, a chain established for data acquisition could divert a portion of physical books to resale.
Why Are Books So Valuable?
Explaining AI companies' interest in books solely through data volume is misleading.
The internet is an extremely large source of text. However, size does not mean quality: repetitive text, SEO-generated content, automated material, incomplete sources, etc.
Books, on the other hand, generally have a different structure. Publishing a book involves not only the author's work but also processes like editing, proofreading, design, and source verification. Furthermore, older books carry linguistic forms, local narratives, and ideas not represented in today's digital environment. The resulting text offers a refined output with documentary value.
Another issue in the data economy is source tracking. A large-scale audit published in the journal Nature Machine Intelligence showed that licensing information in examined text datasets is often missing or misclassified. Researchers examined over 1,800 datasets and identified serious gaps in tracking sources, producers, licenses, and subsequent uses.
Can Synthetic Data Replace Books?
Synthetic data consists of text, images, code, or structured examples generated by other models. Since it allows companies to produce their own training data, it is thought to offer solutions to copyright, privacy, and data scarcity issues.
However, the debate often runs to two extremes:
- Synthetic data can completely replace human data.
- Models inevitably collapse when trained on their own generated content.
The picture emerging from research is more complex.
The degradation known as model collapse can occur when models are continuously trained on data produced by previous generations of models, and original human data is removed from the system. Rare examples may be lost, diversity may decrease, and errors can be passed on to new generations.
In contrast, accumulating synthetic data while preserving real data can prevent model collapse in some experiments. Other studies also show that training with synthetic data is risky, but degradation can be reduced in controlled mixtures of real and synthetic data.
Data scaling research predicts that if current trends continue, training demand could be met by the public stock of human text during the 2026–2032 period; however, it also notes that pathways such as synthetic data, data efficiency, and transfer learning could sustain scaling.
The Future: A New Data Economy Built Around Books
Based on this information, a new data economy is likely to emerge among libraries, authors, and technology companies.
This economy could develop in three different directions.
Licensed Data Market
AI companies could enter into direct licensing agreements with publishers, author associations, and digital archives. This approach could increase legal security. In contrast, the bargaining power of large publishers could exclude small publishers, independent authors, and low-resource languages from the system.
Growth of the Gray Market
Rising licensing costs or unclear copyright regimes could steer companies toward second-hand markets and intermediary suppliers.
The most significant risk of this model is that not only the ownership but also the accessibility of the purchased book is withdrawn from the market.
Destroying a rare book after scanning it allows the text to live on company servers while ending the physical work's public circulation. The digital text does not preserve the typography, paper, marginalia, or traces of previous owners.
Mandatory Provenance and Collective Licensing
Regulatory bodies could require model developers to provide more detailed disclosures about the origin of training data.
Here, not only copyright law but also competition law and information access policies will gain importance. A few technology companies aggregating the world's vast book collections in closed databases affects not just authors' rights but also the future of the information infrastructure.
The Erasure of the Known Concept of the Book
Meta demonstrated the use of pirated digital archives in model training; OpenAI, the preservation of the origin and records of training datasets; and Anthropic, how purchasing physical books could be transformed into an industrial data production process.
The common thread in these cases is that books have acquired a new status in the AI economy.
The book is no longer merely a cultural product to be read, borrowed, collected, or resold. It has also become:
- a data source that improves model performance,
- a competitive advantage for companies,
- a licensing asset for publishers,
- a source of income and a rights struggle for authors,
- a reproduction and fair use issue for legal systems,
- an object of algorithmic demand for used book dealers,
- a cultural heritage risk for society.
Project Panama, carried out by purchasing millions of books, cutting them apart for scanning, and then sending them to recycling, revealed how this transformation could occur on an industrial and physical scale. Thus, it became clear that the data economy is not solely composed of servers and algorithms.
This economy can also be built on warehouses, shipping networks, hydraulic cutters, high-speed scanners, and second-hand book markets.
The current wave of orders in Europe, however, presents an unfinished story. Unusual purchases and algorithmic patterns can be verified. Some requests come directly from AI-linked actors.
However, there is no direct evidence showing that companies like Zoom Books are acting on behalf of major AI labs or that purchased books are being scanned and destroyed as in Project Panama.
There is a strong resemblance between a previously documented data acquisition model and the market anomaly seen today. But resemblance is not causation. Therefore, the real area to monitor in the future is not the shelves of bookstores, but the paths books take after leaving the shelves.
Because long before the AI era, a book's true value lay not in its cover, edition, or sale price, but in the ideas and knowledge it carried. The only thing that has changed today is that this value is now also recognized by machines.
References
- United States District Court, Northern District of California. Kadrey et al. v. Meta Platforms, Inc.
- United States District Court, Northern District of California. Bartz et al. v. Anthropic PBC
- WIRED. Meta Secretly Trained Its AI on a Notorious Piracy Database, Newly Unredacted Court Docs Reveal.
- In Re: OpenAI, Inc. Copyright Infringement Litigation
- The Washington Post — Inside an AI Start-up’s Plan to Scan and Dispose of Millions of Books
- taz — Artificial Intelligence: Buy, Scan, Feed
- SRF — Hunting for Old Books: AI Companies Buy Out Antiquarian Bookshops – and Destroy the Books
- Publishers Lunch — Zoom Books Responds to Accusations of Selling Books for AI Scraping
- NL Times — Rare Book Dealers Fear Tech Firms Are Destroying Obscure Editions to Train AI Models
- Longpre et al. — A Pretrainer’s Guide to Training Data
- Villalobos et al. — Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data
- Gerstgrasser et al. — Is Model Collapse Inevitable?
- Seddik et al. — How Bad Is Training on Synthetic Data?


