Author: David, Deep Tide TechFlow
Burning books to raise AI is becoming a global business.
According to a recent report by Fortune, since the beginning of this year, antiquarian booksellers in the Netherlands, Germany, Switzerland, and Spain have received unusual large orders in succession. Buyers want thousands of academic books at once, with a variety that does not resemble normal procurement.
For example, an antiquarian bookseller based in Haarlem, Netherlands, de Vries, recently received an outrageous email:
The sender is from a Singapore company named 2077AI, claiming to be working on a multilingual book collection project; the email attachment is an Excel spreadsheet listing a procurement list of more than 3,000 books on various subjects, each marked with a publication number, and requesting a quote to be sent to China.
De Vries thought it was too absurd at the time, so he ignored it.
A few weeks later, when a Dutch reporter came knocking, he realized that behind this "scam email" was connected to the industry's insatiable hunger for training data in the AI sector.

As for who is using this data to train which models, the company buying the books refused to respond, according to Fortune.
Not just this one. Another Canadian company, Zoom Books, places bulk orders on European booksellers' online stores every day between three to five in the morning, buying obscure academic books that have no relevance to each other.
Some booksellers told the media that shipping costs are more expensive than the books themselves, but the buyers seem unconcerned.
These books are obviously not bought for reading.
AI companies need physical books published before 2022 because these texts guarantee they have not been contaminated by AI-generated content, making them the cleanest corpus for training large language models.
After acquiring the books, the common industry practice is "destructive scanning," which involves cutting off the spines, rapidly scanning all pages, and then destroying the physical originals.
Abroad, top players in the AI large model industry have already run through this entire process in a more "wolf-like" manner.
Buy pirated books, raise AI
The complete record of this matter comes from a copyright lawsuit in the United States.
In August 2024, American authors Andrea Bartz and two others sued Claude's parent company, Anthropic, accusing it of using their works without authorization to train AI models. The case progressed to early 2026, when over 4,000 pages of internal documents from Anthropic were unsealed by the court, and The Washington Post detailed this batch of documents.
The story revealed by these documents is much more brutal than one might imagine:
Between 2021 and 2022, Anthropic downloaded over 7 million books from the piracy ebook sites LibGen and Pirate Library Mirror via BitTorrent, storing them on internal servers. The documents quote management as believing that negotiating authorization with publishers "was too cumbersome."

Image: A corner of the publicly disclosed court documents
The three authors who initiated the lawsuit were just the first batch. The case soon expanded into a class action lawsuit involving approximately 500,000 works. In June 2025, Judge Alsup made a split ruling:
Training AI with books counts as fair use, but downloading and permanently storing them from piracy sites does not. In other words, the act of training AI is not illegal, but the method is questionable.
The maximum compensation Anthropic faced was $150,000 per book, with 500,000 works, which theoretically amounts to astronomical figures. Ultimately, both parties settled for $1.5 billion, averaging about $3,000 per book.
On July 20 of this year, the court finally approved it, making it the largest settlement in U.S. copyright history.
First steal, then burn
If the story ended here, it would just be a simple script of being caught stealing and paying up.
But according to a January 2026 report by The Washington Post, during the court proceedings, the judge unsealed over 4,000 pages of Anthropic’s internal documents that recorded another plan.
At the beginning of 2024, Anthropic launched an internal operation codenamed "Panama Project," specifically hiring former heads from the Google Books era to purchase physical books in bulk from second-hand book platforms, planning to process between 500,000 to 2 million books within six months.
Buy, cut off the spines, scan, and destroy the originals.
The internal documents defined this project as "destructive scanning of books worldwide," and explicitly noted "We do not want the outside world to know we are doing this."
Compared to directly downloading from piracy sites, the difference is that this time the books are purchased with money.
The judge later ruled this practice legal because buying one book and destroying the physical version while keeping the digital version means there is only one copy, not constituting additional copying, thus conforming to copyright law.
So, Anthropic's case effectively paved a way for the entire industry to bypass legal risks through using antiquarian books to train AI:
Since stealing books from piracy sites is liable to lawsuits, then exploit the loophole, buy the legitimate ones and burn them, and it avoids trouble.
A hidden industrial chain
The path paved by Anthropic now has supporting infrastructure.
According to Decrypt, a U.S. federal judge has made similar fair-use rulings in separate copyright cases involving OpenAI and Meta. "Burning books to train AI" is no longer an isolated case for a single company but is gradually being "legitimized" as a common practice in the industry.
Where there is demand, there is supply.
According to 404 Media, a book database company with over 110 million book data, previously released a marketing page aimed at AI companies on its website, titled "Providing physical book corpus procurement services for your AI large model." The page reportedly offered bulk purchases from 1,000 to 1 million books and promised to protect buyers' identities through strict confidentiality agreements.
After the report emerged, ISBNdb quickly took down the page in the face of public backlash, stating on their official website that the related service "was never actually launched,” and that page was only a "market interest test."

Image: "I offered the service, but you all criticized me, so I took it down."
Believe it or not, but "testing" just revealed that market interest is indeed very high.
According to an anonymous bookseller quoted in the same report by 404 Media, since April of this year, his weekly sales have skyrocketed from about 20 books to hundreds of books. Buyers place precise orders using ISBN numbers, purchasing all unrelated obscure books, the only commonality being that all were published before 2022.
Why 2022?
Some of the selling points on the previously taken-down marketing page by ISBNdb actually exposed the issue, as physical books published before 2022 do not contain any AI-generated content, making them the cleanest training corpus, so it specifically chooses to sell such books to you.
So, let's sort out the upstream and downstream of this industrial chain:
- Database service providers supply selection lists and anonymous procurement channels,
- Second-hand booksellers across the globe provide the supply,
- Scanning service providers are responsible for cutting book spines and high-speed digitization,
- Destruction companies handle the paper remnants, and ultimately, the clean digital corpus enters the training pipeline of AI companies.
This practice has also drawn significant resistance and condemnation. Elon Musk expressed his opposition and sarcasm towards this "modern book burning" behavior on July 27 via X platform:

"I have instructed the SpaceXAI team to keep any rare books processed in a library and use 'dumb methods' (non-destructive ways) to scan them, instead of directly cutting off the spines and then scanning."
But without the public unsealing of court case documents, the names of the buyers might never appear in any of the processes, and no one would know about this secretive and slightly barbaric industrial chain.
Destruction is not the best solution
Scanning books for training is understandable, but the methods don't have to be so violent.
When Google was doing Google Books, it scanned over 40 million books using patented non-destructive scanning techniques, returning the books to the libraries without destroying a single one.
In July this year, Musk publicly stated on social media that he required his xAI team to use lossless scanning for rare books, preserving the original books in libraries.
This is technically entirely doable, just a bit slower and more expensive.
Anthropic chose the faster route. According to analysis from a now-deleted article on the ISBNdb blog, the reason destructive scanning became the industry’s preferred choice is simple: it’s cheaper and faster.
Non-destructive scanning requires special equipment and more manpower, with each book taking several times longer to process than destructive scanning.
In commercial competition, this choice is understandable. The arms race in the AI industry waits for no one; whoever gets more high-quality corpus first is ahead. The law hasn't really stopped it; the judge's ruling is clear: buying books and burning them is legal.
But between legality and reasonableness, there is sometimes a significant distance.
A physical book can be resold, lent, and inherited, lying for twenty years in a corner of an old bookstore before being discovered by a stranger. After being scanned and burned, it becomes private training data on a company’s server, never to be read by anyone again.
Knowledge shifts from a public good to private assets, and the completion of this transformation is not a complex technological barrier, but rather a cold cutting machine and a court ruling.
The author uses AI daily for writing, research, and organizing thoughts. These AI products are indeed useful and are changing the working methods of many people.
But if you look into how the training data behind these models came about, you'll find that many shiny products have some elements that do not stand up to scrutiny.
Each of these matters has its commercial logic and legal basis when viewed in isolation.
But viewed together, it’s hard not to feel something is awry, yet one cannot help but sigh.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。
