The revelation that artificial intelligence companies are destroying printed books to train their models has unleashed outrage among librarians, booksellers, and culture enthusiasts. While tech giants buy massive lots of obscure titles and shred them after scanning, the Internet Archive demonstrates that there is another way: carefully scanning each page by hand without destroying the original.
The destruction of books to train artificial intelligence models has become an increasingly common practice in Silicon Valley. Tech companies in fierce competition to develop the most advanced models are resorting to the cheapest and fastest method: buying printed books, cutting their pages, scanning them, and then shredding the copies.
This dynamic was uncovered by a lawsuit in the summer of 2025 that revealed Anthropic destroyed millions of printed books to train its models. The company used the codename 'Project Panama' to conceal the operation from the public.
The revelation sparked a wave of outrage among those who value the preservation of bibliographic heritage. The idea of piles of book spines feeding wood chippers is particularly painful for lovers of rare editions.
Destructive scanning is the cheapest way to process books en masse. AI companies are in a race to improve their models with long, high-quality texts available only in book format.
Preservation advocates fear that the practice may occur on a much larger scale than reported. And worst of all, some physical copies of rare books could be lost forever.
The destruction of books does not have to be inevitable. Google patented a non-destructive scanning technology in 2009 that AI companies could implement if they were willing to slow down and invest more resources.
However, Google's method is not perfect. Studies have found that the curvature of pages can distort text, and some pages may be omitted during the process.
When workers move too quickly, additional failures occur. There have been documented instances of disembodied hands obscuring pages in scans, compromising the quality of the digitized material.
The Internet Archive has long understood that scanning ancient texts requires time and special attention. The organization helps libraries preserve aging collections through a process that prioritizes the physical integrity of the book.
The Internet Archive's method may seem almost alien to AI companies looking for shortcuts. But it is precisely that dedication that protects copies that could never be replaced if destroyed.
Eliza Zhang, a book scanner at the Internet Archive since 2010, described her work in a 2021 post. For her, scanning 'the hard way' has an intrinsic value that transcends efficiency.
Zhang has scanned over 3 million pages during more than a decade of work. She has also processed over 14,000 foldouts and 18,000 items, mostly books.
Her goal is to achieve 'zero errors.' Each page requires intense concentration because old books have extremely fragile leaves, 'thin as paper.'
Andrea Mills, who helps oversee scanning operations, explained that 'clean, dry human hands are the best way to turn pages.' Automated scanners with vacuum arms do not work well for rare volumes.
Internet Archive tested commercial automated scanners, but the results were disappointing. For fragile books and special collections, human intervention remains irreplaceable.
Zhang takes his time with each work. He carefully lifts the scanner's glass with a pedal, adjusts the cameras, and checks that each page is legible before proceeding.
He also records any fold-out pages, setting reminders to return and scan additional inserts. This ensures that supplementary materials are not lost during the documentation of the main work.
If a page is missed or an image is blurry, Internet Archive's proprietary software halts the process. Zhang must re-scan the defective material before moving on.
Chris Freeland, director of library services at Internet Archive, confirmed that the 2021 publication remains 'the best description of our scanning process today.' The method has not changed fundamentally.
A video of Zhang's work reached 1.5 million views on Twitter at the time. The caption resonated with book lovers: 'We never destroy a book by cutting its binding.'
Since the lawsuit against Anthropic revealed the mass destruction of books, culture enthusiasts have mobilized. The greatest fear is that AI companies will carelessly pulverize copies that can never be replaced.
Last week, two recent reports intensified the alarm. The newspaper The Telegraph accused Silicon Valley of destroying millions of rare books and 'grinding the originals.'
Media outlets like 404 Media reported that ISBNdb, a book database, was promoting its services to help AI companies acquire books in bulk. The company allegedly warned its clients that the optics were poor.
Anthropic began using the codename 'Project Panama' to conceal its destructive scanning operation. There is no evidence that the company has destroyed rare books, and it has denied this in public statements.
However, rare book sellers continue to detect strange orders that raise suspicions. Legitimate buyers tend to look for individual books, not large lots of very varied selections.
Just this week, the Irish bookstore Kennys flagged a 'bananas' order of 5,000 obscure titles. The buyer did not attempt to haggle on the price, which has become a 'red flag' of possible AI involvement.
Tomás Kenny of Kennys Bookshop stated that AI is 'terrifying our industry.' He plans to reject orders likely intended for training models to protect the rights of authors and readers.
Kenny asserted that booksellers are often at the forefront of moral and ethical dilemmas. But in the case of AI taking intellectual property, 'Kennys does not want to be part of that.'
Some booksellers fear that selling to AI companies that indiscriminately destroy their collections is 'signing their own death warrant.' The short-term benefit does not justify the long-term harm.
The community of readers and librarians is calling for transparency about which books are being purchased and for what purpose. However, the lack of regulation leaves AI companies operating in a gray area.
Some tech companies are collaborating with institutions to preserve rare books. OpenAI and Microsoft are working with Harvard librarians on an initiative to train models with approximately 1 million public domain books from the 15th century.
Elon Musk's xAI has publicly declared that it will not destroy rare books to train AI. Musk asked his team to preserve any rare books and scan them 'the hard way' instead of cutting the spine.
However, critics are not entirely convinced. Musk did not say that his team would not destroy any books, and it is unclear how his company defines a 'rare' book.
Copyright law allows companies to destroy copies of books they have purchased. This provides a legal framework, albeit ethically questionable, for the destructive practice.
Public reaction could pressure big tech companies to adopt non-destructive methods like those of the Internet Archive. Social pressure has previously forced AI companies to modify controversial practices.
However, the economic incentive still favors destruction. Scanning millions of books using non-destructive methods requires more time, money, and trained personnel.
The Internet Archive has demonstrated that it is possible to digitize at scale without sacrificing the originals. But its model relies on a preservation mission, not on maximizing the speed of model training.
For book lovers who feel distressed by headlines about destruction, the Internet Archive's method offers a balm. There is an alternative that respects both knowledge and the physical object.
The appeal of rare books goes beyond their content. The experience of holding an old volume, with its yellowed pages and century-old binding, connects the reader with past generations.
Zhang, the scanner from the Internet Archive, occasionally pauses her work to read snippets from the titles she processes. Each collection is important to her.
When asked what she liked most about her work, she responded with contagious enthusiasm. 'Everything! I find everything interesting,' she said, adding that she does not feel the work is boring.
That enthusiasm stands in stark contrast to the industrial logic of shredding books to extract data. The difference in approach reveals two opposing visions of knowledge and its preservation.
If the destructive practice continues unchecked, rare books will become even scarcer. Future readers may lose access to physical copies that we today consider irreplaceable.
The community of booksellers, librarians, and readers demands accountability. Meanwhile, companies like Kennys are making the decision to reject suspicious orders, even if it means forgoing income.
The debate is far from resolved. The demand for mass destruction of books could set important legal precedents for the AI industry.
Preservation advocates hope that transparency and public pressure will force companies to change. Each shredded book represents an irreversible loss for shared cultural heritage.
The Internet Archive's method demonstrates that efficiency does not have to be at odds with respect for the originals. True innovation might consist of scanning more slowly, not faster.
For those who value written culture, the battle against book destruction is also a battle for the future of knowledge. What is at stake transcends artificial intelligence.
The final decision will rest with companies, regulators, and consumers. For now, booksellers like Tomás Kenny have chosen a side: that of books, not data.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.





























