By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Booksellers Suspect AI Firms Destroy Rare Books for Training

Booksellers are voicing growing suspicions that artificial intelligence companies are engaging in the practice of purchasing rare books and subsequently destroying them to expedite the training of their AI models. This concern stems from the perceived race among AI firms to develop more advanced models by utilizing high-quality, long-form text found in books. The process, as described by those in the book trade, involves scanning pages after physically damaging the books, a method considered the fastest and most cost-effective for large-scale data ingestion. This destructive approach is particularly distressing to bibliophiles and collectors who value the historical and cultural significance of these physical artifacts.
The fear is that this practice is occurring on a much larger scale than currently reported, potentially leading to the permanent loss of unique and irreplaceable literary works. The fragility of old books, with their yellowing pages and delicate bindings, makes them vulnerable, and their destruction for AI training represents a significant cultural loss. Booksellers are particularly pained by the thought of these historical documents being reduced to raw data, losing their physical integrity and historical context in the process.
Furthermore, the book community highlights that such destructive methods are not the only viable option for AI training. Alternative, more preservation-minded approaches exist for digitizing book content without necessitating the physical destruction of the original copies. These methods, while potentially slower or more resource-intensive, would allow for the preservation of the books' physical form and historical value. The current trend, however, suggests a prioritization of speed and scale in AI development over the conservation of rare literary materials.
The implications of this suspected practice extend beyond the immediate loss of individual books. It raises broader questions about the ethical considerations in AI development, particularly concerning the acquisition and use of copyrighted and historically significant materials. The potential for irreversible damage to cultural heritage underscores the need for greater transparency and ethical guidelines within the AI industry regarding data sourcing and model training methodologies. The book trade is actively observing this situation, concerned about the long-term impact on literary preservation and the availability of rare texts for future generations.
Original source — read the full reporting at the publisher:
Read on Ars TechnicaGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.