Interestana
Home/News/AI Developers Destroy Millions of Books for Training Data
Decrypt2 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

AI Developers Destroy Millions of Books for Training Data

AI Developers Destroy Millions of Books for Training Data

Artificial intelligence developers are acquiring millions of physical books, processing them into digital training data for large language models, and subsequently discarding the original copies. This practice, described as "AI book burning," involves the physical destruction of books to extract their content for AI training datasets. The process entails slicing apart books, scanning each page, and then disposing of the originals, effectively removing them from circulation and potential future access.

This method of data acquisition is being employed by companies seeking to enhance the capabilities of their AI models, particularly large language models (LLMs) that require vast amounts of text to learn and generate human-like responses. The sheer volume of material needed for training these sophisticated AI systems has led developers to explore diverse data sources, including copyrighted works. The destruction of physical books represents a significant, albeit controversial, approach to fulfilling these data requirements. The scale of this operation is substantial, with "millions of books" being processed in this manner.

The implications of this practice extend beyond mere data collection. Critics and observers are raising concerns about the potential loss of cultural heritage and the long-term accessibility of literature. When physical books are destroyed, they are no longer available for public access through libraries, archives, or for future scholarly research. This raises questions about the sustainability of such data acquisition methods and the ethical considerations surrounding the preservation of intellectual and cultural works. The practice highlights a tension between the rapid advancement of AI technology and the imperative to safeguard existing cultural artifacts.

While the specific companies engaged in this practice are not explicitly named in the provided information, the trend is described as being undertaken by "AI developers." The motivation behind this approach is the need for extensive and diverse textual data to train LLMs. The process itself is described as a form of "book burning," a term that evokes historical instances of censorship and destruction of knowledge, albeit for different technological ends. The ultimate fate of the scanned data and the original books underscores the irreversible nature of this data acquisition strategy.

Original source — read the full reporting at the publisher:

Read on Decrypt

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next