By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Amazon Destroys Rare Books for AI Training
Amazon is reportedly engaging in the destruction of rare and valuable books to extract their content for training artificial intelligence models. This practice, revealed through internal documents and corroborated by sources familiar with the matter, involves digitizing the text of these books and then discarding the physical copies. The primary motivation behind this initiative is to feed AI models with a broader and more diverse dataset than what is readily available on the public internet. Large language models (LLMs) have historically been trained on vast amounts of text scraped from the web. However, as the internet's readily accessible content becomes saturated, AI developers are seeking out new, unique, and often out-of-print materials to enhance model capabilities and prevent them from becoming overly reliant on common information. Rare books, with their unique language, historical context, and specialized knowledge, represent a valuable source for this purpose. The process reportedly involves scanning or otherwise digitizing the pages of these books, converting the text into a format that AI models can process. Once the content is extracted, the physical books are often pulped or otherwise destroyed, rather than being preserved or returned to circulation. This approach has drawn criticism from book collectors, librarians, and cultural heritage advocates who view the destruction of these artifacts as a significant loss. They argue that rare books hold intrinsic historical, cultural, and artistic value that extends beyond their textual content. The act of destroying them for data extraction is seen as a disregard for their heritage and a permanent erasure of unique physical objects. Preservationists emphasize that while digital access is important, it should not come at the cost of irrevocably destroying original materials. Concerns have also been raised about the ethical implications of such a practice, particularly regarding the potential for devaluing physical media and the cultural heritage it represents. The long-term consequences of this method for AI training are also a subject of debate, with some questioning whether the unique qualities of rare books can be fully captured and utilized by AI without the context and materiality of the original object. Amazon has not officially commented on this specific practice, but the company has been investing heavily in AI development and has access to extensive libraries and archives through its various ventures, including its acquisition of MGM which includes a vast film and television library that could also serve as a data source. The practice highlights a growing tension between the insatiable demand for data in the AI industry and the preservation of cultural heritage.
Original source — read the full reporting at the publisher:
Read on TechCrunchGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.