Interestana
Home/News/AI Companies Accused of Destroying Books
The Atlantic3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

AI Companies Accused of Destroying Books

AI Companies Accused of Destroying Books

Concerns are mounting that artificial intelligence companies are systematically using copyrighted books to train their large language models (LLMs) without obtaining appropriate licenses or compensating authors and publishers. This practice has ignited a debate about copyright infringement, fair use, and the economic viability of creative professions in the age of AI. Critics argue that the vast datasets used to train models like OpenAI's GPT series and Google's Bard often include millions of books, many of which are still under copyright. The process involves scraping text from digital sources, and it is alleged that these companies are not distinguishing between publicly available works and those protected by copyright law. This has led to accusations that AI developers are essentially benefiting from the intellectual property of authors and publishers without contributing to their livelihood.

Several lawsuits have already been filed by authors and publishing groups against prominent AI companies. For instance, authors like Paul Tremblay and Mona Awad, along with the Authors Guild, have initiated legal actions. These lawsuits contend that the unauthorized use of copyrighted material constitutes a violation of intellectual property rights. The core of the legal argument often revolves around whether the training of AI models on copyrighted text falls under the doctrine of fair use, a legal principle that permits limited use of copyrighted material without permission for purposes such as criticism, comment, news reporting, teaching, scholarship, or research. AI companies, however, maintain that their use of data for training is transformative and falls within fair use parameters, enabling the development of innovative technologies.

The economic implications for authors and the publishing industry are significant. If AI models can generate content that directly competes with human-created works, or if the value of intellectual property is diminished by its free use in AI training, it could undermine the ability of writers to earn a living. Publishers also face the prospect of decreased sales if AI-generated content becomes a widespread substitute for books. The debate extends beyond legal interpretations to ethical considerations regarding the value of human creativity and the sustainability of artistic careers. Organizations like the Authors Guild are advocating for clearer regulations and licensing frameworks that ensure authors are fairly compensated for the use of their work in AI development.

Further complicating the issue is the opacity surrounding the exact datasets used by many AI companies. While some companies have disclosed general information about their data sources, the specific inclusion of copyrighted books and the extent of their use often remain unclear. This lack of transparency makes it difficult for rights holders to monitor and enforce their copyrights effectively. As AI technology continues to advance rapidly, the legal and ethical frameworks governing its development and data usage are struggling to keep pace. The ongoing legal battles and public discourse highlight the urgent need for a resolution that balances technological innovation with the protection of intellectual property and the livelihoods of creators.

Original source — read the full reporting at the publisher:

Read on The Atlantic

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next