By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Microsoft, OpenAI Execs Acknowledged AI Training Theft Concerns
Internal communications from Microsoft and OpenAI executives, unredacted in a court filing for the New York Times v. OpenAI/Microsoft case, reveal significant concerns about the ethical and legal implications of training large language models (LLMs) on vast amounts of internet content. These documents suggest that individuals within these companies acknowledged that the practice could be perceived as a substantial form of theft and a threat to the creators of the training data. One such communication, from Brent Hecht, Microsoft's director of applied science, stated, "Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions." Hecht further elaborated that "Almost no one intended for content they created to be used in this fashion, nor are they compensated for its use." A Microsoft spokesperson responded to these unearthed comments by stating they "reflect one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views," and reiterated Microsoft's legal stance that "transformative uses are consistent with copyright law and why Copilot is not a substitute for publishers’ journalism."
Beyond the "astonishing theft" sentiment, other internal documents highlighted the potential for generative AI to destabilize the very industries that produce the data it consumes. One Microsoft internal communication identified the "real risk" that generative AI could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained." A separate policy document articulated this concern more starkly, stating, "LLMs are a product that destroys its supply chain." These internal acknowledgments stand in contrast to the public legal arguments made by AI companies, which generally assert that their training practices fall under the legal doctrine of fair use. The publishers, led by The New York Times, argue that this uncompensated use of their content undermines the news industry, which is essential for generating the data that trains these AI models. The unredacted filing provides a glimpse into internal debates and potential awareness of the broader societal and economic impacts of AI development.
The New York Times initiated legal action against OpenAI and Microsoft, alleging that the companies unlawfully used copyrighted material from the newspaper to train their AI models. This lawsuit is part of a broader trend of legal challenges from content creators and publishers who are seeking to address the use of their work in AI training without consent or compensation. The unredacted portions of the court filing offer a more detailed perspective on the internal discussions and potential reservations held by some within the AI industry regarding the methods of data acquisition and the potential consequences for content creators. The case is ongoing and could set significant legal precedents for the future of AI development and copyright law. The documents suggest a complex internal landscape where the rapid advancement of AI technology was being weighed against its potential to disrupt established industries and creator livelihoods.
Original source — read the full reporting at the publisher:
Read on Nieman LabGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.