Interestana
Home/News/Microsoft Exec Calls AI Training 'Largest Theft of Labor'
Fast Company3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Microsoft Exec Calls AI Training 'Largest Theft of Labor'

Microsoft Exec Calls AI Training 'Largest Theft of Labor'

Newly unsealed court documents have revealed internal concerns from a Microsoft executive regarding the training of artificial intelligence models on copyrighted human-created content. In a legal brief unsealed this week, Brent Hecht, Microsoft's Director of Applied Science, described OpenAI's practice of training its models on news articles and other published works as "the largest theft of labor in human history." Hecht further characterized this process of using copyrighted materials to train AI models as "an astonishing theft of unprecedented proportions." These statements emerged within the context of a copyright lawsuit filed in 2023 by The New York Times against Microsoft and OpenAI. The lawsuit alleges that the companies infringed on copyright law by using millions of The Times's news articles, investigations, opinion pieces, and reviews to train their AI products, specifically naming Copilot and ChatGPT. The New York Times contends that these AI models are designed to be "largely substitutive," meaning they aim to replace the original source material, thereby undermining the publishers' business models. The brief also disclosed that OpenAI's head of ChatGPT acknowledged that publishers like The New York Times "face an existential threat" from AI products. Hecht's assessment extended to describing large language models as "a product that destroys its supply chain," highlighting a potential self-defeating cycle where the AI's development depletes the very resources it relies upon. The New York Times's lawsuit further accuses Microsoft and OpenAI of seeking a "free ride," leveraging the newspaper's significant financial investment in producing quality journalism to develop substitutive products without authorization or compensation. The publisher specifically noted that while the defendants copied from numerous sources, they placed particular emphasis on content from The Times when building their large language models, indicating an acknowledgment of the value of these works. This legal action underscores a growing tension between AI development and intellectual property rights, as AI companies increasingly rely on vast datasets of existing content for training, raising complex questions about fair use, copyright infringement, and the future of content creation and distribution in the digital age. The unsealed documents provide a rare glimpse into the internal discussions and strategic considerations within leading AI organizations and their partners regarding the ethical and legal implications of their data acquisition and model training methodologies.

Original source — read the full reporting at the publisher:

Read on Fast Company

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next