Interestana
Home/News/Microsoft Exec Called AI Scraping "Largest Theft of Labor"
Ars Technica3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Microsoft Exec Called AI Scraping "Largest Theft of Labor"

Microsoft Exec Called AI Scraping "Largest Theft of Labor"

Internal documents unsealed on Thursday in a lawsuit brought by news organizations, including The New York Times, reveal that a Microsoft executive described the practice of scraping news content for AI training as "an astonishing theft of unprecedented proportions" and potentially "the largest theft of labor in human history." These documents, part of a motion for summary judgment, were intended to expose how Microsoft and OpenAI perceived the risks to news organizations before launching AI products like ChatGPT and Copilot. Brent Hecht, identified as a Director of Applied Science at Microsoft, repeatedly voiced these concerns in internal communications. Hecht's statements directly challenged the companies' legal arguments that training AI on news content constitutes fair use. In one document, he reportedly stated that the plan to extensively scrape news "made a complete mockery of the idea of ‘fair use.’" The news plaintiffs allege that these internal communications demonstrate a clear understanding by Microsoft and OpenAI of the copyright and ethical implications of their data acquisition methods. For years, Microsoft and OpenAI have sought to keep such internal details confidential amidst ongoing legal battles with news publishers who accuse the AI firms of copyright infringement by using vast amounts of news content to train their models. The unsealing of these documents provides significant new evidence for the plaintiffs in their ongoing fight against the AI companies. The lawsuit centers on allegations that the AI firms systematically violated copyright laws by acquiring and utilizing copyrighted news articles without proper authorization or compensation to train their artificial intelligence systems. The core of the dispute lies in whether the AI companies' use of copyrighted material for training constitutes fair use under copyright law, a defense that Hecht's internal statements appear to undermine. The plaintiffs argue that the AI models developed by Microsoft and OpenAI are directly competing with news organizations, leveraging the very content that was allegedly stolen. This legal confrontation highlights a broader industry-wide debate regarding the ethical and legal boundaries of AI development, particularly concerning the use of publicly available but copyrighted data. The outcome of this case could set significant precedents for how AI companies source training data and how news organizations are compensated for their intellectual property in the age of generative AI. The unsealed documents suggest a level of internal awareness within Microsoft regarding the problematic nature of the AI training data practices, contradicting public statements or legal defenses presented by the companies.

Original source — read the full reporting at the publisher:

Read on Ars Technica

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next