By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Microsoft Called AI Scraping 'Theft' in Unsealed Filings
Newly unsealed court filings have revealed that Microsoft executives privately described the practice of artificial intelligence companies scraping data as "the largest theft of labor in human history." These filings emerged in the context of ongoing legal disputes involving The New York Times and OpenAI, where allegations of unauthorized data usage for training AI models have been central. The documents indicate that Microsoft, despite its close partnership with OpenAI, acknowledged the severe implications of this data acquisition strategy on content creators and publishers.
Specifically, the unredacted filings show that both Microsoft and OpenAI engaged in scraping paywalled content from The New York Times. This data was then utilized to build datasets for training their respective AI models. Internally, Microsoft executives expressed concerns that such practices would significantly undermine the business models of publishers. The company's internal communications, now brought to light, suggest a recognition of the ethical and economic ramifications of using copyrighted material without explicit permission or compensation for AI development. This revelation adds a layer of complexity to the ongoing debate about fair use, intellectual property, and the future of content creation in the age of generative AI.
The legal actions, including the one initiated by The New York Times against OpenAI and Microsoft, focus on the alleged infringement of copyright. The Times claims that its articles were used to train AI systems, including OpenAI's ChatGPT and Microsoft's Copilot, without authorization. The unsealed documents provide further evidence of the internal discussions and acknowledgments within Microsoft regarding the nature of this data collection. The company's private admissions of the practice being "theft" and its potential to "gut publishers" underscore the significant challenges faced by the media industry as AI technologies continue to advance.
These internal statements from Microsoft highlight a potential conflict between the company's public stance on AI development and its private assessments of the methods employed. The legal ramifications of these revelations could be substantial, potentially influencing ongoing litigation and shaping future regulations around AI data sourcing. The broader implications extend to the entire AI industry, prompting a re-evaluation of data acquisition strategies and the need for more transparent and ethical practices. The unsealed filings serve as a critical piece of evidence in the escalating legal and ethical battles over AI training data.
Original source — read the full reporting at the publisher:
Read on TechCrunchGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.