By Interestana AI Editorial — AI-drafted, human-overseen. How we report
AI Firms Acknowledged Content Theft Concerns Internally

Internal documents from OpenAI and Microsoft, made public on September 17 as part of The New York Times' lawsuit, indicate that executives at both companies were aware of the controversial nature of using copyrighted content to train large language models. Microsoft's Brent Hecht, a director of applied science, warned in a memo in early 2023 that "millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft," and characterized the practice as "the largest theft of labor in human history." This statement highlights a significant internal acknowledgment of the ethical and legal challenges surrounding the sourcing of training data for generative AI.
The documents also reveal direct communications between OpenAI personnel regarding the circumvention of publisher paywalls. In one instance, when a researcher described successfully bypassing The New York Times' paywall, OpenAI President Greg Brockman responded with "ah nice." Furthermore, Nick Turley, OpenAI's head of ChatGPT, reportedly referred to chatbots as an "existential threat" to publishers, suggesting an awareness within the company of the disruptive impact their technology could have on the media industry. These internal communications underscore a sophisticated understanding of the value of content and the potential negative consequences of its unauthorized use for AI development.
The revelations come at a critical juncture in the ongoing legal battle between The New York Times and OpenAI, along with its partner Microsoft. The lawsuit centers on allegations that OpenAI and Microsoft unlawfully used millions of copyrighted articles to train their AI models, including ChatGPT. The internal memos and emails now provide evidence that key figures within these AI companies understood the implications of their data sourcing methods, potentially contradicting earlier public stances or perceptions of their awareness. This internal acknowledgment of 'theft' and 'existential threat' could significantly influence the legal proceedings and the broader discourse on AI ethics and copyright.
These internal discussions suggest that while the public narrative around AI development often focused on technological innovation, there was a concurrent, albeit sometimes private, recognition of the value of the raw material being used. The content scraped from the internet, including that behind paywalls, is now understood to have inherent value that varies based on its uniqueness and the intended use by AI customers. The consensus, as indicated by these internal exchanges, is that this value is not zero, a fundamental point that publishers have been advocating for in their legal challenges and public appeals. The legal case is expected to further explore how these internal acknowledgments will factor into the determination of liability and fair compensation for content creators.
Original source — read the full reporting at the publisher:
Read on Fast CompanyGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.