Unsealed court documents from a lawsuit led by The New York Times reveal that executives at Microsoft and OpenAI were internally aware of the profound impact AI scraping would have on news publishers.
PLUS ULTRALawsuitsMicrosoftOpenAI
Unsealed Documents Reveal Microsoft and OpenAI Executives Acknowledged AI Scraping's "Theft" and Existential Threat to News
PLUS ULTRA by Amenoyomi
Microsoft Director of Applied Science Brent Hecht reportedly described news scraping for AI training in a January 2023 internal memo as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." Hecht also suggested that widespread scraping made the concept of "fair use" a "complete mockery."
At OpenAI, ChatGPT head Nick Turley described in an internal message that publishers would face an "existential threat" from commercial products trained on news content, noting that these products are "largely substitutive."
Microsoft's own data highlighted the practical consequences of this substitution. For certain news plaintiffs, Microsoft recorded click-through rate drops of 83% to 93% due to its Copilot "answer engine." Internal documents referred to this phenomenon as a "doom loop" that threatens to hurt both model performance and the entire web by undermining the economic foundations of the content supply chain.
The filings also detail how the companies allegedly obtained content, including using methods to bypass paywalls. Documents show that when an OpenAI researcher informed President Greg Brockman about a "hack" to circumvent The New York Times' paywall, Brockman responded, "ah nice." Furthermore, documents suggest efforts were made to strip copyright notices from training data to prevent them from appearing in model outputs.
While Microsoft defended its AI products as transformative fair use that does not substitute for news sites, internal documents acknowledged that "almost no one intended for content they created to be used in this fashion, nor are they compensated for its use."
PLUS ULTRAby Amenoyomi
The conflict focuses on the legal interpretation of "fair use," specifically whether AI training transforms content or serves as a market substitute. While Microsoft publicly maintains that its AI products are a transformative fair use, internal documents from Director of Applied Science Brent Hecht describe the practice as a "complete mockery" of the fair use concept and the "largest theft of labor in human history."
Internal data highlights the practical impact of this substitution. Microsoft recorded click-through rate drops between 83% and 93% for some news publishers, and between 51% and 94% for others, when users utilized Copilot's answer engine instead of traditional Bing search. This is described internally as a "doom loop," a systemic risk where the end-product undermines the economic foundations of its essential "content supply chain," potentially harming both the performance of future models and the broader web.
The documents also detail specific methods used to acquire training data. OpenAI employees reportedly developed "hacks" to bypass paywalls, an action President Greg Brockman acknowledged with the phrase "ah nice." Furthermore, the firms allegedly stripped copyright notices from training datasets to prevent the models from outputting such notices to users. The scale of this collection is substantial; mid-training datasets contained over 91,692 copies of works from the New York Times, Daily News, and Center for Investigative Reporting, while another dataset included more than 2 million documents from nytimes.com alone.
In response to these revelations, a Microsoft spokesperson stated that Brent Hecht's comments reflect an "individual perspective" and do not represent the company's legal views. CEO Satya Nadella testified that content behind paywalls should be licensed for grounding or training and acknowledged that chatbots have acted as substitutes by providing information directly on the AI platform rather than directing users to the original source.
Sources
- Microsoft exec called AI scraping the “largest theft of labor in human history” (Ars Technica AI, 2026-09-17)
- Microsoft幹部がAIによるデータスクレイピングを「人類史上最大の労働窃盗」と表現 (GIGAZINE, 2026-09-18)