Newly unsealed court filings show Microsoft privately called OpenAI's data practices "theft" while both companies scraped paywalled Times content, built datasets from it, and warned internally it would gut publishers.
By Nexvoro Tech Wire
PUBLISHED THU, SEP 17, 2026 8:00 PM UTC • 6 MIN READ
Primary Journalistic Dispatch & Direct Reporting
Disrupt 2026: OpenAI, Anthropic, Replit, and more take over 6 industry stages. 25% off tickets now
Per the lawsuit, a top Microsoft executive privately described the companies' AI training practices as "theft," and OpenAI's own leadership said its AI models posed an "existential threat" to the publishers and journalists whose work trained them.
It's worth noting that much of the new information comes from The Times' own brief, not the underlying exhibits, which remain sealed. The quotes below are presented without their original context.
In-Depth Developments & Factual Context
Several of the new admissions, however, run counter to OpenAI's fair use defense, particularly the rule's requirement that use doesn't substitute or harm the market for the original work.
For example, Microsoft's own data shows its Copilot "answer engine" caused click-through rates for The New York Times' domain to drop as much as 93% compared to traditional Bing search. An internal Microsoft presentation written by Microsoft's Director of Applied Science, Brent Hecht, in January 2024 describes the decline as a "doom loop" that would "hurt the performance of our models and the entire web at the same time."
"It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain,'" reads the Microsoft document, as quoted in the filing.
Industry Impact & Strategic Analysis
Microsoft CEO Satya Nadella also testified in a deposition earlier this year that "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training," and made clear that, if he "had been made aware that OpenAI had scraped and trained on information that was behind a paywall," he would have "invoked [Microsoft's right to] require OpenAI to retrain its models."
Other admissions cut against different pillars of the fair-use test: OpenAI's Head of ChatGPT, Nick Turley, wrote in internal communication that publishers face an "existential threat" from products like the chatbot, which are "largely substitutive" and "will get more and more substitutive as they get better."
OpenAI President Greg Brockman described the models as "excellent at news." Nadella agreed under oath earlier this year that conversing with chatbots "has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source."
Forward Outlook & Market Perspective
That kind of language speaks to how the technology could directly compete with, rather than transform, the original work.
A Microsoft document states that there is a "real risk" that generative AI could "significantly disrupt the employment of the very people who generated the data on which the foundation model was trained."
The sheer scale of the copying is striking. The documents reveal for the first time that OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting. A Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone.
In a January 2023 internal memo, Hecht called it "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history."
The filing lays out in new detail how OpenAI and Microsoft went about acquiring the plaintiffs' content, including scraping it from the Bing Index.
"OpenAI delivered the entire GPT-3 training dataset to Microsoft, which Microsoft used to evaluate how to implement OpenAI's models within its own commercial products," the filing reads. "Microsoft similarly provided training data to OpenAI through initiatives called Project Taxi and Project Mango."
Reporting synthesized and verified under Nexvoro.tech editorial guidelines. Full primary records referenced via TechCrunch.
Reporting synthesized under Nexvoro.tech Editorial Standards • Referenced via TechCrunch
Verified Dispatch