← All posts / Policy

‘The Largest Theft of Labor in Human History’: Unsealed Filings Catch Microsoft and OpenAI Speaking Out of Court

Newly unredacted filings in the NYT v. OpenAI/Microsoft copyright case reveal a Microsoft exec privately called AI scraping ‘the largest theft of labor in human history,' Nadella testified paywalled content should be licensed, and OpenAI's own leadership admitted chatbots pose an ‘existential threat' to publishers.

‘The Largest Theft of Labor in Human History’: Unsealed Filings Catch Microsoft and OpenAI Speaking Out of Court

On September 17, 2026, the federal courthouse in Manhattan quietly released one of the most damaging sets of documents in the three-year history of The New York Times v. Microsoft and OpenAI. Judge Sidney Stein unsealed the latest round of filings in the landmark copyright case, and the unredacted text shows something no polished press release ever would: the defendants’ own executives, in their own words, describing the economics of AI training in terms that read like the plaintiffs’ opening statement.

The headline quote belongs to Brent Hecht, Microsoft’s Director of Applied Science. In a January 2023 internal memo, Hecht described large-scale scraping of news content as “an astonishing theft of unprecedented proportions” — and, in the phrase now travelling around the world, “the largest theft of labor in human history.” Three years later, that sentence is headline exhibit material in a case where Microsoft’s official position is that training on copyrighted text is transformative fair use.

What the unredacted filings actually say

The newly public material comes primarily from The Times’ own brief rather than the underlying exhibits, which remain sealed — a caveat worth keeping in mind, since the quotes are presented without their original context. But even on the plaintiffs’ characterization, the admissions are remarkable.

The “doom loop” presentation. An internal Microsoft presentation written by Hecht in January 2024 contains data showing that Microsoft’s Copilot “answer engine” caused click-through rates for the nytimes.com domain to drop by as much as 93 percent compared with traditional Bing search. Hecht described a “doom loop” that would “hurt the performance of our models and the entire web at the same time” — a strikingly explicit acknowledgment that answer engines starve the very sources their models depend on. Another line from the same document is almost academic in its candor: “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”

Nadella under oath. Microsoft CEO Satya Nadella, deposed earlier this year, testified that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training,” and stated that if he “had been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.” He also agreed that chatting with a bot “has substituted…giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”

OpenAI’s internal language. Nick Turley, OpenAI’s head of ChatGPT, wrote in internal communications that publishers face an “existential threat” from products like the chatbot, which are “largely substitutive” and “will get more and more substitutive as they get better.” President Greg Brockman described the models as “excellent at news.” A Microsoft document separately conceded a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”

The scale, and the paywall hack

The filings also quantify, for the first time, how much publisher content sat inside the training pipeline. OpenAI’s mid-training datasets alone allegedly contain more than 91,692 copies of works published by the NYT, the New York Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset included more than two million documents from nytimes.com alone. And the data flowed both ways: OpenAI delivered the entire GPT-3 training dataset to Microsoft for product integration, while Microsoft supplied training data back through initiatives code-named “Project Taxi” and “Project Mango” — the latter assembled into a dataset containing copies of at least 160,903 unique works from the news publishers.

Most damning of all is the paywall episode. When OpenAI researcher Nick Ryder told Brockman he had found a “hack to get around nytimes paywall,” Brockman’s reply, as quoted in the filing, was two words: “ah nice.” The filings further allege that employees deliberately stripped copyright notices from training data before it reached the models, because researchers “wouldn’t want model outputting” those notices to users.

Why it matters for fair use

Fair use in the United States turns on four factors, and the fourth — whether the use harms the market for the original work — has always been the publishers’ strongest card. Internal admissions that products are “largely substitutive,” paired with Microsoft’s own 93 percent click-through data, speak directly to that factor in a way outside experts never could. They also complicate the defense’s argument that AI output transforms rather than replaces journalism.

The timing is delicate. Earlier this month, the Trump administration filed a brief defending OpenAI’s unlicensed use of copyrighted material, and courts so far have been largely sympathetic to fair-use arguments in AI training cases. But as Steven Lieberman, counsel for the New York Daily News, put it in a statement: “The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong.” Neither company returned requests for comment on the unsealed filings.

The gap between public positioning and private language is the real story. Microsoft markets Copilot as a partner to publishers; internally, its own scientist was writing memos about theft and doom loops in the same period. Whatever Judge Stein ultimately rules on summary judgment, the unsealed record has already changed the public narrative — and given every licensing negotiation in the industry a new, uncomfortable starting point.