← All posts / Policy

24 Matches in 8.2 Million Chats: Microsoft Bets the Entire AI Copyright Case on One Number

Microsoft's summary judgment filing says an expert found just 24 Copilot responses matching authors' books across 8.2 million logs — a 0.00029% rate it calls proof that LLM training is transformative fair use.

24 Matches in 8.2 Million Chats: Microsoft Bets the Entire AI Copyright Case on One Number

On September 4, 2026, Microsoft walked into the federal courthouse in Manhattan and laid a single number on Judge Sidney H. Stein’s desk: 24. Out of 8.2 million Copilot conversations surrendered in discovery, only twenty-four responses contained at least thirty consecutive words matching the books that authors are suing over. That is a reproduction rate of 0.00029 percent, and Microsoft wants it to end the largest copyright battle in the history of artificial intelligence.

The filing is a motion for summary judgment in In re: OpenAI, Inc. Copyright Infringement Litigation — the multidistrict proceeding before Judge Stein in the Southern District of New York that consolidates claims from the Authors Guild, named fiction and nonfiction authors, The New York Times Company, and the Center for Investigative Reporting. Filed alongside a parallel motion targeting the News plaintiffs’ claims and a judgment-on-the-pleadings bid on contributory infringement, it is Microsoft’s most aggressive legal move since the case began, and arguably the most consequential single filing to date in the broader war between AI labs and content owners.

What the log analysis actually found

The heart of the memorandum is a forensic audit of Copilot’s real-world behavior, conducted by the plaintiffs’ own expert, Dr. Shawn Shan. Two datasets matter.

The adversarial lab test. Shan ran approximately 5.3 million deliberate extraction attempts — feeding the GPT models verbatim book passages hundreds of words long and repeatedly prompting until the model emitted matching text. Even under that adversarial protocol, fewer than one percent of attempts produced any 30-word match. Microsoft characterizes this as a best-case scenario for the plaintiffs: a researcher who already possesses the book, targets a known passage, and tries over and over to force a regurgitation, mostly fails.

The production logs. Outside the laboratory, Shan reviewed 8.2 million real Copilot conversations and found the 24 matching responses. For 202 of the 212 asserted books, there was no regurgitation in the logs at all. On the news side, the numbers run slightly higher but tell the same story: 59,545 conversations — under one percent — shared at least sixteen words with news content used to ground the model, and the Center for Investigative Reporting’s expert identified just 51 instances of substantial overlap with CIR journalism.

There is a critical caveat, and Microsoft discloses it rather than hiding it: the 8.2 million logs were not a random sample. They were pre-filtered on keywords tied to the plaintiffs’ websites, meaning the dataset was deliberately stacked toward the conversations most likely to contain infringing material. Even inside that rigged sample, matches were vanishingly rare. Microsoft’s lawyers clearly calculated that volunteering the selection bias makes the headline number stronger, not weaker.

The fair use argument

The legal core of the brief is that training a large language model on copyrighted books is fair use as a matter of law. Microsoft calls the use “profoundly transformative,” leaning on two prior AI training rulings — one against Meta, one against Anthropic — where courts described LLM training in similar terms. Books are written to be read by humans; OpenAI’s use of texts like John Grisham’s The Street Lawyer and Stacy Schiff’s Cleopatra served, in the filing’s framing, a technological purpose: building a system that generates natural-language responses to prompts.

On the messiest factual issue — OpenAI’s downloading of books from Library Genesis, a shadow repository of unauthorized copies — Microsoft draws a hard line. The undisputed record, it says, establishes that Microsoft did not download from LibGen and had no involvement in acquiring those datasets, and it argues that provenance of training data does not change the fair use analysis because every step served the same training purpose. It also addresses Microsoft’s transfer of portions of the Bing search index at OpenAI’s request, conceding the evidence is inconclusive but noting the plaintiffs’ expert identified only 8 to 24 asserted works in the transferred portions.

On market harm, the traditional fourth fair use factor, Microsoft goes on offense. It cites the authors’ own sales evidence showing they are selling just as many books as they would have had ChatGPT and Copilot never existed, quotes playwright David Henry Hwang saying he does not believe the defendants’ LLMs are hurting the market for his plays, and points to research finding that the overwhelming majority of consumers would not buy an AI-generated book even at a drastically lower price. A licensing requirement for training data, the brief warns, would force developers to clear rights with millions of authors — a compliance wall that only the largest companies could afford to climb.

The plaintiffs’ rebuttal, in advance

The Times’ lead counsel, Ian Crosby, has already rejected the framing: discovery, he says, established that Microsoft and OpenAI “built commercial products designed to substitute for Times journalism and compete directly with its business.” The plaintiffs’ theory has always been that raw reproduction counts understate the harm — that partial substitution at scale siphons readers, referrals, and advertising revenue from the outlets whose reporting fed the training data, whether or not full articles ever appear verbatim.

The counter-arguments write themselves. Sixteen matching words is a low bar for overlap but a high bar for proving substitution — and a keyword-filtered discovery sample, plaintiffs will argue, says nothing about reproduction rates among the billions of Copilot queries that never touched plaintiff-related keywords. A chatbot that summarizes a Times investigation in different words may still displace the click to nytimes.com. That is a market-harm argument, not a regurgitation argument, and it is precisely the theory Microsoft is asking the judge to bury.

There is also the political backdrop. Days earlier, the Trump administration’s Justice Department filed a statement of interest urging the judge to rule for OpenAI and Microsoft, arguing that training AI on the newspaper’s content does not violate copyright law. Federal weight now sits behind the fair-use position.

Why this filing matters beyond one case

Whatever Judge Stein decides, the filing’s structure — a discovery dataset, a matching-word analysis, and a fair use argument anchored to concrete numbers rather than abstract claims about model behavior — becomes a template. Anthropic settled with the Authors Guild in a separate action earlier this year. OpenAI faces parallel suits from multiple newspaper chains. A grant of summary judgment here would not bind those cases, but it would hand every AI defendant a discovery-tested playbook: measure the outputs, show the rate approaches zero, and argue transformative purpose.

A denial cuts the other way. Discovery continues, and the substitution theory gets tested before a jury — a far riskier posture for AI defendants, given the sums involved and the emotional resonance of named journalists and authors describing what happened to their work. Responses to the motions are due October 5, 2026, with opposition briefs publicly refiled by October 15 under the sealing calendar Judge Stein signed on September 3.

Every lab shipping a retrieval-grounded chatbot is watching. The ruling will set the operating cost of grounding models on the open web for years — and for now, the entire defense rests on a number small enough to say out loud: twenty-four.