← All posts / Research

"That Did Not Happen": TU Dresden's Andreas Thom Accuses OpenAI of a Misleading Answer on His Private ChatGPT Chats

The group theorist whose 2019 paper underpins Astra's non-sofic group proof says OpenAI's Mark Sellke answered only half of his question about whether his private ChatGPT conversations entered the training pipeline — the third public misconduct allegation against OpenAI's math program in a week.

"That Did Not Happen": TU Dresden's Andreas Thom Accuses OpenAI of a Misleading Answer on His Private ChatGPT Chats

A new front has opened in the escalating credit war between working mathematicians and OpenAI — and this one reaches back months before the Navier–Stokes controversy that has dominated headlines all week. Andreas Thom, professor of geometry at TU Dresden and one of the two authors behind the technique that anchors OpenAI’s most celebrated mathematical result, has gone public with an accusation that a senior OpenAI researcher gave him a misleading answer about whether his private conversations with ChatGPT ever entered the company’s training pipeline.

The claim, laid out in a three-part post on Mathstodon on September 10, 2026, lands as the third public misconduct allegation against OpenAI’s mathematics program in a single week — and unlike the previous two, it is made by the researcher whose own published work sits underneath the result in question.

The result at stake

In August 2026, OpenAI announced that its Astra model had constructed the first explicit non-sofic group, resolving a 27-year-old open problem in geometric group theory first posed by Mikhail Gromov in 1999. The construction was the headline item in a batch of ten AI-generated results that OpenAI published with Lean-formalized proofs, and it was pitched as evidence that frontier models could produce genuinely new mathematics.

There was a catch that emerged within days: the proof’s central technical step leaned directly on a 2019 paper by Gábor Kun and Andreas Thom, along with a related 2016 result by Kun. Cambridge group theorist Francesco Fournier-Facio and Yeshiva University’s Steven Miller separately argued that two of Astra’s ten flagship results — including the non-sofic construction — depended on existing papers far more heavily than OpenAI’s announcement acknowledged. Miller told Scientific American the pattern of unattributed prior work “points to research misconduct.” OpenAI subsequently edited its blog post to soften language claiming the problems had seen “no progress… for at least a decade.”

But the attribution question, it turns out, was only half of Thom’s concern.

The email and the answer

Shortly after Astra’s announcement, Thom says he emailed OpenAI’s Mark Sellke and Sébastien Bubeck — the researcher who leads the company’s mathematics efforts — with a disclosure of his own: he and a colleague in Dresden had spent months discussing the expander matching problem and extensions of the Kun–Thom work with ChatGPT itself. Before the proof existed, the ideas that would underpin it had passed through the very product OpenAI now wanted to celebrate.

Thom’s email asked two separate questions: whether those conversations had entered the model’s training data, and separately, whether they had been accessible to the system while it was working on the proof.

Sellke’s reply, as Thom quotes it, was six words long: “Regarding your conversations with ChatGPT: that did not happen.”

Thom’s argument is that this answer addresses only the second question — direct access during the solving process — while staying silent on the first and more consequential one: whether his conversations shaped the training data itself. He now calls the response materially misleading rather than honest, and his reasoning is difficult to dismiss as pedantry. A flat denial that sounds categorical can, on careful reading, deny only the narrowest interpretation of the question — a pattern of answer that gives the recipient exactly the reassurance they appear to be seeking while leaving the substance untouched.

Why the Navier–Stokes fallout changed everything

What transforms Thom’s months-old exchange from a private grievance into a public accusation is what OpenAI said — and did not say — in the separate dispute that erupted this week.

In the Buckmaster–Alpöge affair, NYU’s Tristan Buckmaster and Anthropic researcher Levent Alpöge accused OpenAI of learning about their unpublished Navier–Stokes work and racing to publish a competing proof, with Bubeck allegedly pushing to strip Alpöge’s name from a joint paper because he works at a rival lab. OpenAI’s eventual written response said no researcher or agent saw the pair’s work before publication, and that no specific user data was accessed.

But the company added a carve-out it had never offered Thom: it “cannot rule out that de-identified data derived from their usage of our products helped improve” its models.

That sentence, Thom argues, is precisely the distinction his original email drew — and precisely the distinction Sellke’s flat denial erased. OpenAI’s own language in a higher-stakes dispute implies that “that did not happen” was never a full answer to the training-data question. If de-identified usage data can help improve models, then a denial of access during the solve says nothing about whether private conversations fed the pipeline that trained the system in the first place.

What Thom is actually asking for

Notably, Thom is not asking anyone to reverse-engineer OpenAI’s training pipeline. His demand is narrower and harder to refuse: a categorical denial needs to come with a disclosed basis. He has called on OpenAI to specify what its account settings say, what training checkpoints show, and — most pointedly — to provide a plain-English definition of what “de-identified data derived from usage” actually covers.

He also notes a structural problem that every ChatGPT user shares: opting out of training is a forward-looking control with no audit mechanism. Thom says he opted out on June 29 — a date he can verify, but a setting that tells him nothing about what happened to his conversations before that date. If the only entity that holds the relevant data is the same entity being asked to deny using it, the denial’s evidentiary value depends entirely on the denier’s willingness to show its work.

The mathematical argument

Thom’s suspicion is sharpened, though not proven, by the mathematics itself. He observes that the Kun–Thom approach was never considered the leading candidate for resolving non-soficity — other lines of attack involving quantum games looked more promising to specialists. That is part of why he found it notable that Astra homed in on his specific technique rather than the field’s consensus direction.

He also makes a point that cuts deeper than the specifics of this case: de-identifying a conversation strips out a name, not the underlying mathematical idea. Privacy safeguards engineered for personal data — names, addresses, identifying patterns — do not obviously do any work when the “data” in question is unpublished research reasoning. A scraped chat about the expander matching problem remains a scraped chat about the expander matching problem no matter how thoroughly it has been anonymized. The training-data question for intellectual labor is not a privacy question, and treating it as one may be exactly how a “that did not happen” stays technically true while missing the point.

A pattern, and a question of norms

Thom’s post ties the episodes together explicitly. He argues that Bubeck’s conduct in the Buckmaster–Alpöge affair, combined with Sellke’s narrow answer to him months earlier, “deepen[s] the concern that there is a loss of moral compass” among the people setting OpenAI’s norms for how it treats mathematicians’ unpublished work. He now reads Sellke’s original answer as, at minimum, unjustifiably broad, and at worst deliberately misleading.

OpenAI has not yet issued a specific response to Thom’s posts. But the trajectory of the past week suggests the pressure will mount. The company already walked back parts of its framing in the Buckmaster–Alpöge case once the written record surfaced, and the exchange with Thom predates that controversy by weeks — meaning it cannot be dismissed as a misunderstanding born of this week’s heat.

The stakes extend well beyond one proof or one lab. Frontier models are increasingly trained on the aggregate text of their own users, and the people most likely to have deep, substantive, unpublished thinking in their chat histories are researchers themselves. If “de-identified data derived from usage” can legally and technically include a mathematician’s months of private exploration, then every conversation between a researcher and a chatbot is a potential unpaid, unattributed contribution to the next model’s capabilities. Thom’s three-part post is, at bottom, an insistence that this question be answered in plain language rather than deflected with a sentence that sounds like an answer.

For a company that just declared its “automated research intern” milestone met — and that is racing toward a fully automated researcher by March 2028 — the credibility of its human norms around research credit is not a side issue. It is the foundation the entire automation project stands on. On the evidence of the past week, that foundation is being tested by exactly the people best equipped to test it.