Andreas Thom · 2026-09-09 · major
Andreas Thom — a second mathematician questions OpenAI on his private chats
Andreas Thom says he spent months discussing the expander matching problem with ChatGPT, then asked OpenAI whether those chats reached the model that produced its non-sofic group result. He calls the answer he got incomplete.

A working mathematician asks OpenAI to say plainly whether his private ChatGPT sessions fed the model that solved his field's problems.
Quick facts
| Who | Andreas Thom, mathematician |
|---|---|
| Posted | September 9, 2026, on Mathstodon |
| The ask | Did his ChatGPT chats enter training data? |
| OpenAI's reply | "That did not happen" — Mark Sellke |
| OpenAI's other wording | Cannot rule out de-identified data helped |
| Earlier case | Tristan Buckmaster (NYU) and Levent Alpöge |
| Discussion | 364 points, 457 comments on Hacker News |
What is it?
Andreas Thom posted on Mathstodon that he had discussed the expander matching problem with ChatGPT over months, and that OpenAI later announced a model result constructing the first non-sofic group. Thom asked OpenAI two questions: whether his conversations entered the training data, and whether they were reachable by the system that produced the result. He says the reply he received only answered the second.
How does it work?
The reply Thom quotes came from OpenAI researcher Mark Sellke — "Regarding your conversations with ChatGPT: that did not happen" — which rules out direct access during the solve but says nothing about training. Thom sets that against OpenAI's own written statement in the parallel Navier–Stokes case, where the company said it "cannot rule out that de-identified data derived from their usage of our products helped improve our models". He treats the two statements as pulling in opposite directions.
Why does it matter?
Thom is the second mathematician in a week to ask this. Tristan Buckmaster of NYU, working with Levent Alpöge, says information about their progress reached OpenAI before they published, and OpenAI answered that "no specific user data was accessed" while adding that it "cannot rule out that de-identified data derived from their usage of our products helped improve our models". Researchers who put unpublished work through a frontier lab's tools now have a concrete reason to ask for that answer in writing.
Who is it for?
researchers using AI tools on unpublished work
Frequently asked questions
- What exactly has OpenAI said about user data and its math results?
- OpenAI's public statement reads: "We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed." The company then added: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models." Andreas Thom argues those two sentences pull against each other.
- Is Andreas Thom the only mathematician raising this?
- No. Tristan Buckmaster, a mathematics professor at NYU working with Levent Alpöge, said that "information about our progress had been passed to OpenAI" and that an "entire team had been working on the problem" with "an insane amount of compute". Andreas Thom's Mathstodon post is a separate case, about the expander matching problem and OpenAI's non-sofic group result.
- What did OpenAI's Sébastien Bubeck say to Tristan Buckmaster?
- By Buckmaster's account, OpenAI mathematician Sébastien Bubeck asked him to drop a collaborator's credit as a "compromise", then said "Why would you ruin your career?" When Buckmaster pushed back, he says Bubeck replied: "If you don't want me to be nice, then I don't have to be nice." These are Buckmaster's characterisations of private conversations, reported by TechCrunch.
- Should researchers stop putting unpublished work into ChatGPT or Codex?
- Andreas Thom does not tell anyone to stop, but his post makes the gap clear: a denial that a system read your files during a solve says nothing about whether your text became training data earlier. His position is that labs should answer both questions separately and in writing. Researchers with unpublished results can ask for exactly that before uploading drafts.