business 5 min read

Who Owns Math When an AI Proves It? OpenAI Faces a New Kind of Authorship Crisis

After OpenAI announced breakthrough math results, two researchers independently accused the lab of benefiting from their unpublished work without acknowledgment. The dispute reveals how AI training transparency is collapsing into an authorship crisis.

  • OpenAI
  • AI Training Data
  • Intellectual Property
  • Mathematics
  • AI Transparency

The Proof Is in the Prompt

Two mathematicians. Ten results. One company that refuses to say where it got the ideas.

OpenAI’s recent announcement of ten mathematical proofs was meant to showcase the growing power of AI in formal reasoning. Instead, it became the latest flashpoint in a dispute that cuts deeper than any data scraping controversy before it — because mathematics, unlike text or code, has a centuries-old infrastructure of authorship, priority, and attribution. When an AI model produces a proof, the question is no longer just who trained it. It’s who owns the thinking behind it.

The first crack appeared when Tristan Buckmaster, a mathematics professor at New York University, publicly questioned whether OpenAI’s Codex had benefited from his own interactions with the system. Buckmaster’s concern was straightforward: he had used Codex as a research tool, sharing nascent ideas in conversations that were never meant to be part of a training corpus. If those exchanges shaped the model’s outputs, then OpenAI was converting private intellectual labor into proprietary capability — a dynamic that mirrors debates about code and literature, but with a crucial difference. Mathematical ideas don’t belong to individuals in the same way a poem does. They belong to a community that tracks authorship through publication, citation, and peer review. Breaking that link doesn’t just raise copyright questions. It raises epistemic ones.

Then came Andreas Thom.

Thom, a mathematician based in Europe, noticed that one of OpenAI’s ten results concerned non-sofic groups — a domain he and his colleague Gábor Kun have spent years working on. Non-sofic groups are infinite mathematical structures that resist approximation by finite objects, a deceptively simple definition that hides one of the most stubborn problems in modern algebra. The connection was not coincidental. Thom observed that OpenAI’s result demonstrated a level of familiarity with his and Kun’s techniques that felt too precise to be coincidence, especially since those techniques were neither the most obvious nor the most promising routes to a solution at the time.

What Thom found alarming was not that an AI might produce a correct proof. It was that the proof seemed to carry the fingerprints of his unpublished research — ideas that had not yet entered the public record, ideas that OpenAI had no legitimate way of accessing through conventional channels.

The Mathematics Community Doesn’t Play By Tech Rules

There is a reason this dispute matters more than the typical AI-data controversy. In computer science, a dataset can be scraped, redistributed, and argued over in court. In the humanities, fair use doctrines provide some framing. Mathematics operates under a different regime entirely. Priority matters more than ownership. A proof is published, cited, and attributed. If someone arrives at the same result independently, both are credited. If one arrives first, the naming rights are non-negotiable.

This system assumes that ideas become public through publication. It does not assume that ideas can be extracted from private conversations, training logs, or model weights. OpenAI’s behavior — announcing results built on unpublished work and then quietly amending its writeup after criticism — suggests either a profound ignorance of how mathematics works or a deliberate strategy to test how much attribution it can omit before pushback forces a correction.

Neither interpretation is flattering.

The amendment itself is telling. OpenAI originally failed to acknowledge Thom and Kun’s contributions in its announcement. Only after criticism from the mathematical community did it quietly revise the text. This is not the pattern of a company confident in its transparency. It is the pattern of a company that knew it had crossed a line and preferred to fix it invisibly rather than address it openly.

The Real Question: What Does “Trained On” Mean?

The core of the dispute is not whether OpenAI can prove it did not use Thom and Kun’s work. It is whether OpenAI can prove it did. The company’s models are trained on trillions of tokens. There is no practical way to audit which specific ideas surfaced in which specific conversations. The burden of proof, in this case, should not rest entirely on the accusers. When a company releases a product that claims revolutionary capability and benefits from an opaque training process, it has an obligation to demonstrate that the capability did not emerge from uncompensated, uncredited labor.

Instead, OpenAI has offered silence and edits. That is a non-answer. It is also, unfortunately, a pattern.

The implication extends far beyond mathematics. If a research community’s unpublished ideas can be absorbed into a frontier model without acknowledgment, then the boundary between tool and extractor collapses. Mathematicians use AI tools. AI tools learn from mathematicians. The line between the two is thinner than anyone at OpenAI seems willing to admit.

Who Wins, Who Loses, and What Comes Next

Thom and Buckmaster have forced a conversation that OpenAI would have preferred to avoid. They have exposed the gap between the company’s public claims of openness and its actual practices around attribution. The mathematics community, which has historically been slow to engage with questions of AI authorship, has responded with unusual speed and unanimity. That response matters. It signals that the field is watching, and it is not satisfied with vague assurances.

OpenAI, for its part, wins nothing from this dispute except a record of defensive edits. The company’s credibility among researchers — already fragile — takes another hit. Its competitors, particularly those positioning themselves on transparency, have a clear opening.

What happens next will depend on whether OpenAI and other labs choose to treat attribution as a compliance problem or a design principle. If the former, we will see more quiet amendments, more legal maneuvering, and more researchers choosing to keep their ideas offline. If the latter, we might see a genuine shift toward transparent training data disclosure, model cards that cite source contributions, and revenue-sharing mechanisms for communities whose work shapes frontier models.

Neither outcome is guaranteed. But the mathematicians have drawn a line. The question is whether the companies will respect it.