business 5 min read

OpenAI's Math Claims Expose AI's Credibility Gap

OpenAI released hundreds of math proofs that fell short of the very standards it claimed to follow. The gap between AI's marketing and verifiable results is widening — and mathematicians are calling it out.

  • Artificial Intelligence
  • OpenAI
  • Tech Industry
  • AI Benchmarks
  • Mathematics

The advisory group’s guidelines were clear. OpenAI mostly ignored them.

When OpenAI released hundreds of claimed solutions to some of the world’s hardest math problems this week, the lab said it had consulted an advisory group of elite mathematicians — presumably to avoid the controversy that accompanied its last major problem-solving announcement. It didn’t help.

OpenAI fell short of the very standards it claimed to follow, particularly where the mathematicians emphasized human understanding as a non-negotiable component of any real mathematical result. That gap between declaration and delivery is becoming the defining story of AI’s credibility problem.

The Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton University’s Institute for Advanced Study, released guidelines for frontier labs in late September. Nine prominent researchers worldwide compose the group. Their first recommendation was blunt: stop testing advanced mathematical problems on proprietary models. OpenAI’s release explicitly said it was doing exactly that.

The lab followed some principles — releasing results quickly, including information about how models reached their conclusions. But selectively. Just 10 of the 719 manuscripts included the model’s chain of thought. For the papers humans don’t yet understand, AGMAI recommended formalization through tools like Lean, a programming language that verifies proofs by compiling them as code. Only 42 percent of OpenAI’s proofs had undergone that process.

More importantly, the group asked OpenAI to help fund the human mathematicians who would be required to make those solutions meaningful. No indication the lab committed to that. When AGMAI asked for machine-readable metadata correlating natural language proofs with formal artifacts — essentially a map so humans could follow the translation — OpenAI didn’t provide it.

A mistranslation that shouldn’t be ignored

A paper this week from mathematicians at the University of Cambridge and King’s College London documents the concrete risk of skipping human oversight. The researchers found at least two discrepancies between the natural language explanation of a solution and the Lean code behind OpenAI’s proposed proof for a problem derived from the Navier-Stokes equations — the famous fluid dynamics equations that come with a million-dollar Millennium Prize.

The discrepancies don’t necessarily disprove the solution. But they do demonstrate that frontier labs cannot simply automate the translation from human-readable math to machine-verifiable code and call it done. The process is fragile. The paper’s authors wrote that these autoformalized Lean proofs “should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to.”

That’s a low bar. Every published mathematical result goes through peer review before it enters the canon. OpenAI’s releases bypass that entirely.

The understanding problem

Terence Tao, the Stanford mathematician and Fields medalist who has been openly critical of OpenAI’s approach, put it plainly on social media after the release: problems are being solved autonomously by “AI prompters who have no interest in the broader field itself once their initial target is solved, and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field.”

That last point is the core issue. When a human mathematician discovers a result, they take responsibility for it. They publish it. They present it at seminars. They field questions from colleagues. The community scrutinizes it, extends it, applies it. Understanding deepens. The result becomes part of the shared knowledge base rather than a black-box output.

“When a model is prompted to solve a hard problem and spits out a solution, there is not human understanding of them at the point of release, and now the work begins,” said Harvard mathematics professor Melanie Wood. In other words, the model’s release isn’t the end of the process — it’s the beginning. And right now, no one is picking up the ball.

Why this matters beyond mathematics

The OpenAI math episode isn’t just about whether a handful of proofs check out. It’s a case study in what happens when an industry accelerates past verification in pursuit of narrative wins. The lab wanted to position itself as responsible — consulting an advisory group, releasing quickly, providing transparency. But the substance didn’t match the framing.

This is the pattern repeating across AI benchmarking. Models claim scores on evaluations that measure something close to reasoning but not the full thing. Results get published before peer review. Controversial claims circulate because they’re exciting, not because they’ve survived scrutiny.

The mathematicians at Cambridge and King’s College are clear about what’s at stake: if you want AI-generated proofs to be accepted as real contributions to the field, they need to survive the same processes that real contributions survive. Not a press release. Not an advisory group photo op. Actual human engagement with the work.

AGMAI didn’t respond to TechCrunch’s request for a more thorough evaluation of OpenAI’s latest release. That silence may be telling. The group’s guidelines were specific. OpenAI’s results were publicly available. The gap between the two is now a matter of record.

What happens next will depend on whether the mathematical community treats these releases as starting points for collaboration or as curiosities to file away. The first path benefits everyone. The second turns AI-generated math into exactly what Tao warned about: outputs without owners, solutions without understanding, and claims without accountability.