business 6 min read

AI Solved Math's Hardest Problem. Now Nobody Knows What It Found

OpenAI's AI produced a verified proof of the Navier-Stokes problem — but the 166-page manuscript is unreadable to humans. The episode raises urgent questions about what happens when machines solve problems we can't understand.

  • Artificial Intelligence
  • OpenAI
  • AI Research
  • Mathematics
  • Navier-Stokes

The Proof Is Right. Nobody Knows Why.

OpenAI spent roughly $15 million computing a solution to one of the hardest problems in mathematics. The company marshaled 10,000 AI agents over 88 hours, burning through 300 billion tokens to generate a 166-page proof that the Navier-Stokes equations behave the way mathematicians have suspected for centuries.

The proof compiles. A program called Lean — a formal verification language — checked it line by line and found no errors. By every technical standard available, the problem is solved.

And yet, as James Maynard of Oxford University put it bluntly: “So far it’s been very difficult to really extract any human understanding from this new AI proof.”

This is the paradox of a moment that will define the next era of scientific discovery. AI has crossed a threshold where it can verify truth without producing insight. The mathematical community now faces a question it has never encountered: what do you do when a machine can prove something true, but no human can follow the logic?

How We Got Here

The Navier-Stokes problem is one of the seven Millennium Prize Problems, each carrying a $1 million bounty from the Clay Mathematics Institute. Established in 2000, these problems represent the most stubborn obstacles in pure mathematics. The Navier-Stokes equations describe how fluids move — water, air, molten steel — yet no one can prove whether their solutions always remain smooth or whether they can unexpectedly break down into chaos.

Tristan Buckmaster, a mathematician at New York University, had been inching toward a solution with his collaborator Levent Alpöge, who works at rival AI company Anthropic. Other顶尖研究者, including Fields Medal winner Martin Hairer at EPFL, had noted that among the seven Millennium Problems, “everyone agreed that this would be the next one that got solved.”

Then OpenAI caught wind.

On September 1, the company announced it was throwing its full compute weight at the problem. Ten thousand agents. Eighty-eight hours. A proof that appeared to arrive faster than the human researchers who had been working toward it for years.

The timing raised suspicions. Buckmaster later told NPR that OpenAI appeared to know more about his approach than the company let on — information he suggested could have come not just from human intelligence but from the agents themselves, scraping public conversations or reading the prompts Buckmaster and Alpöge had used in their own AI-assisted work.

OpenAI denied using their proofs or prompts. But the company’s actions spoke louder. Buckmaster accused them of trying to force him to drop Alpöge from any resulting paper — a move that would have split a collaboration and potentially given OpenAI exclusive access to the human researchers who understood the problem best.

Maynard, who signed a statement alongside 24 other Fields Medal winners, called the whole episode “misaligned goals” — a conflict between AI companies racing for prestige and mathematicians who care about understanding, not just answers.

The Paper Nobody Can Read

The solution sits on the internet now. Anyone can download it. Few have finished it.

Javier Gómez-Serrano, a Brown University mathematician who uses AI in his own research, said the paper is “not written for humans.” Buckmaster agreed: “The paper doesn’t explain what parts are important. What parts are routine? How does the idea feed into other places?”

These are not minor complaints. In mathematics, a proof is not merely a certificate of correctness — it is a story. It reveals why something is true, not just that it is true. Good proofs illuminate. They introduce new tools, new ways of thinking, new connections between previously unrelated ideas. The value of a proof lives in its readability, in what it teaches the reader about the structure of the problem.

OpenAI’s proof teaches almost nothing.

Gómez-Serrano offered a sliver of hope: with “some serious re-writing,” the proof could eventually help advance the field. But as of now, the answer to a question that mathematicians have wrestled with for generations has been produced by a machine that apparently understands it in a way no human does — and can no longer convey that understanding back.

Verification Without Comprehension

The Lean formalization is what gives the community confidence that the proof is correct. Lean is a programming language designed for one purpose: checking mathematical proofs with absolute rigor. If the code compiles, the proof is valid. There is no ambiguity, no loophole, no hope that a subtle error was overlooked by tired human eyes.

This is the real shift happening here. For centuries, mathematical truth depended on human comprehension. A proof was accepted when experts could read it, trace its logic, and convince themselves — and each other — that it held. The social contract of mathematics was built on shared understanding.

That contract is fracturing.

Lean doesn’t need to understand. It needs to compile. And OpenAI’s proof compiles.

This creates a new category of knowledge — verified but opaque, true but incomprehensible. The mathematical community has never had to decide what to do with it before.

What This Means for Science

The implications extend far beyond one problem. If AI can now produce verified proofs that humans cannot parse, then every field that relies on proof — not just mathematics, but theoretical computer science, physics, even cryptography — faces a future where machine-generated results may be correct but inaccessible.

Consider what this means for the scientific method. Science has always depended on reproducibility and understanding. A result that cannot be followed, step by step, by another researcher is not truly science — it is something else. AI-generated mathematics risks becoming a black box: outputs that work, but whose workings are lost behind layers of computation no single human can hold in mind.

Martin Hairer summed up the distinction that matters most: “It wasn’t about only answering this problem, it was about the human understanding behind it.”

The Race Is Already Losing

Perhaps most troubling is what happened here. Two human researchers, working with AI tools, were close to a breakthrough. A company with vastly more resources swooped in, produced a verified result, and left the human mathematicians scrambling to publish their own — still incomplete, still unreadable — work just to stay relevant.

Buckmaster said the rush forced him to publish preliminary results that were “not at the level I’m happy with,” though at least the introduction contained the key ideas. The race for publication has always existed in academia, but AI has accelerated it to a point where the slow, careful work of understanding is being replaced by the fast, brute-force work of generating answers.

Maynard and the other Fields Medal winners called for better alignment between AI companies and the mathematics community. The statement they released on September 11 was a plea for collaboration over competition — a recognition that the most valuable tool in mathematics may be the human mind that can make sense of what the machine produces.

The Question Ahead

OpenAI proved it can solve a Millennium Prize problem. It spent $15 million and 300 billion tokens to do it. The proof checks out.

But in the six months since AI exploded onto the mathematics scene, as academic researchers told NPR, the community still cannot read what the machines have written.

This is not a failure of the AI. It is a failure of something older and harder to fix: the gap between verification and understanding.

A machine can know that something is true. It cannot yet know why — and it cannot yet tell us.

Until that gap closes, the most profound discoveries of the AI era may also be the most incomprehensible ones.