OpenAI's Math Problem Solve Costs $40 Million — and It May Not Count
OpenAI claims its agents cracked the Navier-Stokes Millennium Prize Problem after burning through 10,000 concurrent agents and an estimated $40 million in compute. But two leading mathematicians say the path to that solution runs straight through their unpublished research data — raising the question of whether this is a proof, or just expensive interpolation.
The Price of a Proof
OpenAI said on Tuesday that its AI agents had solved the Navier-Stokes equations, one of seven Millennium Prize Problems that have stumped mathematicians for decades. The Clay Mathematics Institute is offering $1 million for a correct solution. OpenAI’s team says they’ve found it.
The cost was staggering. Ten thousand concurrent AI agents worked for 88 hours, exchanging 2.7 million messages. Another 17 hours went into verification. The output tokens alone — at OpenAI’s average consumer price — would run roughly $6.5 million, according to an estimate by LisanBench, an LLM benchmark evaluator. Factoring in the far larger volume of input tokens, the total could reach $10 million to $40 million. Sam Altman, when asked, joked on X that artificial intelligence was a bubble.
But the numbers don’t tell the whole story. What really matters is not how much it cost — it’s how they got there.
The Mathematicians Who Say They Were Used
Two researchers, Tristan Buckmaster and Levent Alpöge, immediately raised a concern that goes beyond technical nitpicking. They said OpenAI’s training data includes de-identified information derived from their published and unpublished work on related problems, including the clustering conjecture. The company did not deny it. It said only that it “cannot rule out” that data from users engaging with OpenAI products — which might include conversations about their research — helped improve the models that produced the proof.
In other words, OpenAI’s path to the solution may have been built, in part, on the intellectual labor of the very people who have spent years trying to reach it. The firm maintains it has no direct access to their unpublished work. But the distinction between direct access and indirect ingestion is vanishingly thin when you’re training on the aggregate output of a million users.
What Makes Navier-Stokes So Hard
The equations describe how fluids move — air over a wing, water through a pipe, blood through an artery. They are simple to write. They are brutally difficult to solve. The question is whether smooth solutions always exist in three dimensions, or whether they can develop singularities — points where the velocity becomes infinite. No one has proved either possibility.
A correct solution would transform fluid dynamics, climate modeling, aerospace engineering, and countless fields that depend on predicting how things flow. An incorrect one, dressed in the language of mathematics, is just expensive word salad.
That’s the tension AI brings to pure math. The system produces something that looks formal. It generates lemmas, corollaries, definitions. It follows the shape of a proof. But the question — the actual question that mathematicians care about — is whether the steps are valid. Not whether they sound right. Not whether they’ve seen something like them before. Whether they hold.
Is This a Proof or a Performance
There is a structural difference between discovering a result and verifying one. AI is phenomenally good at the first part, at least when the first part means generating plausible-looking text. It is not clear it is good at the second, which requires checking each step against axioms that the model may not actually understand.
The 105 hours OpenAI logged is impressive for a coordination problem. It is not clear it is impressive for a mathematical proof. The verification time — 17 hours — also deserves scrutiny. Who verified it? Other AI agents, or human mathematicians? If the former, you have a system checking its own work. If the latter, you should hear from those mathematicians. OpenAI has not released the proof for public review.
This matters because the Navier-Stokes problem has resisted not just effort but creativity of a particular kind. It has required insights that no amount of pattern matching can substitute for. It requires, in the language of computer science, a new primitive — a way of thinking about partial differential equations that simply does not appear in the existing literature. The question is whether the model found that primitive, or whether it assembled a collage of existing ones.
The Second-Order Implications Are Real
Even if the proof turns out to be flawed, the episode reveals something important about where the field is headed. Within months, an AI system has attempted to solve a problem that has stood for 90 years. It burned through more tokens in a day than most research groups generate in a year. It cost millions, yes, but the marginal cost of another attempt is close to zero.
The next system will be faster, cheaper, and better at formal verification. The next proof will arrive in hours, not days. The barrier to entry for mathematical discovery — historically a slow, incremental, deeply human process — is collapsing. That is not science fiction. It is happening in real time, with real money, and real reputations at stake.
The question is whether mathematics can absorb that shock without losing what makes it trustworthy. A proof is not just a conclusion. It is a chain of reasoning that any competent human can follow. If the chain is too long, or too alien, or too opaque, it ceases to be a proof in the sense the field has understood since Euclid.
The $1 Million Question
The Clay Institute has not yet acknowledged the submission. It has not invited review. It has not said anything. That silence is itself telling. The institute has spent decades building a standard for what counts as a solution. It will not abandon that standard because a company can afford to generate text quickly.
OpenAI says it will not claim the prize. That raises its own question: why announce a solution you do not intend to defend? The answer may be strategic. The announcement establishes a fact — or at least a narrative — before anyone else gets a chance to. It shifts the timeline. It makes the next step look inevitable.
Sam Altman’s joke about a bubble may have been meant to deflect. But the deeper irony is that the bubble exists precisely because the capability is real, even if the proof is not. The models can do things they could not do before. They can reason, however superficially, about problems that have resisted human attention for generations. The cost is high. The risk is real. And the world is not ready for what comes next.
The next problem will not be Navier-Stokes. It will be something harder, something more consequential, something that no single mathematician could ever solve alone. And the question will not be whether AI can attempt it. The question will be whether we can trust the answer.