business 6 min read

OpenAI and the 88-Hour Proof That May Not Be a Proof

OpenAI claims it solved a 90-year-old math problem in 88 hours using 10,000 AI agents. A rival team says its own progress was fed to OpenAI first. The episode reveals what AI-augmented mathematics really looks like—and how much of it is theater.

  • OpenAI
  • AI Mathematics
  • Navier-Stokes
  • Millennium Prize
  • AI Verification
  • Theorem Proving

The headline that is not the headline

OpenAI says it solved a 90-year-old Millennium Prize problem in 88 hours. The company is also saying, repeatedly, that it does not intend to claim the prize. Those two statements cannot coexist quietly. One is a product launch. The other is damage control.

What actually happened is more interesting than the headline and far less settled than OpenAI wants you to believe.

The claim

On September 5, OpenAI announced that a cluster of roughly 10,000 AI agents—bots trained on an internal, unreleased model significantly more capable than any public release—had found a partial solution to the Navier-Stokes existence and smoothness problem. The system exchanged nearly three million messages. It consumed 130 billion output tokens. By OpenAI’s own pricing, that compute footprint would have cost roughly $10 million if billed externally.

The result resolves two of the four statements the Clay Mathematics Institute requires for the full proof. OpenAI says it is publishing the finding to demonstrate the capability of its models. It says it will not pursue the $1 million prize.

In the AI sector, that sentence reads like a plea bargain.

The accusation

Same day. Hours before. Tristan Buckmaster, a professor at New York University, posted a detailed account. He and Levent Alpoge, a mathematician at Anthropic, had been working on the same problem using OpenAI’s Codex. Buckmaster claims that on September 3, information about their progress reached OpenAI. The very next day, OpenAI published a Navier-Stokes result.

Buckmaster’s account includes email exchanges in which he challenged OpenAI’s timeline and methods. His central grievance is procedural rather than mathematical: whether OpenAI derived work product from users of its own platform, however indirectly.

OpenAI’s response was calibrated. The company congratulated the “concurrent work” of Buckmaster and Alpoge. It said it had not seen their work through any means before its public release. It acknowledged a possibility it called “unlikely”: that de-identified data derived from their usage of Codex improved the models powering OpenAI’s internal tool. It then noted that its proofs differ significantly and that even the precise results proved are different.

The word de-identified does heavy lifting here. It is the standard disclaimer that lets companies admit a risk without accepting liability. If Buckmaster’s interactions with Codex influenced the weights of an internal model that then generated the agents tasked with Navier-Stokes, OpenAI may have used the outputs of unpaid researchers without their knowledge. That is not theft in the criminal sense. It is the business model running as designed, and the problem is that the design has a moral blind spot.

What the numbers actually say

Let us be concrete about what 10,000 agents and 130 billion tokens represent. This is brute-force search dressed in a conversational interface. The agents talk to each other. They propose lemmas. They verify steps. They backtrack. The process mirrors how large-scale theorem proving already works in systems like Lean and Coq—except at industrial scale and without human curators guiding the proof structure.

Two of four statements is a partial result. The Clay Institute does not award prizes for partial results. It also does not accept proofs generated by AI without human-verified, machine-checkable formalizations. OpenAI has not provided either. The company’s announcement is therefore not a mathematical proof. It is a demonstration of a capability that, if scaled and paired with formal verification tooling, could one day produce one.

That distinction matters because the market is pricing this announcement as if it were the former.

The real question: verification, not generation

The episode exposes the reliability frontier of foundation models in domains where truth is binary. A language model can generate text that looks like a proof. It can also generate text that looks like a proof and contains a subtle error in step 47 of 200. The only way to know is to run the argument through a formal checker. That is slow. That is expensive. That is also the entire point of a proof.

OpenAI’s 88-hour run did not include a formal verification step. It produced a human-readable argument that two sub-statements of the full conjecture appear to hold. Independent mathematicians have not examined the detail. The Clay Institute has not weighed in. A professor who believes his own work was siphoned into the pipeline is publicly disputing the company’s timeline.

This is not a failed proof. It is an unverified claim. The difference is everything.

Who wins, who loses

OpenAI wins attention. The stock-adjacent narrative around AI-as-researcher gets another chapter. The company’s internal model now has a flagship demo that no public benchmark can match. That demo will be cited in fundraising decks, enterprise pitches, and regulatory testimony for months.

Buckmaster and Alpoge lose time. Their work entered a system they did not sign up to fund. Whether that system is legally culpable depends on the terms of service they accepted when they started using Codex. Whether it is ethically sound is a separate question, and one the industry has not answered.

The broader math community loses standards. When a partial result from an unreleased model getsheadline treatment equal to a century-old conjecture, the signal-to-noise ratio in mathematical communication degrades. Every future claim will be judged against this moment. Some will be more rigorous. Most will be less.

What happens next

Three scenarios are plausible.

First, a formal verification effort runs over the next few months and either confirms the partial result, finds a gap, or both. That would be the honest path and the slowest.

Second, OpenAI releases a formalized version of its argument in a proof assistant. If the machine checker accepts it, the mathematical community will engage seriously regardless of how the result was discovered. The origin story is secondary to the output. This is the best-case outcome for everyone except the researchers who feel their contributions were appropriated.

Third, the claim fades into the background noise of an industry that produces one breakthrough narrative per week. The agents go back to cheaper tasks. The internal model gets deployed for coding assistants and customer support flows. The Navier-Stokes result becomes a footnote with a press release attached.

OpenAI says it does not want the Millennium Prize. But the company has already monetized the claim. That is enough. The question now is whether the math community will treat this as a method or as a marketing event—and whether it can afford to do either without setting a precedent that rewards speed over rigor.