science 5 min read

OpenAI Claims 100 Solved Math Problems, But Can Anyone Verify Them?

OpenAI says its latest model cracked 100 unsolved math problems including a Millennium Prize难题. Mathematicians fired back—then OpenAI quietly built a review panel. The credibility gap between claiming a result and proving it just got wider.

  • OpenAI
  • AI Governance
  • Millennium Prize
  • AI and Mathematics
  • Academic Advisory

The Claim That Outpaced the Proof

OpenAI says its internal model, which began training on August 28, solved over 100 long-standing unsolved math problems. Among them: the Navier-Stokes existence and smoothness problem—one of the seven Millennium Prize Problems worth $1 million each. The company announced on September 9 that it cracked Navier-Stokes in 88 hours. The full list of 100+, still unreleased in detail, is arguably more consequential than any single result.

Mathematicians did not wait for the papers. A letter signed by 25 of the field’s most prominent figures, including June Huh of Princeton and Fields Medalist Terence Tao, warned that the race to solve problems via AI could inflict long-term damage on the discipline. OpenAI responded not with a rebuttal but with a structure: an independent advisory group of nine mathematicians, based at the Institute for Advanced Study in Princeton, to review and co-ordinate the timing of AI-generated results.

The structure is a smart public-relations move. The underlying tension is far harder to resolve.

What This Actually Means for Math

The immediate issue is not whether OpenAI’s claims are true—it is whether anyone can verify them. A mathematical proof is not an answer. It is a humanly inspectable chain of reasoning. An AI can produce a string of symbols that encodes a proof. It can also produce a string of symbols that looks like a proof and is wrong. The difference is not always visible to a trained mathematician in real time, let alone to a machine.

June Huh’s name on that open letter carries particular weight. He won the Fields Medal in 2022 for work connecting combinatorics to algebraic geometry—fields where intuition and human insight have traditionally been inseparable from discovery. His signature signals that the concern is not Luddism. It is the fear that a flood of AI-generated results, verified at the speed of corporate release rather than the pace of peer review, will reshape how mathematics is done, not just what gets done.

The advisory group’s mandate reflects that fear. It reviews AI-generated results and helps coordinate when and how they are made public. It cannot influence OpenAI’s development pace. It receives no pay. It is, in every operational sense, a brake on publication, not a brake on production.

That distinction matters. OpenAI is still releasing claims faster than the independent panel can review them. The group is a review board, not a gatekeeper.

The Verification Bottleneck

Here is what the global AI industry should understand: OpenAI’s math milestone, if true, sits at the same inflection point as its earlier claims about coding, science, and reasoning. Each time, the pattern repeats. An AI produces results at a scale and speed no human lab can match. The results are impressive. The verification gap is enormous.

The Millennium Prize Problems were chosen precisely because humans could not solve them—not because they are obscure puzzles but because they sit at the edge of mathematical understanding. Solving one requires not just computation but conceptual innovation. If OpenAI’s model has actually solved even one, the implication is staggering. If it has produced a proof that cannot yet be inspected, the implication is equally staggering and far more dangerous.

The advisory panel’s existence acknowledges this without closing the gap. Nine mathematicians, working independently and without compensation, cannot verify 100 results before the next corporate announcement lands. The timeline is structurally asymmetric. OpenAI controls the release. The panel controls the review. The review will always lag.

Who Wins, Who Loses

The companies win credibility. They demonstrate that their models can touch problems previously thought immune to automation. Investors and customers see a concrete milestone, not a vague capability claim. The stock market and the narrative both reward the announcement, regardless of whether the proofs survive scrutiny.

The mathematics community loses leverage. Every new AI-generated result pulls the conversation toward a deadline the company sets, not a review cycle the field controls. Young researchers who might have spent years on a single Millennium Problem now face a landscape where a company claims the solution and publishes it in a press release.

The broader AI ecosystem gains a template. This is the second time OpenAI has faced this exact pressure cycle—claim, backlash, advisory panel. The same structure could repeat for biology, chemistry, and any domain where AI outputs are rapidly outpacing human verification capacity. The advisory group is not a unique solution. It is a replicable damage-control mechanism.

The Second-Order Story

The real story here is not that AI solved a hard math problem. It is that OpenAI needed a damage-control mechanism so quickly. The company anticipated the backlash and built the panel before any proof was independently examined. That foresight is telling. It suggests OpenAI understands the legitimacy risk better than the field does—and that the risk is structural, not situational.

For English-language readers outside Korea, the June Huh connection is the bridge. A Korean-American Fields Medalist signing a joint letter with Terence Tao places this squarely in the global conversation about AI and mathematical integrity. It also highlights a demographic that is often invisible in Western AI coverage: the mathematicians outside the U.S. and Europe who are most directly affected by the shift toward automated proof generation.

What Comes Next

Three things to watch. First, whether any of the 100 claimed results produce independently verified proofs within a year. Second, whether the advisory panel gains actual veto power over publication timing, or remains advisory in name only. Third, whether other AI labs adopt the same template—claim, face backlash, build an independent review body, repeat.

If the proofs survive scrutiny, mathematics enters a new era of assisted discovery. If they do not, the episode becomes a case study in the cost of outpacing verification. The advisory panel does not determine the outcome. It only determines how the world sees the failure—or the success.

OpenAI’s move was always going to be judged by results, not structures. The structure is in place. The results are not yet public. The next 12 months will decide which side of that gap the industry stands on.