OpenAI's 722 Math Papers Don't Prove What It Claims
OpenAI published 722 AI-generated math papers on GitHub, claiming breakthroughs on problems related to the Riemann Hypothesis. The catch: many lack formal proof, the advisory group that advised against this very practice, and a pattern of overreach.
The Claim That Outgrows the Evidence
OpenAI announced on October 6 that its internal frontier model produced 722 math manuscripts, many backed by Lean formal proofs, and released them on GitHub. The press emphasis landed on the Riemann Hypothesis. The reality is more complicated and, in several ways, more interesting.
The model did not prove the Riemann Hypothesis. It produced results about related problems — specifically, regions where zeros of the zeta function cannot exist. One paper on that topic was edited by a human for readability. The rest were generated using a consistent pipeline on an unpublished model, averaging the compute equivalent of three hours of ChatGPT Pro reasoning per manuscript.
That matters. This is not a single breakthrough. It is a volume operation.
The Filter Was the Point
OpenAI started with roughly 4,000 problems. It kept 722 after filtering for significance. That means nearly 80 percent of outputs were discarded. The company did not publish the full set. It did not describe the rejection criteria in detail. It offered no count of problems it could not solve.
This is the first red flag for any skeptical reader. In normal mathematical practice, negative results and failed approaches are as valuable as successes. In OpenAI’s release, they are absent. The 722 represent selection, not completeness.
The selectivity itself is not inherently dishonest — every journal practices it — but the asymmetry is striking. A research group publishing 722 verified results alongside zero failed attempts creates a distorted picture of capability. Readers unfamiliar with mathematical culture will absorb a narrative of relentless success. That narrative has marketing value. It also has epistemic cost.
Second-order consequences are already visible. Twitter threads and blog posts are citing the release as evidence that AI has “solved” or “cracked” the Riemann Hypothesis. The gap between the company’s careful framing and the public’s interpretation is widening faster than anyone inside OpenAI can correct it. By the time a clarification circulates, the headline has already done its work.
Lean Is Not a Guarantee
Many papers include formal proofs checked by Lean. That is notable. Lean proof verification is one of the few methods that can remove human ambiguity from a mathematical argument. A verified Lean proof is close to a certificate of correctness.
But not all results are formalized. OpenAI acknowledged this openly and said problems may exist in non-formalized work. The company promised fixes. That is responsible on its face. It is also a deferred audit. Anyone who wants to verify the output must now comb through 722 manuscripts, identify which lack Lean coverage, and check them manually. The burden has shifted from the publisher to the community.
This is a structural problem, not a moral one. OpenAI cannot formally verify everything it generates at this scale. But it also cannot claim proof for everything it publishes. The ambiguity between the two camps is where the controversy lives.
There is a deeper tension here as well. Lean formalization requires translating informal mathematical reasoning into a machine-readable format. That translation process itself can introduce errors — omitted lemmas, misattributed hypotheses, or subtle logic gaps that slip through automated checks. The fact that a proof exists in Lean is strong evidence of correctness, but it is not infallible. The community has seen examples where formally verified proofs later required revision when edge cases emerged.
The Advisory Group That Said No
OpenAI consulted the Advisory Group on Mathematics and Artificial Intelligence, or AGMAI, before publishing. The group’s September recommendations are instructive. They called for independent repositories, full disclosure of models and prompts, timing and cost transparency, formal verification where possible, and counts of unsolved problems alongside solved ones.
Then AGMAI issued a clarifying statement the same day as the release. It said its advice did not endorse OpenAI’s results. It said the group does not represent the mathematics community. It said the evaluation belongs to mathematicians, not advisers.
The subtext is clear: AGMAI recommended standards that OpenAI only partially met, and the group is already distancing itself from the outcome. The advisory relationship looks less like a seal of approval and more like a liability boundary.
What this episode reveals is the growing friction between AI labs and the institutions that might hold them accountable. Advisory groups like AGMAI occupy an awkward middle ground — they lend credibility without commanding authority. Their recommendations are easy to ignore when convenient and easy to cite when useful. Neither outcome strengthens their position. The mathematics community, already suspicious of tech companies co-opting its language, is watching closely. If OpenAI treats AGMAI’s guidance as optional rather than foundational, other labs will interpret that as permission to do the same.
The Pattern Is the Problem
This is not the first time OpenAI has stumbled into overreach. In September, the company claimed its internal model had solved the Navier-Stokes existence and smoothness problem, one of the Millennium Prize Problems. Mathematicians working in the field pushed back hard. The claim did not survive scrutiny.
In August, OpenAI reported results from an internal version of its next flagship model called Astra on 10 open problems. In October 2025, an executive’s post about solving Erdős problems with GPT-5 was deleted after criticism.
The pattern is consistent. OpenAI generates ambitious outputs, announces them publicly, and then faces pushback from experts who find gaps in the reasoning or the framing. Each episode erodes trust. Each cycle normalizes the idea that AI can solve hard problems, even when the evidence is thin.
The cumulative effect matters more than any single incident. When AI companies repeatedly announce half-truths about mathematical progress, the broader public loses the ability to distinguish genuine breakthroughs from performative ones. This creates a credibility trap: even when OpenAI later produces something truly significant, the audience trained on sensational headlines will default to skepticism. The company is spending its reputation capital on incremental announcements.
Who Wins, Who Loses
OpenAI wins attention. The GitHub repository is searchable. The 722 papers are indexed. Researchers can mine them for ideas, lemmas, or computational approaches. That is real value. Mathematical research has always borrowed from adjacent work. An AI that produces large volumes of formally checked output could accelerate that process.
Mathematicians lose something subtler. Their authority over what counts as proof is being diluted. When a frontier model generates results that partially pass verification tools, the line between machine assistance and machine authorship blurs. Peer review cannot keep pace with the output volume. The system will adapt, but the adaptation will be messy and slow.
The broader public loses clarity. Headlines will say OpenAI proved something about the Riemann Hypothesis. They will not say the company proved nothing of the sort. The gap between the headline and the document will widen with each release.
There is also a less discussed loser: the mathematical community’s collective patience. Every time a company misframes its results, mathematicians must spend time correcting the record instead of doing their own work. That time is non-renewable. Over repeated cycles, it becomes a real resource drain — one that disproportionately affects early-career researchers who are already stretched thin.
What Comes Next
OpenAI plans a workshop and special programs to help mathematicians understand AI-generated results. It says it is working toward responsible general release of the model. It is considering community-run publication venues beyond GitHub. These are sensible steps. They do not erase the problems with this release.
The most important question is not whether AI can help do mathematics. It can. The question is whether the companies building these systems will adopt standards that let the community verify, replicate, and reject their claims at scale. OpenAI’s latest output is a milestone. It is also a warning. The model can produce. The infrastructure for accountability is still catching up. Until that infrastructure matures, the math community should treat these releases with the same cautious optimism it applies to any preprint — grateful for the signal, unconvinced by the hype, and ready to hold the line when the two are confused.