science 6 min read

AI Solved a 40-Year Math Problem in 2 Months. What Happens Next for Math?

A Korean-American mathematician used AI to crack a four-decade-old problem called M23 in just two months. The result is forcing the math world to confront a question it has avoided: if machines produce proof, who gets credit?

  • Korean Science
  • AI Research
  • AI Mathematics
  • Fields Medal
  • Inverse Galois Theory

The problem that resisted humanity for four decades did not yield to genius. It yielded to a tool.

At the top of a preprint paper on arXiv last month sat a 23rd-degree polynomial with coefficients stretching to 14 digits — a numerical monster that mathematicians had tried and failed to construct for 40 years. Two months later, it existed. The problem, known as M23, sits inside inverse Galois theory, one of the most abstract branches of pure mathematics. It had been ranked among the 26 most stubborn open problems in its subfield. Twenty-five of those problems have since been solved. M23 held out longest, until a team led by Professor Kyu-Hwan Lee at the University of Connecticut bent it to an AI system at the American Institute of Mathematics workshop in late May.

What makes the result unsettling is not merely that it was solved. It is how it was solved, and what the method implies for every institution that measures mathematical greatness.

A miracle, formally classified

The word “miraculously” appears twice in the paper. In a field that treats such language as professional malpractice, its use is itself a data point. Lee told Chosun Ilbo that the term reflects the genuine bewilderment of the human researchers: the AI produced results that were correct but opaque, the way AlphaGo’s famous move 37 stunned professional Go players before the mathematics of the position was understood.

“The AI moved ahead and the humans had to follow behind to understand,” Lee said. That sequence — machine discovery, human comprehension — is the inversion of everything the mathematical tradition has been built on.

The division of labor in the M23 project was clear enough to describe but impossible to quantify. The AI handled brute-force computation, code generation, and hypothesis validation that would have consumed months or years of human labor. Lee estimated that a doctoral student might spend three years on a problem the same AI solved in three days. The humans set direction, assessed plausibility, and eventually built the understanding that turns a computed answer into a mathematical result.

Lee compared the partnership to an orchestra: you cannot fairly separate the conductor’s contribution from the violinist’s and assign a ratio. But the metaphor also hides an uncomfortable asymmetry. The violinist can be replaced. The conductor cannot.

The Fields Medal faces an identity crisis

The most consequential implication of the M23 result is not mathematical but institutional. The Fields Medal, awarded every four years to mathematicians under 40, is the discipline’s highest honor. It recognizes individual insight, originality, and sustained intellectual struggle. All of those values are destabilized when the core labor of mathematical research — computation, verification, exploration of solution spaces — can be offloaded to a machine that operates on a timescale incomprehensible to a human researcher.

Lee cited a prediction circulating among some mathematicians: the 2030 Fields Medal could be the last. The reasoning is stark. If AI collaboration becomes the default mode of research, drawing a meaningful line between human achievement and machine-assisted discovery will become impossible. Awards built on the premise of individual human brilliance will either lose their legitimacy or be forced to redefine what counts as “human.”

This is not speculative. The workshop where M23 was solved took place at AIM, an institution specifically designed to foster collaborative problem-solving. The result emerged from a human-AI loop that no single person could have sustained alone. The peer-review apparatus, the doctoral training pipeline, the tenure process — none of these were designed for a world in which the fastest path to a result runs through a black box.

Five to ten years of noise, then a golden age — or a crisis

Lee forecast a turbulent transition. Young mathematicians trained in the traditional way, who spent years learning to compute by hand and to see patterns through intuition, will find their preparation misaligned with the new reality. The market for their specific skills will shrink. The question of what to train them for will remain unanswered for years.

He is not pessimistic about the outcome. He described a scenario in which mathematicians who treat AI as an “Iron Man suit” — an external capability that extends rather than replaces human judgment — will tackle problems that were previously out of reach. The constraint is no longer computational capacity. It is imagination.

But imagination without the ability to verify is dangerous. Lee emphasized that the stronger AI becomes, the more critical human mathematical reasoning becomes as a control mechanism. A mathematician who cannot audit an AI’s output is not a partner to the machine. They are its passenger. The future elite, in his view, will be those who can think fast enough to stay in the driver’s seat.

Who wins, who loses

The winners are mathematicians with deep foundational training who can operate at the level of proof structure and conceptual insight while delegating computation. They will solve problems that were previously infeasible and build reputations on judgment rather than grind.

The losers are those whose value proposition is computational labor — the slow, careful, human calculation that AI now performs in minutes. This includes a significant portion of graduate training, where students are expected to spend years mastering techniques that will soon be automated. A problem that defined a doctoral student’s three-year trajectory can now be completed in a weekend.

The institutions most exposed are those built around human-only achievement metrics: the Fields Medal, the Clay Mathematics Institute’s Millennium Prize Problems, tenure committees that weigh publication count and speed. None of these frameworks have categories for “supervised machine assistance.”

The deeper question: what is a proof when no human wrote it?

The M23 result did not produce a proof that AI generated and humans stamped approvingly. It produced a result that AI computed and humans eventually understood and verified. The gap between those two scenarios is currently wide. It will not stay wide.

As AI systems absorb more of the intermediate steps of mathematical research — the lemmas, the computational checks, the exploratory calculations — the boundary between human-generated and machine-assisted knowledge will blur. The question that will eventually force itself into the open is whether a proof that no single human mind can fully trace is still a proof in the classical sense. Mathematics has always accepted computer-assisted verification, as in the Four Color Theorem. But that was a case of humans delegating tedious checking. M23 points toward a different future: humans delegating discovery itself.

Lee’s assessment is worth taking seriously because he has lived the transition. A Seoul National University-trained pure mathematician specializing in algebra and number theory, he has watched the field’s timeline compress from years to months in less than six months. His warning is not anti-AI. It is pro-judgment. The mathematicians who thrive will not be those who compete with machines on computation. They will be the ones who can ask the right questions and recognize when the machine has answered them wrong.

The 40-year problem that fell in two months is not just a result. It is a signal. The signal says the center of gravity in mathematics is shifting, and the institutions that measure it have not yet noticed.