
Is this the “mathocalypse”? Why the latest publication of OpenAI results shocked mathematicians
Earlier this week, OpenAI released a trove of hundreds of mathematical results that it said could “push the frontiers of human knowledge.” Produced largely by a never-before-seen artificial intelligence (AI) model, the 722 articles focus on 372 open-ended problems covering everything from algebra and geometry to theoretical computer science.
Several papers have since been retracted or edited, but the scale and form of the publication – variously described as a drop, a dump, a carpet bomb and a mathocalypse – sent shockwaves through the mathematics community.
Some papers claim significant advances on high-profile and long-standing problems, including the Riemann hypothesis and the Birch-Swinnerton-Dyer conjecture, each of which carries a US$1 million (A$1.4 million) bounty if solved as one of the seven Millennium Prize problems.
Why did OpenAI do this? And can we be sure that the claimed solutions are correct? None of these questions seem to have particularly obvious answers.
A strange way to enable progress
OpenAI says its motivation in releasing the enormous amount of results is to “enable further advances in mathematics.” However, the distribution methods are not necessarily conducive to this objective.
An experienced colleague of mine found one of his favorite problems among the solved ones and tried to read the accompanying document. He told me it was so unintelligible that if he had received it as an editor at a mathematics journal, “it would have gone straight into the trash.”
Even OpenAI’s large ChatGPT language model (GPT-5.6 Sol, to be precise) was skeptical when I posed it, describing at least one of the much-publicized results as “a serious hallucination…(which) should never be cited, submitted, or released as evidence without a full expert audit.”
In other words, despite the revolutionary nature of some of the results obtained, the way in which at least some of the accompanying articles are written is not understandable, even to experts.
Change the conversation
OpenAI may also be looking to move the debate forward in relation to its last major mathematical publication. In September, the company released a controversial solution to the Navier-Stokes problem, another $1,000,000 millennium problem.
Mathematicians Tristan Buckmaster and Levent Alpöge, who worked on the problem themselves, claimed that OpenAI accessed their data and used their ideas, then devoted about $15 million of computing power to find a solution. OpenAI has denied these allegations.
Another possible broader motivation for OpenAI is the hot topic of artificial general intelligence (AGI). This is a hypothetical type of AI that could outperform humans in virtually all cognitive tasks.
Mathematics – especially pure and abstract mathematics – is a discipline that relies heavily on often very complex arguments to prove the truth of mathematical statements. This is often considered a very difficult subject for most people.
Therefore, the ability of OpenAI models to perform research-level mathematics may lend credence to the idea that the models are getting closer to the holy grail of AGI. This idea could benefit OpenAI before the company’s planned IPO. The company hopes to achieve a valuation of up to $1.4 trillion, or about 4% of the entire U.S. gross domestic product.
How can we be sure the math is correct?
One of the key characteristics of mathematics (especially pure mathematics) is its basis in objective truth. Statements are either true or false, and there is almost never any ambiguity.
For centuries we have recognized mathematical truth by consensus among mathematicians. In particular, new works are evaluated by experts who check their accuracy before being published and becoming part of the “literature”.
This process can take a long time. Some longer technical documents may take years to review.
Thus, the final test of OpenAI’s results will be what the mathematical community makes of it after proper review.
OpenAI has already withdrawn three of the papers due to a basic error and edited several others due to errors that invalidated their results. This suggests, at the very least, a lack of sufficient control before publication.
Machines to control machines
Over the past decade or more, there has been a trend toward “formalizing” results with software systems such as Lean. These systems build mathematics from basic principles or axioms, meaning the software can verify each logical step of a result.
Until recently, formalization efforts were largely focused on theory and already solved problems. An early example is the 2008 formalization of the decades-old four-color theorem in graph theory, concerning the minimum number of colors required to color a map.
Formalizing a result requires a lot of knowledge and effort. However, AI systems are now being used for “autoformalization” – automatically translating a result into a machine-verifiable proof.
To encourage confidence in the accuracy of its results, OpenAI claims to have formalized 300 of the key findings at the time of writing.
However, not all formalizations are implemented correctly – and there is evidence of problems with AI-generated formalizations, notably that of Navier-Stokes. It is currently unclear whether the mathematical community will accept OpenAI’s formalizations.
Enthusiasm – and concern
Reactions to OpenAI’s downfall have been mixed, ranging from “obviously the most important moment in the history of mathematics” to “mathematicians are not so thrilled about this giant pile of turd dumped on our doorstep.”
There is certainly great excitement about progress in solving important mathematical problems in various fields, as experts attempt to comb through the results.
For many, this enthusiasm is also mixed with uncertainty and worry: what does the volume and speed of these discoveries mean for the future of mathematical research, and how do we contribute to it? And who should take on the task of trying to make sense of arguments that initially read like AI crap?
These papers, along with other AI-assisted developments, have forced the mathematics community to think about what we value. This is particularly important for students and early career researchers, who must grapple with the idea of their research being co-opted by an ambitious AI user, or how their work will be valued.
And this conversation will continue long after the dust has settled on this latest tranche of discoveries.
Gn bussni