Mathematicians Weigh In on OpenAI’s 722 Math Papers: Promising, But Verification Will Take Years
Four days after OpenAI released 722 AI-generated mathematics manuscripts, The Verge spoke with over 30 mathematicians who concluded that the papers contain genuine breakthroughs, though discerning which results are reliable could take years.
Initial Reactions and Concerns
Mathematicians’ initial assessment of OpenAI’s batch of AI-generated math papers is that they contain “real stuff,” but it will likely take years to determine which papers are correct. OpenAI uploaded the 722 AI-produced math manuscripts to GitHub on October 7th, only to retract three of them the following day in an update log.
The Verge’s interviews with more than 30 mathematicians revealed descriptions of the work as “astonishing,” “unprecedented,” and “surreal.” The majority agreed that simply understanding what OpenAI has released will take years, and expressed concern that OpenAI might release another batch before the mathematical community has had a chance to digest the current one.
The Challenge of Verification
Currently, no one can definitively answer the question of accuracy, and OpenAI itself admits uncertainty. The README states that the model was tested on approximately 4,000 problems during its evaluation, which OpenAI then grouped and filtered to produce the current catalog of 719 manuscripts, organized into 372 groups. These numbers cannot be directly divided to calculate a success rate, as a single problem might be incorporated into multiple groups after categorization.
The degree of verification varies significantly. Lean, a programming language for computer-assisted proof checking, offers a higher degree of logical certainty for results that have been formalized in Lean. OpenAI’s update log indicates that 300 out of 719, or about 42%, of the main results have formalizations, less than half. The README also concedes that results without formalization may be problematic.
Furthermore, formalization does not guarantee correctness. Interviewees pointed out that the statements verified by computers do not always align with the claims made in the papers, requiring individual verification.
Kevin Buzzard of Imperial College London noted that in his field of algebraic number theory, only about six of the papers immediately appeared noteworthy, and almost none of these had undergone Lean verification. He faces the choice of reading “potentially incorrect garbage” or waiting for others to read and formalize the work before confirming its accuracy.
Some researchers have privately commented that certain papers seem to retrace paths already explored by others, while others mentioned that results originally intended for hundreds of pages have been condensed into a few dozen.
Errors and Retractions Emerge
Errors have already surfaced. The update log reveals a sign error in a paper discussing Weil classes, rendering its argument invalid. Two other papers related to the Hodge conjecture, which built upon this work, were also retracted. An additional 14 papers have been revised. Crucially, none of the results have yet undergone peer review.
High Standards, Despite Presentation
Setting aside the presentation, multiple interviewees believe the overall standard of the work is high. Before the advent of AI, many of these findings would have been sufficient for publication in top journals and would have secured academic positions for their authors. Jared Duker Lichtman of Stanford University even suggested that “dozens” of the results are strong enough to make their authors serious contenders for the Fields Medal, one of mathematics’ highest honors. These include advancements towards the Riemann Hypothesis, special cases of the Hodge Conjecture, and the four-dimensional Kakeya Conjecture. This is his personal assessment.
However, OpenAI’s README notes that the Riemann and Hodge conjecture papers were not produced through the same standardized process as the other results, and does not explain the differences. One version of the Riemann conjecture paper also underwent manual editing.
Scott Armstrong of Sorbonne University emphasized that these are not obscure problems; many are well-known challenges that have eluded mathematicians for decades.
Impact on the Research Community
For some mathematicians, OpenAI’s release of these papers feels like their own research projects have been rendered obsolete overnight. Colva Roney-Dougal of the University of St Andrews stated, “A group of my friends and colleagues just had their research grant applications completely wiped out.” Armstrong is aware of an entire research program by a team that has been nearly eliminated. Tristan Buckmaster of New York University has also heard of three individuals whose projects have been erased.
The most severely affected are PhD students, early-career researchers, and those without tenure, as unsolved problems often form the basis of doctoral dissertations, grant applications, and job searches. Bartosz Naskręcki of Adam Mickiewicz University in Poland commented, “This is a social problem that AI labs are completely ignoring.”
Mixed Feelings on AI in Mathematics
The majority of interviewees do not oppose the use of AI in mathematics, and many use AI tools themselves.
The unease stems from OpenAI’s approach: a high-profile announcement, preliminary drafts, leaving the mathematical community to digest the findings, and then immediately moving on to the next objective. Rumors are already circulating among mathematicians about a potential next wave of releases, and OpenAI has not responded to The Verge’s inquiries about whether another batch is forthcoming. Armstrong remarked, “Something will come along in two months to overshadow this.” Despite being one of the more optimistic mathematicians about AI and using it himself, he admitted, “I have trouble sleeping at night,” with his immediate goal being to complete his current research “before OpenAI beats me to it.”



