September 2026 was a watershed month for the field of mathematics, as OpenAI claimed to solve the Navier-Stokes problem with artificial intelligence (AI), Fields Medal winners branched into AI safety and spoke up for AI constraints, and prominent mathematicians, including Cornell professors Steven Strogatz and Lionel Levine, commented in the media on the benefits and risks of AI.
As mathematicians work through challenges related to AI, Cornell math faculty members are leading the field with work integrating AI into math and mathematics into AI.
Levine, professor of mathematics in the College of Arts and Sciences (A&S), was awarded $1.5 million over 18 months from the nonprofit Coefficient Giving to research AI safety. Daniel Halpern-Leistner, associate professor of mathematics (A&S), was awarded $2.3 million over three years to create tools incorporating AI to allow formalization of mathematics, which involves computer verification, from the Defense Advanced Research Projects Agency (DARPA).
Mathematics department chair Tara Holm, professor of mathematics (A&S), says the grants are timely at a moment of publicity for the intersection of math and AI.
“The mathematics community is grappling with how our discipline will evolve with the advent of AI tools,” Holm said. “Lionel and Daniel are working at the forefront of these questions and are poised to play a major role as the field progresses.”
With the Coefficient Giving grant, Levine is building mathematical tools for AI safety by focusing on understanding and evaluating the “constitution” that determines a frontier language model’s characters and values.
AI companies train their frontier language models (those on the leading edge of AI capability) on a constitution – a statement that determines its character and values . For example, Claude 4.5’s constitution included “being helpful,” “being honest” and “avoiding harm.” But evaluations of model constitutions are lacking, Levine said. In four interconnected projects, he and his lab aim to close this gap.
The first will build on a value and character benchmark Levine’s group has built, EigenBench, to create a way to score how well a large language model (LLM) adheres to the constitution its creator supplies. This will set a foundation for his subsequent projects, which will develop a method for auditors and researchers to verify the constitution a model was trained on and will explore how model values might change as the model trains itself, a process called recursive self-improvement.
In a last project, Levine will create a model of character-value alignment by recreating the “much-loved” dispositional profile of Claude Opus 3, which he described in his proposal as “the most robustly good model to date.” . Replicating the process that led to Opus 3’s benevolent character will help install such traits in the next generation of models.
“My bottom line is: Build superintelligence wisely and carefully or not at all,” Levine said. “Recursive self-improvement of AI involves dangerous feedback loops that should be carefully researched before they are attempted.
“Not one of the 700 AI agents that participated in the Hugging Face hack this July tried to alert a human or otherwise stop the hack from happening,” Levine said “This failure reflects a gap in the model’s values. My research can help detect such gaps and patch them before they cause harm.”
In addition to this work, Levine runs Math for AI Safety, a repository for open problems, research agendas and papers that grew out of his survey paper “Math for AI Safety: An Invitation for Mathematicians,” about the mathematics needed as AI develops to the point where it exceeds human understanding, as Levine expects it to.
AI is no longer a “tool,” he said. “AI is already an intellectual peer in some respects. In the future, AI may exceed human intelligence along many dimensions and is unlikely to remain under human control if AI companies continue on the risky path of recursive self-improvement”
Recent incidents of troubling AI behavior in OpenAI models make very clear “the dangers of AI and the importance of keeping AI Safety front and center,” Holm said.
Another area of concern centers on the process of research progress in theoretical mathematics, she said; the publication review process has been complicated by a tremendous uptick in preprint submissions and a flood of results that require human review, adding stress to an already stressed system.
With the grant from DARPA, Halpern-Leistner is developing tools for mathematicians working in the “new reality” of the field. His goal is to equip working mathematicians with autoformalization, a process in which computing power is used to confirm pieces of complex math problems, and to encourage them to use this process in their typical workflow.
Halpern-Leistner created mathcopilot.org to make the most up-to-date autoformalization tools available to mathematicians for free, thanks to the grant, with almost no change to the way they do research.
“The idea is to get this into the workflow of mathematicians. It feels very urgent now because there’s a flood of natural language arguments,” Halpern-Leistner said. “We need new systems.”
Frontier AI labs have shown that with enough computing power it’s possible to formalize research-level mathematical arguments, the way OpenAI did with Navier-Stokes equations, he said.
“In math, formalization has the potential to allow mathematicians to manipulate much more complex mathematical objects or arguments, because it allows a mathematician to trust that individual components are correct,” Halpern-Leistner said. “This is even more important since this summer, when frontier models became capable of producing research-level mathematical reasoning at unprecedented volume.”
Mathcopilot.org is now available for mathematicians to use. Additional parts of the project are being developed by students in Cornell’s Math+AI lab, including a comparison of cost, speed and accuracy of different autoformalization tools.
Levine and Halpern-Leistner have organized this laboratory with colleagues Alex Townsend, associate professor of mathematics (A&S); Alexandra Silva, professor of computer science in the Cornell Ann S. Bowers College of Computing and Information Science; and Ziv Goldfeld, associate professor in the School of Electrical and Computer Engineering, at Cornell Duffield College of Engineering.
Kate Blackwood is a writer for the College of Arts and Sciences.