OpenAI This week, hundreds of solutions claiming to solve the world's most difficult mathematical problems were released. The cutting-edge laboratory stated that they consulted a group of advisors consisting of top mathematicians in order to avoid the controversy that arose last time when their model solved a long-standing unsolved mathematical problem.
However, OpenAI does not meet these standards, especially in the areas where mathematicians emphasize the need for human understanding of mathematical results. This concern is even more pronounced after the publication of a new paper that points out a gap between the natural language description and the formal expression of a problem that the OpenAI model is purported to have solved, a problem valued at one million US dollars.
The Mathematics and Artificial Intelligence Advisory Group ( Advisory Group on Mathematics and Artificial Intelligence , abbreviated as AGMAI ) is organized by the Institute for Advanced Study at Princeton University, and its members include 9 renowned researchers from institutions around the world.
At the end of September, the organization released guiding principles for cutting-edge laboratories to solve mathematical problems. In a statement regarding the latest batch of proofs, AGMAI stated: "In the end, whether our recommendations have been successfully followed will still need to be assessed by the mathematical community."
However, the organization's primary request is to "stop testing advanced mathematical problems on proprietary models." Yet, the release of OpenAI clearly states that they are using open research problems in the field of mathematics to evaluate their own proprietary models.
When TechCrunch sought more detailed evaluations regarding the latest publications by OpenAI, the advisory group did not respond. It is apparent that the laboratory followed some of these principles, including releasing results as soon as possible and providing information on how the model reached its conclusions. However, not all were met: among the 719 manuscripts, only 10 included a disclosure of the model's thought process.
For papers that are difficult for people to understand, mathematicians suggest that the proofs should be formalized; however, in the proofs published by OpenAI, 42% have not yet gone through this process.
Overall, it is still unclear whether OpenAI fulfills the responsibility of ensuring human understanding when publishing proofs, in accordance with the principle of AGMAI. AGMAI suggests that OpenAI should help to support the work of human mathematicians, as these mathematicians will be necessary to make the laboratory's solutions meaningful in any practical sense.
"The problems are being solved autonomously by those AI prompt operators who have no interest in the field itself; once their initial goals are 'solved,' they no longer care, and their understanding of the AI outputs is not sufficient to answer questions, prepare reports, or interact with others in the field in any other way," noted the renowned mathematician Terence Tao on social media after the release. Tao had previously criticized the practices of OpenAI.
This issue is reflected in a paper published this week. Written by mathematicians from the University of Cambridge and King's College London, the paper questions the way leading laboratories handle such challenges.
When the AI models solve mathematical problems, they first generate a segment of “natural language” explanation, and then attempt to write the results in Lean. This is a programming language, and in theory, its accuracy can be confirmed by compiling the proofs into code.
However, there may be issues with the way the model translates natural language proofs into code; this paper records at least two inconsistencies between the natural language proof and the code in the solution provided for a problem originating from the Navier-Stokes equations, denoted as OpenAI. The Navier-Stokes equations are used to describe the complex behavior of fluids.
These inconsistencies do not necessarily negate the solutions of any particular version, but they do raise questions: is it possible to formalize the answers solely relying on models without human involvement? This is one of the reasons why AGMAI requires OpenAI to “include machine-readable metadata to correspond between natural language and formalized outputs,” yet this cutting-edge laboratory did not do so in these releases.
"As there are instances of mistranslations – as emphasized in this paper – OpenAI and other automatically formalized proofs, such as Lean, should not in principle be trusted without undergoing the same peer review process and scrutiny as other proofs," concluded the author of the paper "Lost in Translation."
Mathematicians emphasize that when humans discover new results, they take responsibility for those findings and interact with a broader community through papers, lectures, and seminars. This process deepens the understanding of the solutions, leads to strategies that can be applied to solve other problems, and enables the new knowledge to be put into practical use in various fields.
Harvard University mathematics professor Melanie Wood told TechCrunch that when a model is prompted to solve a difficult problem and comes up with a solution, “there is no human understanding at the time of publication, and now the real work has just begun.”












