OpenAI has published 722 mathematical manuscripts produced by an unreleased internal model, releasing the collection on GitHub. The papers are organised into 372 families of related results. A family can include a principal result, companion arguments, consequences or alternative proofs. They came out of an evaluation in which the model was posed roughly 4,000 problems, according to the repository. The release also includes Lean formalizations of many proofs and abridged summaries of the model's reasoning for 10 result families.
How were the results produced?
OpenAI says most results followed one fixed procedure, averaging about three hours of ChatGPT Pro thinking compute each. Exceptions include work on a zero-free region for the Riemann zeta function and a proof of the Hodge Conjecture for CM abelian varieties. The zeta writeup was human edited for readability.
Why is it a big deal?
Scale is the first reason. Earlier AI maths demonstrations were typically isolated results. This is hundreds of manuscripts released at once, spanning areas the reasoning summaries list, including number theory, operator algebras, statistical physics and fluid-related equations.
The second is cost. Each result took about three hours of Pro-level compute, which points to a very different economics of research output.
The third is verification. Lean lets a computer check a proof, which matters when volume could otherwise overwhelm human referees. OpenAI says it expanded these evaluations after performance on its existing maths benchmarks saturated.
Not every manuscript has a Lean formalization. OpenAI itself says some unformalized results could contain issues and promises quick fixes. Corrections will be recorded as new versions, with earlier versions kept public. The company says it consulted the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to release the work. The claims are OpenAI's and have not been independently verified.
What comes next?
OpenAI says it will fund workshops and conferences on understanding major AI-produced results. It is also working to release the model responsibly, though it has given no date.
