🔍 Read the full analysis: 722 Proofs, One Uncertain Outlook For OpenAI’s AI Mathematics on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
OpenAI published 722 mathematical manuscripts produced by an unnamed, unreleased model, covering 372 families of results drawn from about 4,000 problems. The catalogue includes claims about major open problems, but outside mathematicians have not confirmed them, and OpenAI warns that some unformalized results may contain issues. The key test is whether researchers can verify and use the work, not just whether a claim appears to be proved.
OpenAI published 722 mathematical manuscripts on Monday, presenting results produced by an unnamed, unreleased model across 372 families of related work. The catalogue includes claims concerning several famous open problems, but OpenAI chief executive Sam Altman said the results have not been confirmed by outside mathematicians, leaving their reliability and potential impact unsettled.
According to OpenAI’s post and its GitHub repository, the manuscripts span number theory, geometry, operator algebras, topology, theoretical computer science and mathematical physics. The work was selected from roughly 4,000 problems posed to the model; OpenAI filtered the results for what it considered an appropriate level of significance. The published collection is available under the Apache-2.0 license.
The company says the results are arranged into 372 families, with an average result taking about three hours of ChatGPT Pro thinking compute. Many, but not all, have Lean formalizations, a machine-checkable representation of mathematical proof. The repository README cautions that some unformalized results could have issues. OpenAI also supplied ten abridged reasoning summaries, a small subset of the full collection.
Among the manuscripts are claims involving the Unique Games Conjecture, Hilbert’s tenth problem over the rationals, isomorphism of nonabelian free group factors, a zero-free region for the Riemann zeta function to the right of Re(s) = 11/12, and the Hodge conjecture for CM abelian varieties. These are claims in the published work, not independently established breakthroughs. The source material says the Riemann write-up was edited by humans for readability and identifies it, along with the Hodge result, as an exception to the usual process.
722 proofs, one question: will any of OpenAI’s AI mathematics actually lead anywhere?
An unreleased, unnamed model produced claimed proofs of results that would each define a career. Sam Altman calls them “claims not yet confirmed by outside mathematicians.” The real question isn’t whether it’s impressive. It’s whether answers nobody understands become discoveries anyone can build on.
Same day: Alon, Bloom, Gowers, Litt, Sawin post a digested, human-verified version. The model for success.
Connes rigidity counterexample challenged within a day — constructed groups fail the required condition. Three rival machine “counterexamples” from different labs now circulate.
~10,000 agents, 88 hours, est. ~$22M at retail. Priority dispute; 25 Fields Medalists sign “A Severe Misalignment” — not saying it’s wrong, saying it’s not understood.
Altman now hedges at announcement — a shift from September. Verification has barely started.
Humans extract the technique, write it up, build on it. This is where downstream discovery comes from.
The question is answered; nobody learns anything reusable. Closes a door without opening a field.
The proof breaks, or proves a statement that doesn’t match the conjecture as mathematicians mean it.
The Unique Games Conjecture is the clearest case. Results like the optimality of Goemans–Williamson for Max-Cut are proved assuming UGC. A correct proof converts them all — no understanding required. A zero-free strip for zeta works the same way for prime-distribution results. Free group factors, Kadison, Mahler would redirect whole programmes — but how depends on the method, which means digestion.
Technology. A Navier–Stokes blow-up proof doesn’t change how anyone designs aircraft; engineering turbulence models never depended on the answer. Near-term consequences are mathematical, not industrial. “AI will cure cancer next” skips several steps.
“Verification abundance, adjudication scarcity” — making proof-checking cheap doesn’t reduce the burden of deciding what’s true and what matters. 722 manuscripts land on a review system built for a trickle, filtered by a selection nobody outside OpenAI made.
Humans re-deriving results, like Alon–Gowers et al. in May
Other people’s work building on these manuscripts
How many unformalized results survive expert checking
Do the Lean statements match the real conjectures?
Do any survive peer review?
Some of it, yes — where a literature is waiting (UGC), a correct proof pays off immediately; where a proof carries a new technique humans digest, it can open a field. Most of it, probably not on its own: at 722 manuscripts with 10 reasoning summaries, the Four Colour pattern is the likely default unless mathematicians are funded and given time. And some will be wrong — OpenAI says so itself. It’s an industry pattern, not one company’s: the forced-Euler result came from an Anthropic researcher, and rival machine-generated Connes “counterexamples” circulate from different labs. The proofs arrived this week. The discoveries, if they come, will arrive at the speed of human understanding.
Verification Will Set the Value
The release matters because its headline results, if correct, could affect active research well beyond the individual problems. The Unique Games Conjecture, for example, underpins many results in theoretical computer science about the limits of approximation algorithms. A verified proof could prompt researchers to revisit conclusions that rely on the conjecture. But until specialists check the argument and its exact assumptions, those consequences remain conditional.
In mathematics, a correct answer is not always the same as a useful discovery. The source material contrasts results that researchers can digest and build on with proofs that settle a question but offer little reusable insight. OpenAI’s May release concerning the Erdős unit-distance conjecture is described as an example of machine output subsequently converted into a human-verified result by mathematicians. That process—not the number of manuscripts—is the more meaningful measure of whether this collection changes the field.
There is also a question of research attention. Hundreds of manuscripts can put a substantial review burden on mathematicians, especially when the results range across specialties and some lack formal verification. The material’s central uncertainty is therefore practical as well as mathematical: whether researchers can identify sound arguments, understand their methods and find worthwhile ideas among them.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A Mixed Record of Math Claims
This is described in the source material as OpenAI’s fourth major mathematics release of the year. Its earlier releases offer both a positive example and reasons for caution. In May, the company’s model produced a counterexample to the Erdős unit-distance conjecture. Five mathematicians posted what they called a digested, human-verified version on the same day, illustrating how machine output can be checked and made useful to the research community.
OpenAI’s August collection, called “Ten Advances,” had a more disputed outcome. The source says a claimed counterexample to Connes’s rigidity conjecture was challenged within a day because the groups constructed did not meet a condition required by the conjecture. It also reports that similar machine-generated counterexamples from separate labs have circulated. This history makes careful checking of definitions and hypotheses especially relevant to the new release.
In September, OpenAI announced a Lean-formalized result concerning finite-time blow-up for the Navier–Stokes equations, a Millennium Prize problem. The source describes a dispute over priority and reports that 25 Fields Medalists signed a declaration criticizing the use of famous problems as AI benchmarks without human understanding. Their objection, as presented in the material, was about the purpose and practice of mathematical research—not a finding that the Navier–Stokes proof was false. These episodes underline the distinction between producing a proof candidate and establishing its validity and value.
formal proof checkers for mathematicians
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Proof Status Remains Open
The material does not report independent verification of the 722 manuscripts, nor does it identify which claims have been checked by specialists. It is also unclear how many of the results have complete Lean formalizations, what standards OpenAI used to rank the roughly 4,000 problems, or how much review each manuscript will require. OpenAI controlled the selection of work judged significant, so the published catalogue should not be read as an independently chosen or externally ranked set.
The source lists headline claims but provides no outside assessment of their proofs. It does not establish whether any manuscript proves a result in precisely the form mathematicians have posed it, whether errors or missing assumptions will be found, or whether correct arguments will yield reusable methods. Until those questions are answered, the claims should remain attributed to the manuscripts and their publisher.
AI-powered mathematical problem solver
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Independent Review Comes Next
The immediate next step is for specialists in the relevant fields to inspect the manuscripts, test their reasoning and check formalizations where available. Some results may need to be rewritten or translated into arguments that researchers can assess; others may be challenged, corrected or set aside. The source material does not give a timetable for that review or identify a formal independent evaluation process.
For readers, the clearest milestones will be published assessments from mathematicians, corrections to the manuscripts, and evidence that researchers can use the methods in follow-up work. The catalogue’s longer-term significance will depend on those outcomes: verification first, then possible uptake. Until then, OpenAI’s release is evidence that its model generated a large body of mathematical claims, not confirmation that those claims resolve the problems named in them.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI publish?
OpenAI published 722 mathematical manuscripts grouped into 372 families. According to its post and repository, the work was selected from roughly 4,000 problems posed to an unnamed model that has not been released.
Have mathematicians confirmed the results?
The source material says the results have not been confirmed by outside mathematicians. The repository also warns that some unformalized results could have issues, so the claims should not be treated as established findings.
What is Lean formalization?
Lean is a proof assistant used to encode mathematical arguments in a form that can be checked by software. The source says many, but not all, of the results have Lean formalizations; formalization does not by itself explain a result’s wider value to researchers.
Why is the Unique Games Conjecture claim attracting attention?
The conjecture is connected to many results in theoretical computer science about the limits of approximation algorithms. If a proof is verified, researchers may need to revisit work that depends on it; that consequence is conditional on the manuscript being correct and proving the conjecture as stated.
When will the manuscripts be judged?
No review timetable is provided in the source material. The next meaningful developments would be independent assessments, corrections or follow-up research from mathematicians working in the relevant areas.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
