OpenAI Publishes 722 AI-Written Math Manuscripts, Including a Quasi-Riemann Proof
TL;DR
At 21:47 UTC on October 6, OpenAI created a public repo, openai/math, holding 722 manuscripts grouped into 372 result families, all written by an unnamed, unreleased internal model. The claims include a zero-free half-plane Re(s) > 7/8 for the Riemann zeta function and every Dirichlet L-function, a proof of the Unique Games Conjecture, a matrix multiplication exponent of at most 9/4, and an integer multiplication algorithm below n log n. 235 of the 372 families ship Lean formalizations; the other 137 do not, and the README says some of those "could have issues." OpenAI says the average result cost about three hours of ChatGPT Pro thinking compute. A week earlier, the Advisory Group on Mathematics and AI, the independent panel OpenAI itself convened in September, had asked labs to "stop testing advanced mathematical problems on proprietary models." OpenAI cited the group's advice in its release post and did it anyway.
From ten results to 722 manuscripts in nine weeks
This is the third escalation in OpenAI's math program. On August 1 it published ten results with 548,207 lines of Lean in the openai/ten-proofs repo. On September 21 it announced the advisory group alongside a claim that its model had resolved more than 100 open problems. Today's drop is 372 families and 722 papers, released under Apache 2.0 with a BibTeX block per manuscript and a promise to preserve every public version.
The release post is short and unusually careful. It describes the author as "an internal frontier model," says OpenAI has "been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study," and says it is "working to responsibly release the model that produced these results." No model name, no date.
The headline claims
The manuscript map is 640 KB of abstracts. These are the families that will be fought over first:
- Quasi-Riemann hypothesis (family 003). Every Dirichlet L-function, including zeta, is zero-free for Re(s) > 7/8, and Landau-Siegel zeros are excluded uniformly. Lean-formalized. A companion paper gives an independent proof at 11/12, and the README notes that writeup was "human edited for readability." The Riemann hypothesis itself needs 1/2, so this is not that, but no fixed strip to the left of Re(s) = 1 had ever been cleared before.
- Unique Games Conjecture (102). A deterministic polynomial-time reduction from 3SAT to Unique Games, plus direct proofs that beating the Goemans-Williamson ratio for Max-Cut and factor two for Vertex Cover are NP-hard. Lean-formalized.
- Matrix multiplication exponent (107). ω ≤ 9/4 over the complex numbers. The field has spent a decade shaving decimals off 2.37. Lean-formalized.
- Integer multiplication below n log n (109). A deterministic multitape Turing machine algorithm in O(n (log n)^(1-κ)) with κ = 2^-182, which would disprove the Schönhage-Strassen optimality conjecture that Harvey and van der Hoeven's n log n result made plausible. No Lean.
- Kakeya (074). The maximal conjecture in three dimensions and full Hausdorff dimension in four. No Lean.
- Hodge conjecture for CM abelian varieties (032). One of two results the README says did not follow the fixed procedure. No Lean.
- Hilbert's tenth problem over the rationals (004). Negative: no algorithm decides whether an integer polynomial has a rational zero. No Lean.
- Disproofs. Hadwiger's conjecture fails even for fractional coloring (157), Kaplansky's zero-divisor conjecture gets a counterexample (196, Lean), and Barnette's conjecture is proved (180, Lean). The plane cannot be five-colored (158, Lean), leaving six or seven.
If the zero-free half-plane is hard to picture: think of the critical strip as a hallway running from Re(s) = 0 to Re(s) = 1. The Riemann hypothesis says every zero stands exactly on the center line. Until now, the best anyone could do was rope off a lane along the right wall that gets thinner and thinner the farther down the hallway you walk. The new claim ropes off the entire right eighth of the hallway, at every height, forever.
On the integer multiplication result, the exponent κ = 2^-182 means the speedup over n log n is a factor of (log n)^(1/2^182). For any input that fits in the observable universe, that factor rounds to one. It is a galactic algorithm in the purest sense: a conjecture killed by a knife too thin to measure.
What is checked and what is not
The repo's own numbers are the ones that matter for anyone deciding how much to trust this. 235 of the 372 families have a Lean scope page under lean/docs/. The formalization catalogue lists 162 papers with a formalized main result. The README states it plainly: "This collection includes results at different stages of verification. Not all have accompanying Lean formalizations. Some of the unformalized results could have issues."
A Lean file is only as good as the statement it proves, and the scope pages are where OpenAI draws that line. The page for the quasi-Riemann family says the formalization covers the 7/8 bound for zeta, every Dirichlet L-function, and finite-order Hecke L-functions over Q(√-3), while "the paper's later applications are not included." To check that the formal statement matches the English one, each family has challenge files for the Comparator tool, and the instructions are three commands:
lake update
lake exe cache get
lake env comparator ComparatorChallenges/QuasiRiemannHypothesis.json
That is the part worth copying regardless of what you think of the math. The artifact you can trust is not the PDF and not the blog post. It is the machine-checked statement plus a tool that confirms the statement says what the abstract says.
Three hours of Pro, 4,000 problems, three days of output
The README describes one fixed procedure for "the vast majority" of results: an unreleased internal model, roughly 4,000 problems posed, and "on average, each result used three hours of ChatGPT Pro thinking compute." Roughly 9 percent of posed problems became a result family, though OpenAI frames that as output after "requiring an appropriate level of significance," not a solve rate. The two exceptions to the fixed procedure are the Riemann zero-free region and the Hodge result.
The manuscript directory names carry dates, and they cluster hard. 460 of the 722 papers are dated September 23 to 25, and another 112 are dated October 5, the day before release. Compare that with the Navier-Stokes result OpenAI announced in September, which took roughly 10,000 concurrent agents over 88 hours. Scientific American reports that this batch came from single agents working from single prompts. The README also explains why the problems were open ones at all: "We expanded these evaluations after performance on our existing mathematical evaluations saturated." The unsolved problems of mathematics are now an eval set.
What the mathematicians asked for, and what they got
The Advisory Group on Mathematics and Artificial Intelligence is nine mathematicians, among them Timothy Gowers, Martin Hairer, Edward Witten, Ravi Vakil, and Camillo De Lellis, hosted at the Institute for Advanced Study and unpaid by any lab. On September 29 it published guidelines informed by more than 600 survey responses. The key line: "we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models." It also asked that for each result "the AI lab should make public the name of the model, the prompts used, a (summarized) chain of thought, the time taken, and the estimated cost of computation," that results be deposited in scholarly repositories that "should not be controlled by any AI lab," and that labs fund "the development of human understanding of the AI mathematical output that they release."
OpenAI's post says it has "drawn on their advice and public recommendations," and the funding commitment is real: "a series of workshops, conferences, and special programs around the understanding of major results produced by AI." The rest is partial. Ten reasoning summaries cover ten of 372 families. Compute is an average in Pro-hours, not a per-result figure in dollars or GPU time. The repo lives under the openai GitHub organization while the company is "continuing to explore other community-hosted alternatives." And the central request, to stop running open problems on a model nobody outside can query, was declined by the act of publishing.
The reaction
Scientific American's headline calls it "a field already in shock." MIT's Andrew Sutherland told the magazine: "Until and unless they release the model and people can replicate their results, I think you should treat any claims about one-shotting problems with a single agent as unverified. We should ask for receipts." Toronto's Daniel Litt took the other side: "If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us. To me, it's going to be a good thing for mathematics." The magazine also reports Terence Tao has criticized the pace as "insane." The Hacker News thread passed 500 points and 420 comments within five hours. The comment that stuck: one mathematician who had spent 24 years on Barnette's conjecture said learning it was solved felt "like hearing an ex-girlfriend died suddenly in a car crash."
Why builders should care
Three things transfer out of pure mathematics into any agent pipeline you run.
First, the verification split is the real headline. OpenAI could have published 722 PDFs and a press release. Instead it published 235 Lean scopes, 162 formalized main results, and a Comparator harness, and it wrote down which claims lack them. If your own agents ship work without an equivalent of the lean/docs page, the parts without one are exactly the parts the README warns "could have issues."
Second, the disclosure gap is a template for what not to do. An average of three Pro-hours tells you nothing about the tail. The advisory group asked for per-result prompts, time, and cost. Those are the same fields you would want in any agent run log, and OpenAI, with every incentive to look good, still did not ship them.
Third, the model is coming, eventually. "Working to responsibly release the model that produced these results" is a commitment with no name and no date attached. When it lands, the openai/math repo is the eval set it was trained against, and 137 families are still waiting for a Lean file.
Key Takeaways
- OpenAI published 722 manuscripts in 372 result families to the public openai/math repo on October 6 under Apache 2.0, all written by an unnamed, unreleased internal model.
- Headline claims include a zero-free half-plane Re(s) > 7/8 for zeta and all Dirichlet L-functions, the Unique Games Conjecture, ω ≤ 9/4 for matrix multiplication, and integer multiplication below n log n.
- 235 of 372 families have Lean formalizations and 162 papers have a formalized main result; Kakeya, Hilbert's tenth over Q, Hodge for CM abelian varieties, and the sub-n-log-n multiplication claim have none.
- OpenAI says the average result used about three hours of ChatGPT Pro thinking compute across roughly 4,000 posed problems; 460 of the 722 papers are dated September 23 to 25.
- The IAS-hosted advisory group asked labs on September 29 to stop testing open problems on proprietary models and to publish model names, prompts, and per-result cost; OpenAI cited the group and met one of its asks in full.
- OpenAI says it is "working to responsibly release the model," with no name or date.
Sources: openai/math repository (README, CONTENTS.md, lean/), OpenAI: Sharing AI progress in mathematics, AGMAI: Responsible Release of AI-Generated Mathematics, Scientific American, TechCrunch, Hacker News