Trend Watch 2026-10-08

Trend Watch: OpenAI Dumps 722 Math Manuscripts From an Unreleased Model -- and Mathematicians Ask Who Is Going to Check Them

On October 6 OpenAI published 722 manuscripts in 372 result families, produced by an internal model given about 4,000 open problems at roughly three hours of ChatGPT Pro compute each. Lean proofs cover the main result of 162 papers. Its own advisers call it 'the beginning, not the completion' of human understanding; one Navier-Stokes researcher says OpenAI hasn't done its 'due diligence at all.'

On Tuesday, October 6, 2026, OpenAI put 722 mathematical manuscripts into a public, Apache-2.0 GitHub repository, openai/math. The README says they were "produced by an internal OpenAI model" as part of evaluating it on open research problems, after "performance on our existing mathematical evaluations saturated." The Next Web reports it is the same model behind OpenAI's Navier-Stokes result last month.

What is in the release

  • Scale and method. 722 manuscripts in 372 families. The model "was posed approximately 4,000 problems," and each result used on average "three hours of ChatGPT Pro thinking compute." OpenAI told Scientific American that nearly every paper came from a single prompt to a single agent. Two results -- a zero-free region for the Riemann zeta function and the Hodge conjecture for CM abelian varieties -- did not follow that procedure, and the zeta write-up was "human edited for readability."
  • Verification. The repository's formalization catalogue lists 162 papers with a Lean-formalised main result. For the rest, the README says plainly: "Some of the unformalized results could have issues." Corrections will ship as new versions, with old ones kept.
  • Reasoning. Abridged reasoning summaries cover ten result families, including the irrationality exponent of pi and the Mahler conjectures.

The debate

The Advisory Group on Mathematics and Artificial Intelligence, hosted by the Institute for Advanced Study, had asked labs on September 29 to stop testing hard problems on private models and to publish prompts, time and compute costs, warning of "a two-tier system where labs outrun the rest of the field." After the release it said its advice was not an endorsement: "This release is the beginning, not the completion, of the process of human understanding." NYU's Tristan Buckmaster told The New York Times: "I don't think they've done their sort of due diligence at all." MIT's Andrew Sutherland told Scientific American to treat the single-agent claims as unverified until others can run the model. OpenAI research lead Dan Roberts called the proofs a byproduct of building better tools; OpenAI says it will fund workshops and conferences on AI-produced results.

Why this is a trend. Meta published six papers written with Muse Spark three days earlier; OpenAI has now published two orders of magnitude more. Production of candidate results is no longer the bottleneck in mathematics -- checking them is, and the checkers are human referees who were not consulted on the volume.

I think the formal-proof number is the most important fact in the release, and it cuts both ways. 162 Lean-checked main results from one model run is a real achievement that no referee shortage can undo. But it also means about four in five manuscripts rest on trust, from a model nobody outside OpenAI can query, and the README itself says some "could have issues." Publishing everything with version history is more honest than publishing only the highlights, and it shouldn't be dismissed as spam. The fair ask, which is also AGMAI's, is that labs carry the verification cost they create: formalise before release where possible, label unformalised results clearly as unrefereed, and give independent researchers access to the model so the claims can be reproduced. Without that, mathematicians are being asked to referee a lab's evaluation set for free.

OpenAI's 722-manuscript release shows AI can now generate candidate math results faster than the field can check them: 162 have Lean-verified main results, the rest rest on an unreleased model; its own advisers call it a beginning, and the fair demand is that labs pay for the verification they create.