OpenAI has released a collection of new mathematical results produced by an internal frontier model, alongside formal proofs and disclosure about the process used to obtain them. Published on 6 October, the work is intended to support scrutiny by mathematicians while the company develops better standards for presenting AI-assisted scientific results.
The results are available in a GitHub repository with protocols for revisions and citations. OpenAI says it is exploring community-hosted alternatives that meet guidance from the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study.
Lean formalisation supports mechanical checking
Many proofs are being shared in Lean, a programming language and proof assistant that allows mathematical arguments to be checked by computer. OpenAI plans to add more formalisation as it becomes available. Machine checking can catch logical gaps in the encoded proof, although experts must still assess whether the formal statement matches the intended mathematical claim.
Formal proofs also improve reproducibility. Researchers can inspect definitions, dependencies and verification results rather than relying entirely on prose. The repository model makes corrections visible, provided revisions are carefully documented and earlier versions remain identifiable.
Process details accompany the results
OpenAI is publishing ten summaries of the model’s reasoning, estimates of compute expressed in ChatGPT Pro usage and statistics on attempted problems. The company says the average result used compute comparable to roughly three hours of ChatGPT Pro thinking.
These numbers provide useful context but should not be treated as a full measure of research cost. Human problem selection, verification, failed directions, formalisation and editorial work also contribute. Future releases would benefit from consistent reporting of those inputs and the criteria used to decide which results were shared.
Independent advice shaped the release
OpenAI consulted the Institute for Advanced Study group about best practices and says it incorporated the group’s public recommendations. This responds to concerns that major AI-generated claims can arrive without adequate exposition, attribution or channels for correction.
The company acknowledges that citations, mathematical writing and presentation need further improvement. That is important because a valid proof can still be difficult for the community to understand, connect with earlier work or evaluate for significance.
Scientific credit needs durable rules
A GitHub workflow can record model outputs, human revisions and formal files, but scholarly recognition also requires clear authorship and contribution statements. Researchers need to know who selected the problem, verified novelty, wrote the exposition and accepted responsibility for errors.
Citation protocols should prevent later corrections from obscuring which version another paper used. Persistent identifiers, archived releases and transparent change logs would make the results easier to reference than a repository that changes continuously.
Community scrutiny is the real test
OpenAI says it will fund workshops, conferences and special programs around major AI-produced results. Those forums can help experts examine proofs, compare methods and identify where model reasoning contributed something genuinely new rather than rediscovering known work.
The most credible outcome will be independent verification and useful follow-on research. Headline difficulty or compute does not establish mathematical importance. Specialists must judge originality, elegance, generality and whether the ideas illuminate related problems.
Towards responsible scientific models
OpenAI is working to release the model that produced the results, while continuing evaluations in mathematics and other sciences. Wider access could let researchers test capability and failure modes directly, but release decisions need to consider reliability, misuse and the resources required for responsible use.
This publication is valuable less as a final standard than as an experiment in scientific disclosure. Formal proofs, attempt statistics and external advice create a stronger basis for assessment. The next step is consistent community governance that can survive beyond any single company repository.
Mathematics is unusually suited to this process because formal systems can verify logical correctness. Other sciences require empirical data, experimental replication and judgement about measurement. Lessons from this release should therefore be adapted carefully rather than treated as a universal template for AI-generated discoveries.
Novelty checking remains a human challenge
A proof can be correct yet already known in another form. Establishing novelty requires broad literature knowledge, conversations with specialists and careful citation, particularly when terminology differs across fields. OpenAI’s commitment to improve references will be central to avoiding accidental rediscovery or incomplete attribution.
Researchers should also be able to separate the model’s contribution from subsequent human repair. Publishing failed approaches, revision history and verification notes can show whether the system originated the core idea or produced a promising sketch that experts substantially rebuilt.
Workshops funded by OpenAI should include sceptical reviewers and researchers outside the company’s immediate collaborators. Independent replication can identify hidden assumptions and improve norms for future releases. A healthy research process must allow criticism of both the mathematics and the disclosure method.
It should also preserve enough detailed intermediate material for later scholars to reconstruct how each result developed.