First look at mathematics manuscripts from an internal frontier model at OpenAI

We’re releasing a broad range of new mathematical results produced by an internal frontier model.

We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.

Read the blog

The public openai/math repository currently contains 722 manuscripts organized into 372 related result families. It includes papers, supporting proof artifacts, Lean formalizations for many results, and 10 abridged summaries of the model’s reasoning.

What’s included

  • Results from an evaluation involving approximately 4,000 problems
  • An average compute expenditure per result equivalent to roughly three hours of ChatGPT Pro thinking with the internal model
  • Computer-checkable Lean proofs for many, but not all, manuscripts
  • Revision and citation protocols, with previous versions remaining accessible

Independent re-check of the quasi-Riemann result (family 003): it passes

Thanks for publishing the Lean code with comparator configs. It made this straightforward to check.

I re-ran the Lean check for OAI.riemannZeta_ne_zero_of_seven_eighths_lt_re (ζ(s) ≠ 0 for Re(s) > 7/8) from the public repo at commit adc7f12:

  • With your challenge and config, Lean FRO’s comparator printed “Lean default kernel accepts the solution” and “Your solution is okay!”, and exited 0.
  • I ran it again against a challenge statement I wrote myself, with the independent nanoda kernel turned on. Both kernels accepted it.
  • #print axioms lists exactly propext, Classical.choice and Quot.sound.

All 2,924 modules in the proof compiled with no errors and no warnings. The report, the raw logs and a container recipe to reproduce it are on GitHub at davegoldblatt/openai-zeta-proof-check (I can’t post links yet as a new forum member).

Caveats: this checks the Lean proof, not the paper. It’s one run on one machine, and an AI agent (Claude Code) carried it out at my request, so independent reproductions would help.

One suggestion: 404 of the 405 comparator configs in the repo leave the second kernel off. Turning it on by default would make these checks stronger.

One of the results in the release is a new matrix multiplication upper bound showing that ω ≤ 2.25 over the complex numbers. I was surprised that the underlying field mattered, so I tried to understand why the results is restricted to ℂ.

Long story short, Codex extended the proof to arbitrary fields, producing a Lean proof in a fork of the repo. The resulting theorem states that ω(F) ≤ 2.25 for every field F, using the same definition of ω.

Lean verification passes, as do Comparator checks. From reading the formalized statement, I am pretty convinced it is correct, and so are the agents.

Cannot include links, but the proof is on GitHub under selanavot/matrix-multiplication-all-fields

The Snaky game solution is very nice. Since this strategy is a finished program, you should post this as a playable game, for example on Android.

Let me suggest: Open AI needs to set up a WIKI so that humans can comment on their giant output of math alleged-problem-solutions. The rate is so high that it is drinking from a firehose. Humans cannot even digest this rate. With a wiki, humans will be able to collaboratively digest, increasing the digestion rate; and also feedback from the wiki might be usable by the AI to enable it, in future, to rewrite its manuscripts with better writing.

“Better” meaning, more useful for human readers.

I am NOT in favor of holding back machine discoveries so some self-interested committee of human anuses can restrict progress. And I do not like the human arXiv and science journal systems as they presently work - I consider them corrupt and intentionally dysfunctional. But I AM in favor of figuring out how to digest the AI flood and how humankind should make use of it, and we need a wiki as part of that.

Let me elaborate on what science needs. (I’ve been saying this for about 30 years, but the science community refused to give a crap about what I’ve been saying, which is because I’m right and they are wrong and they are corrupt.)

The arXiv is corrupt because it only permits those with arXiv-approved employers to publish on arXiv. You can publish utterly false crap there immediately, no questions asked, IF you are a lifelong academic fraud. FOR EXAMPLE, OpenAI in their “irrationality exponent of pi is 2” paper cite work of Nelson A. Carella. Unfortunately OpenAI neglects to mention that all or nearly all work that Nelson A. Carella has ever published, is fraudulent, including this one, and his “proof” of the Riemann hyporthesis, and his “proof” that INTEGER FACTIORNG is in polynomial time, and his “proof” that there are no odd perfect numbers, and on and on. The reason Carella is allowed to keep publishing his garbage is that he is employed by academic institutions who simply do not care that he is a lifelong fraud. If arXiv had commenting, Carella would have been exposed many years ago and hopefully fired from his jobs. I know I personally tried to point out this fraudulence is this and other cases, years ago, and never received any satisfaction.

Meanwhile if somebody not part of the club tries to publish on arXiv, no matter how important and correct, his/her work is rejected. In particular, the 1905 works of Albert Einstein would today be insta-rejected. Ditto Srinivasa Ramanujan. As a result virtually the entire continent of Africa is near 100% excluded from doing science. Ever. Indeed 99% of humanity is excluded. The cost to human progress is enormous.

Now suppose somebody on arXiv publishes some utterly false fraudulent bullshit. Then I am not allowed to point that out. I repeat, I am not allowed even to provide a COMMENT pointing out an error in somebody’s work. No comments or questions of any kind are allowed by arXiv. This is absurd. Repositories like arXiv offer the ability to distribute science far faster and cheaper than in the pre-internet age. Why not also provide far superior refereeing than attainable in the old days? But no: arXiv chose not only to have zero refereeing, but in fact to FORBID even pointing out an error.

If each arXiv paper had a wiki attached allowing commenting by anybody - not just a member of the “club” but anybody - the result would be effectively infinite refereeing, far faster, far cheaper, ultimately far better than ever before in the whole history of science. But no. They chose to forbid that.

And I want a numerical rating system: readers can RATE the papers they read. So there are good papers rated 100 and bad papers rated 0. And it should be tied into searching. So I can search for papers on topic X using keywords Y, but only if ratings>Z. Places like stackoverflow and allrecipes have already solved the problem of providing ratings and ratings-of-raters, etc.

viXra is better than arXiv in the sense they provide a commenting system and they do not have a “club” of people who are allowed to publish, with everybody else arbitrarily excluded. BUT they have no rating system, and no searching-tied-to-ratings capability, and their paper collection sucks. I estimate 80% of vixra papers are garbage, the opposite of arXiv with 20% garbage. I would not care about 80% garbage fraction on vixra if there were a search+ratings capability enabling me to find the few good papers. Then it would be a useful science repository. But since no such capability, that just makes viXra a pile of garbage.

And now with AI providing a flood of papers, you have the opportunity to SET UP such a system as a place to put your AI-generated papers; and you could let everybody else put science papers there too including those partly-generated using AI. As a side effect you will revolutionize science in a way that will vastly benefit humankind, democratize science, and overcome decades of corruption. Also the potential now exists for AIs to act as referees of human papers with comments, questions, alleged refutations, plagiarism finding, lit-search cite-adds, etc etc. That also might cause dramatic improvements both in the human papers and in the AIs.

Please - I’m begging you - set this science-paper system up. Doing so will be far cheaper and far more useful than any conference you set up where a bunch of blowhards pontificate about how to mishandle AI science.

In the OpenAI papers “Universal computation in forced Navier–Stokes flows,” “Universal computation with eventually stationary Navier Stokes forcing,” etc you do the same thing I already did in my 2005 paper that you did not cite (vixra 2503.0157) except better in these senses: (a) OPENAI claims to be fully rigorous and (b) they unlike me did not need a weird rigid container. But worse in the sense (c) my 2005 thing solved halting problem in bounded time, think OpenAI’s takes longer. I have doubts OpenAI really accomplished a+b because as I said in my paper there are fundamental obstacles to rigorization, and if OpenAI really overcame those obstacles, that is bigger news than this.

Anyhow, that is just one example of the useful discussion that could happen on a wiki. Also note, probably the reason OpenAI failed to cite my paper was it was not on arXiv. Proving I am correct that asrXiv is corrupt and dysfunctional since you are a victim of exactly that, as this and the Carella case both prove.

Want more such proof? Another: in your OpenAI paper “One rational hitting point for noncommutative formulas” fails to cite my 2004 work (vixra 2607.0013), chapter 11 of which shows results very related to yours. Are my results better than yours? Worse? Related exactly how? (Mine seem simpler and better explained anyhow.) Those answers are not immediately obvious to me, but in any case it is clear you would have been better off knowing about my work and citing it.

--Warren D. Smith (PhD), 7 Oct 2026.