Mathematicians Aren’t Sold on OpenAI’s 719 Proofs

OpenAI released hundreds of claimed solutions to some of the hardest open problems in mathematics this week. The mathematicians it consulted beforehand say it didn’t meet their standards. According to TechCrunch AI, the lab said it had worked with an advisory group of leading mathematicians to avoid a repeat of the backlash over its last big math claim. It still fell short on the group’s most important principle: making sure humans actually understand the results.

The short version

  • What happened: OpenAI published 719 manuscripts with claimed solutions to open math problems and said it had followed outside guidelines.
  • The problem: It only partly followed those guidelines. Just 10 of the 719 manuscripts came with the model’s chain of thought, meaning its step-by-step reasoning.
  • New doubts: A paper from researchers at Cambridge and King’s College London found at least two mismatches between the written proof and the machine-checked code for OpenAI’s claimed solution to a Navier-Stokes problem. That’s one of the million-dollar Millennium Prize problems.
  • Why it matters: The idea that formal verification makes AI proofs automatically trustworthy is now under real pressure.

Who set the rules

The Advisory Group on Mathematics and Artificial Intelligence (AGMAI) is hosted by the Institute for Advanced Study in Princeton. Its nine members are prominent researchers at institutions around the world. The group published guidelines for frontier labs working on math problems at the end of September.

Its first request was blunt: “stop testing advanced mathematical problems on proprietary models.” OpenAI’s release says outright that it’s doing exactly that, testing its proprietary models on open research problems.

To be fair, OpenAI did follow some of the guidelines. It published results quickly and included some information on how the models reached their conclusions. Beyond that, the record is patchy:

  • Chain of thought: included for only 10 of 719 manuscripts.
  • Formalization: AGMAI wants proofs that humans don’t understand to be formalized, meaning written as code that a computer can check. TechCrunch cites a 42% figure here, which still leaves a big share of the release unchecked by machine.
  • Linking metadata: AGMAI asked for “machine-readable metadata correlating the natural language and formal artifacts.” OpenAI didn’t provide it.
  • Funding human follow-up: The group suggested OpenAI help pay the mathematicians who will have to make sense of these results. There’s no sign that’s happening.

AGMAI’s public statement was careful. It said “it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully.” The group didn’t respond when TechCrunch asked for a fuller assessment.

Lost in translation

The technical issue here is worth understanding. When an AI model tackles a hard problem, it first writes its proof in ordinary language. Then it translates that proof into Lean, a programming language that checks a proof by compiling it like code. If the code compiles, the proof should be valid.

The catch is that the translation step can go wrong. If the model writes Lean code that doesn’t match its written argument, the code can compile while proving something slightly different from what the paper claims. The Cambridge and King’s College paper found at least two of these mismatches in OpenAI’s Navier-Stokes solution. That’s not proof that either version is wrong. It does show that a model checking its own homework isn’t enough.

The authors’ conclusion is direct: AI-generated Lean proofs “should not prima facie be trusted without the same peer review process and scrutiny that other proofs are subjected to.”

The human understanding gap

This is the core of the dispute. When human mathematicians make a discovery, they stand behind it. They write papers, give talks and answer questions. That process spreads understanding, turns up techniques that work on other problems, and gets new ideas into applied fields.

Terence Tao, a prominent critic of OpenAI’s approach, put it sharply on social media: “Problems are being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is ‘solved’, and do not understand the AI output well enough to answer questions on the result, give talks, or otherwise interact with the rest of the field.”

Harvard mathematician Melanie Wood told TechCrunch that with AI-released proofs, “there is not human understanding of them at the point of release, and now the work begins.”

What to watch

What stands out to me is the gap between solving a problem and actually contributing to the field. OpenAI can produce proofs at a scale no research group can match. But each of those proofs creates review work that human mathematicians have to do, often unpaid.

If you build on AI-generated math or code, here’s what to take from this:

  • Formal verification isn’t the finish line. Check that the formal statement actually matches the claim.
  • Expect pushback on volume releases. Releasing 719 papers at once looks less like a breakthrough and more like a review backlog.
  • Norms are still forming. The AGMAI guidelines are only weeks old, and this is their first big test.

The next few months of peer review will show how many of these proofs hold up and whether OpenAI starts treating human understanding as part of the job. TechCrunch AI has the full report.

Scroll to Top