The popular story says AI’s biggest hurdle in mathematics is raw capability. Can a model actually prove something new? That question is mostly settled. The harder problem now is behavior.
According to The Verge AI, OpenAI, Anthropic, and other labs spent the past year announcing breakthroughs on long-standing math problems. Some results went well past what researchers thought current systems could do, including the resolution of one of the famous Millennium Prize problems. You’d expect a victory lap. Instead, The Verge AI reports, labs have moved through the field “with all the grace of a runaway bulldozer,” and results that should have been celebrated set off a backlash.
The myth: better models fix everything
The industry assumes that if the output is correct, nothing else matters. Mathematics runs on different rules.
A proof isn’t finished when someone posts it. It’s finished when the community has checked it, credited it, and fit it into what came before. That process is slow on purpose. It protects the field’s most valuable asset, which is trust that a published result is true and properly attributed.
AI labs have treated the field like a product launch:
- Announce first, verify later. Claims hit social media before mathematicians have had time to check the work.
- Blur the credit. It’s often unclear how much a human steered the model, or whether the “new” result was already in the literature.
- Ignore the norms. Timing, attribution, and coordination with the people who actually maintain these problems get skipped.
We’ve already seen this play out. In 2025, OpenAI and Google DeepMind both reached gold-medal level at the International Mathematical Olympiad, and OpenAI’s early announcement annoyed organizers who’d asked labs to wait. Later that year, OpenAI researchers suggested GPT-5 had “solved” a batch of Erdős problems. It turned out the model had mostly dug up existing solutions. Google DeepMind CEO Demis Hassabis publicly called it “embarrassing.”
Why the backlash matters now
To me, what stands out is that the results are getting more impressive while the reception gets worse. That tells you capability isn’t the bottleneck anymore.
The pattern matters for three reasons:
- Math is the trust benchmark. Labs use math results as public proof of reasoning ability. If mathematicians stop believing the announcements, the marketing value disappears.
- Verification is the real moat. A claimed proof that nobody trusts is worth less than a modest result that’s fully checked. Formal verification tools like Lean, which machine-check every step of a proof, are becoming the currency of credibility.
- It’s a preview for other fields. Medicine, law, and scientific research all have their own verification cultures. How labs treat mathematicians shows how they’ll treat everyone else.
The Verge AI notes that labs say they’re “learning from earlier mistakes.” That’s a promise. Nobody has shown it yet. Labs have every incentive to rush announcements, and those incentives haven’t gone away.
The case for the labs
The other side deserves a hearing too. Some mathematicians welcome the pressure. AI tools are speeding up literature searches, catching errors, and handling tedious case-checking that once ate months of work. A field that’s slow by design can also be slow out of habit.
Both things can be true. AI can be a real gift to mathematics and still be handled badly by the companies building it.
What practitioners should take from this
If you’re building or deploying AI in any expert domain, the math saga is a cheap lesson:
- Make verification part of the work. Ship results with evidence an expert can check, not just a confident claim.
- Respect the domain’s norms. Experts will judge your tool by how you treat their process, not just by your benchmark scores.
- Be precise about credit. Spell out what the human did and what the model did. Vague claims cost you credibility for a long time.
- Hold the press release. A week of outside review costs almost nothing compared with a public retraction.
What comes next
The next big result will test whether the labs’ new humility is real. Watch for formally verified proofs, credit shared with human collaborators, and announcements that come after expert review rather than before it. If labs get that right, AI could become mathematics’ most useful collaborator. If they don’t, they’ll keep producing breakthroughs that the field won’t fully accept.
The Verge AI is tracking the full timeline of these disputes, and its original piece has more details.