Math has an AI problem, and mathematicians know it

OpenAI just dropped a set of solutions to longstanding math problems, and it landed like a bombshell inside one of the oldest academic disciplines there is. That’s the story The Verge AI unpacks in a new Decoder interview with Robert Hart, its London-based AI reporter, who spent weeks talking to some of the most accomplished mathematicians alive about what’s happening to their field. What he found is an existential crisis, plain and simple.

The short version: AI went from genuinely terrible at math to convincingly good at a professional level in the span of six months to a year. As Hart puts it in The Verge AI’s report, mathematicians are now compressing five years of upheaval into a very short window.

The strange part

Here’s the twist that makes this so odd. These models are still bad at the math a child does. They struggle with counting, with arithmetic, even with knowing what day it is. Hart notes that ChatGPT still can’t reliably tell time, and both he and Decoder host Nilay Patel share the same half-joking theory: the famous “how many R’s in strawberry” fix is basically hard-coded, not learned.

So how can something fail at counting yet succeed at abstract, research-level math? Because advanced math isn’t really about numbers. Open an academic math paper and you often won’t see any. It’s about reasoning, spotting connections between distant ideas, and applying old methods in new ways. That’s exactly the kind of pattern work these newer models have gotten good at.

What stands out here is that “math” isn’t one thing. Hart compares it to biology, which spans everything from watching animals in a field to mapping biochemistry inside a cell. AI is strong in some corners of math and weak in others. Mathematicians have floated topology as an area the models still fumble. Counting, obviously, remains a mess.

Where the crisis actually lives

So which part has mathematicians rattled? The answer, per The Verge AI, is a bit of both ends.

Nobody fears an AI that’s a mediocre mathematician. The worry sits at the frontier, where models are now producing work on par with strong human researchers. That raises questions the field hasn’t had to face before:

  • If AI can solve problems at this level, what happens to the research careers built on solving them?
  • What’s the point of grants and university programs training new generations of mathematicians, if frontier models just answer the open questions?
  • Can labs take skills learned in math and transfer them to other domains entirely?
  • And the cynical one: is all this math hype just a marketing exercise for AI labs that don’t actually care about the discipline?

That last question deserves weight. Math makes for great demos. It’s rigorous, it’s prestigious, and a clean solution to a famous problem is easy to show off. It’s fair to ask whether the attention is about advancing mathematics or about advertising the model.

Why this matters beyond math

I’d put this next to what already happened in software engineering. We’ve been living through the AI crisis in code for a while now, and the pattern rhymes: a tool goes from useless to unavoidable fast, and the people whose expertise defined the field have to figure out what their job even is anymore.

Math is the next domain to hit that wall, and it may be the most telling one. If a system that can’t count its way through “strawberry” can still match trained researchers on abstract proofs, then capability and reliability aren’t the same thing, and the gap between them is where all the confusion lives. That’s a lesson every field adopting these tools should sit with.

The deeper signal is transferability. If the reasoning that cracks hard math generalizes to other domains, the math story is really a preview. If it doesn’t, then the demos are narrower than they look.

Either way, mathematicians are being forced to answer a question the rest of us will face soon enough: what’s left for humans when the machine is good at the hard part and bad at the easy one? The full conversation with Robert Hart is worth your time over at The Verge AI’s Decoder.

Scroll to Top