Copilot Flunked a Simple Math Problem

Picture asking a search engine one simple riddle and watching it fumble the exact same question, twice, in a single day. That’s what happened when a Redditor typed “58 in 1961, how old now?” into Copilot Search, hoping to skip a bit of mental math. This person is upfront about being anti-AI, and was only hunting for a website that could do the arithmetic for them. They hit search, got an answer, checked back a few hours later out of curiosity, and got a completely different wrong answer the second time around.

The original poster, u/Honeycrisp042, shared both screenshots on r/PromptEngineering, and the thread lit up over it. Neither answer was close. The math itself is simple: someone who turned 58 in 1961 was born in 1903, which makes them well past 120 years old today. Copilot Search missed it both times, in two different ways, which is somehow funnier than blowing it just once.

There’s a nice bit of irony baked into this story. This Redditor wasn’t trying to test AI or write a callout post. They just wanted a fast answer to a birthday math problem, the same way anyone would type a question into a search bar and move on with their day.

🤔 Why It Matters

This isn’t really about one bad search result on one lazy afternoon. It’s about what happens when we hand off arithmetic, the one job computers have always been great at, to a language model that’s better at sounding confident than being correct.

Search-integrated AI tools are built to answer fast, not necessarily to show their work. When the model underneath treats a math problem like a language problem, it can produce a fluent, well-formatted answer that happens to be wrong. Because the interface looks so put together, most people never think to double-check it.

One commenter on the thread nailed the real takeaway: they still double-check every simple calculation AI hands them, even ones that feel too basic to mess up. That’s the actual lesson buried in this joke of a search result. Not “AI is useless,” but “verify anything with a number attached.”

There’s also the timing detail that makes this case extra interesting. The two wrong answers came from the same tool, on the same question, hours apart, and they weren’t even wrong in the same way. That points to a model generating a fresh, slightly different guess each time rather than pulling from a fixed, checked calculation.

🛠️ How to Stress-Test AI Math Yourself

Want to see how your favorite AI tool handles a question like this? Here’s a quick way to run the same experiment at home.

  1. Pick a two-step math problem: an age calculation, a percentage change, or a unit conversion all work well.
  2. Ask your AI search tool or chatbot directly, with no extra hints or context.
  3. Work out the answer yourself first, or use a plain calculator, so you have ground truth to compare against.
  4. Ask the exact same question again a few hours or a day later, then compare both answers.
  5. If the tool shows its reasoning, check each step individually instead of skimming straight to the final number.

Run this a handful of times across a few different tools and you’ll get a real feel for where each one holds up. Some models are genuinely strong at arithmetic. Others, like this Copilot Search example, still stumble on a question a ten-year-old could solve on paper.

💡 Tips & Tricks

  • Ask for the reasoning, not just the answer. A prompt like “show your work step by step” exposes errors that a bare number hides.
  • Cross-check with a second tool. If two different AI tools agree, that’s a decent sign. If they don’t, trust neither until you verify by hand.
  • Watch date and year math especially closely. Age, tenure, and “years since” questions trip up models more often than plain arithmetic does.
  • Screenshot the weird results. This entire post exists because someone bothered to save the evidence, and that’s exactly how the community catches patterns like this.
  • Don’t assume consistency. As this example shows, the same tool can hand you two different wrong answers hours apart on the identical question.
  • Treat AI search results as a fast first draft, not a final answer, especially anywhere money, dates, or ages are involved.

None of this makes AI search worthless. It just means treating it like a helpful assistant instead of a calculator you’d trust with your taxes. Save the actual number-crunching for a calculator, a spreadsheet, or your own head, and let the AI handle the parts that involve words instead of digits.

⚓ Give It a Shot

Next time an AI tool spits out a number, run the two-second gut check before you trust it! Head over to r/PromptEngineering to see the original screenshots and the rest of the reactions, they’re worth the laugh.

Frequently Asked Questions

Q: Why should I always double-check AI calculations?

Even advanced AI models like Copilot can confidently produce incorrect answers to straightforward math problems. In this case, Copilot not only got the age calculation wrong, but also changed its answer hours later, suggesting inconsistency rather than a one-off glitch. Always verify important calculations with a calculator or second source before relying on the result.

Q: Can AI handle basic arithmetic reliably?

Unfortunately, no. While AI excels at complex reasoning and synthesis, basic math can be surprisingly error-prone, especially with context-dependent questions like “how old would someone born in 1961 be now?” The fact that results can change between queries makes this even trickier to catch.

Q: What’s the best way to verify AI answers?

Use a multi-layer approach: cross-check with another AI tool or a simple calculator, and for important decisions, consult authoritative sources. This catches inconsistencies and gives you confidence before acting on the information, especially for math, dates, or anything time-sensitive.

Broke Copilot Search by accident
by u/Honeycrisp042 in PromptEngineering

Scroll to Top