Stop Telling AI To Think Harder

Two years of prompt advice just expired.

“Think step by step” used to be the one line that upgraded almost any answer out of a model. For two years it worked because older chat models would jump straight to a guess unless you forced them to slow down first. Current reasoning models already do that slowing down internally before they show you a single word, so the old trick barely moves the needle anymore. A Redditor posting as u/Professional-Rest138 broke down exactly why that happened, and handed over the three prompts that replaced it.

Quick start: you’ll learn why the step-by-step trick stopped working, what OpenAI’s own guidance says about it, and three prompts to swap in today. All you need is five minutes and whatever reasoning model you’re already running.

The Old Way vs The New Way

Old way: you type “think step by step,” and the model performs a long, visible chain of reasoning before answering. It feels thorough. It reads like proof of work, even when it isn’t.

New way: you ask for an answer built to be checked, not performed. Same model, same question, a completely different kind of output.

The original poster points to OpenAI’s own documentation, which says telling a reasoning model to think out loud isn’t a reliable upgrade anymore. These models already reason before they respond, so asking them to narrate that process just adds filler on top. “Think harder” fails for a separate reason: reasoning effort is a setting on the model or the API call, not a phrase you slip into a prompt. Typing it harder doesn’t turn a dial that lives outside the text box entirely. The fix isn’t a better instruction, it’s picking the right setting before you even open the chat.

What actually moves the result, per the post, is asking for output you can verify. Here’s the core prompt, reproduced exactly:

Work out [problem] carefully. Give me the answer with a short explanation of the key steps, the assumptions you made, and any calculations I’d need to verify it.

The difference shows up in what comes back. “Think step by step” hands you a performance. This prompt hands you three things you can use to actually catch a mistake: where the model started, what it assumed, and the numbers behind the answer. That’s the whole shift in one line!

Two More Prompts From the Same Family 🔧

The poster shared two companion prompts worth stealing, in the order given in the original post:

1. Surface the gaps before you get an answer:

For [my request], list the details you’re missing and the assumptions you’d otherwise make. Identify the one missing detail most likely to change your answer, and ask me for it.

This one matters most. According to the post, most bad answers aren’t reasoning failures. They’re the model quietly filling a gap you never mentioned and running with it like it’s confirmed fact.

2. Audit an answer after the fact:

Review [this answer]. Separate what’s supported by what I gave you, what’s assumption, and what’s opinion. Don’t label anything verified unless you checked a source.

One commenter on the thread, u/Existing-Meat-8167, flagged the same pattern from the reader side: the model invents a missing detail and runs with it “like it’s gospel.” The gap-surfacing prompt catches that before it happens. The review prompt catches it after, by forcing a line between fact, assumption, and opinion instead of blending all three into one confident paragraph.

Practical Steps

Here’s how to put this to work right now, in order:

  • Swap “think step by step” out of your templates for the verification prompt above.
  • Before a complex request, run the gap-surfacing prompt first. Answer the one question it asks, then send your real prompt.
  • After a high-stakes answer, run the review prompt to split fact from assumption from opinion.
  • Need deeper reasoning on a genuinely hard problem? Change the model or the effort setting directly in your tool or API call. Don’t ask for it in the words, because the phrase alone won’t touch that setting.

Small shift, real payoff: you stop reading performances of thinking and start reading answers you can actually check against your own numbers.

Next step beyond this post: pick one recurring task you run through AI every week, and replace whatever instruction you currently bolt onto it with the verification prompt above. Run it for a week before judging the difference, one good test beats a dozen opinions about prompting.

The original poster has been turning prompts like these into ready-to-run templates and agents. The full thread is worth a look too, the comments surface a few more variations on the gap-surfacing idea that didn’t make it into this post.

“think step by step” doesn’t do what most people think on reasoning models anymore. openai’s own guidance says it’s not a reliable upgrade. here’s what to ask for instead
by u/Professional-Rest138 in PromptEngineering

Scroll to Top