Models

Chain-of-Thought Prompting: Does It Still Help Reasoning Models?

Chain-of-thought prompting used to be the trick that unlocked better answers. With reasoning models doing that step internally now, here's whether asking one to 'think step by step' still does anything.

Elena Vasquez

Former ML Researcher, Industry Analysis Lead

Published 7 min read
Close-up of hands engaging with wooden brain teasers on a table indoors.
In this story 5 sections

Chain-of-thought prompting, telling a model to 'think step by step' before answering, largely stops mattering once you're using a dedicated reasoning model. Reasoning models already generate an internal chain of reasoning by default, regardless of how the prompt is worded, so the trick that meaningfully boosted older models now does little to nothing on top of what the model was already going to do.

That's a real shift, and it's easy to miss if you learned prompting techniques a couple of years ago and haven't revisited them. Chain-of-thought prompting was arguably the single most influential prompting technique of the pre-reasoning-model era. At Emergent Wire, we still see teams pasting 'think step by step' into every prompt out of habit, on models where it was never going to change the output.

This explainer covers where chain-of-thought prompting came from, what it actually did to a model's accuracy, why reasoning models changed the equation, and when the technique still earns its place in a prompt today. It's written for anyone writing prompts for production use who wants to know which habits are still worth keeping.

  • Chain-of-thought prompting was introduced in 2022 and measurably improved multi-step reasoning accuracy.
  • It worked by making a model's intermediate steps explicit instead of jumping straight to an answer.
  • Reasoning models generate that same kind of internal chain automatically, without being asked.
  • Adding 'think step by step' to a reasoning model's prompt typically adds little to no benefit.
  • The technique still helps on non-reasoning, general-purpose models for genuinely multi-step problems.
  • Forcing reasoning steps on very simple questions can occasionally hurt rather than help.

Where Chain-of-Thought Prompting Came From

Chain-of-thought prompting comes from a 2022 paper by researchers at Google, 'Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.' The core finding was straightforward: when a model was prompted with a few worked examples that showed intermediate reasoning steps, rather than just a question and a final answer, its accuracy on math word problems and multi-step logic tasks jumped substantially compared to prompting it for a direct answer.

The mechanism made intuitive sense. A model generating tokens one at a time has no separate scratchpad to work things out before committing to an answer. Writing out the reasoning steps as part of the output effectively gave the model somewhere to do that work, step by step, in a form it could then build its final answer on top of. Emergent Wire's read on why this mattered so much for that model generation is that it exposed a real limitation in how those models processed multi-step problems, not just a quirk of prompt phrasing.

Algebra equations with symbols on a chalkboard in a brightly lit classroom.

What Reasoning Models Changed

Reasoning models, the category that includes OpenAI's o1 and o3 models and open-weight models like DeepSeek-R1, are trained specifically to generate an extended internal chain of reasoning before producing a final answer, whether or not the prompt asks for it. OpenAI, the lab that publishes its own system cards describing how these models are trained, has described this internal reasoning stage as a core part of the model's training objective, not a response to prompt phrasing. That reasoning process happens by default, as part of how the model was trained to respond, not as a behavior triggered by a particular prompt phrase.

That's the key difference from the 2022 finding. Chain-of-thought prompting worked because it changed the model's behavior from 'answer directly' to 'show your work first.' A reasoning model doesn't have a 'direct answer' mode to switch away from in the first place; showing its work is the default path baked into training, not an optional add-on a prompt can turn on. Our explainer on how reasoning models actually work goes deeper into that training process and why it produces a different kind of output than a standard chat model.

One architectural detail worth understanding here: many reasoning models route the internal chain of thought through a separate, often hidden, reasoning phase before the visible answer, and some use a mixture-of-experts design underneath to keep that longer reasoning process computationally efficient. Our piece on how mixture-of-experts architectures work covers how a model can run a longer reasoning process without a proportional jump in cost for every query.

Testing the Claim: Does It Still Move the Needle?

Testing this directly is simple: run the same benchmark set through a reasoning model twice, once with a plain prompt and once with an explicit 'think step by step' instruction added, and compare accuracy. Across the math and logic benchmark categories we've run this comparison on at Emergent Wire, the gap between the two conditions is consistently small, often within normal run-to-run variance rather than a clear, repeatable improvement.

That result lines up with what the original chain-of-thought researchers themselves later found in follow-up work: the technique's benefit scales with how much a model needs the explicit nudge to reason step by step, and a model already trained to do that by default has little left to gain from being told to do it again. The original chain-of-thought paper, published via arXiv by its Google Research authors, actually predicted a version of this: the technique's gains were largest on the models that reasoned worst by default, and smallest on the strongest models tested, well before reasoning-specific training existed at all.

Testing the Claim: Does It Still Move the Needle?
Model TypeChain-of-Thought Prompt BenefitWhy
Older non-reasoning models (pre-2023)LargeNo default step-by-step behavior to draw on
Modern general-purpose chat modelsSmall to moderateSome implicit reasoning already trained in
Dedicated reasoning models (o1-class, R1-class)Minimal to noneInternal reasoning chain generated by default
Side view of crop anonymous learner reading papers with test while studying alone
Hand placing piece in transparent puzzle, symbolizing leisure and problem-solving.

When Chain-of-Thought Prompting Still Helps

The technique isn't dead, it's just narrower than it used to be. On general-purpose chat models that aren't trained as dedicated reasoners, explicitly asking for step-by-step reasoning still tends to help on genuinely multi-step problems, since those models still default to a shorter, more direct answer path unless prompted otherwise.

We've also found, testing capability differences at Emergent Wire, that chain-of-thought-style prompting still has a real use even on reasoning models: not to trigger reasoning, but to shape its structure. Asking a reasoning model to organize its answer around a specific framework, or to double-check a specific type of error before finalizing, can still change the output meaningfully, even though the blanket 'think step by step' instruction mostly doesn't. Our look at whether AI models can actually do real math covers a related case where the structure of a prompt, not just whether reasoning happens, still measurably changes accuracy.

The practical rule we'd give a team writing prompts today: on a reasoning model, drop the generic 'think step by step' instruction and save the prompt-engineering effort for structuring what the model should check or prioritize. On a standard chat model without built-in reasoning, the older technique is still doing real work and is worth keeping.

The Bottom Line

Chain-of-thought prompting hasn't stopped working so much as it's been absorbed into how reasoning models are trained in the first place. For a reasoning model, telling it to think step by step is mostly redundant. For a standard chat model, it's still a genuinely useful habit. Knowing which kind of model you're prompting is what actually determines whether the technique is worth the extra words.

Emergent Wire covers how AI models are built and trained, for readers who want to know what's actually happening under a model's output, not just what it produces.

What is chain-of-thought prompting?
Chain-of-thought prompting is asking a model to write out its intermediate reasoning steps before giving a final answer, often by adding a phrase like 'think step by step' to the prompt. Researchers introduced it in 2022 and found it meaningfully improved accuracy on math and logic tasks for that generation of models.
Do reasoning models like o1 or DeepSeek-R1 still need chain-of-thought prompts?
Not really. Reasoning models are trained to generate an internal chain of thought automatically before answering, regardless of how the prompt is phrased. Adding 'think step by step' to a reasoning model's prompt typically produces little to no measurable improvement, since the model is already doing that by default.
Does chain-of-thought prompting still help with regular, non-reasoning models?
Yes, for models that don't generate an internal reasoning chain on their own, explicitly prompting for step-by-step reasoning still tends to improve accuracy on multi-step problems, though the effect is smaller on today's stronger base models than it was on the models tested in 2022.
Can chain-of-thought prompting ever make a reasoning model worse?
It can, in a few narrow cases. Forcing an unusual reasoning format can occasionally shorten or distort a reasoning model's own internal process, and on very simple questions, added reasoning steps can introduce an error that a direct answer wouldn't have made.