Reasoning Models Explained: What "Thinking" Actually Means
Reasoning models spend extra computing time working through a problem step by step before answering, which changes both what they get right and how much they cost.
In this story 6 sections
A reasoning model is an AI system trained to generate an internal chain of intermediate steps before producing a final answer, trading extra computing time and cost for better performance on complex, multi-step problems.
When people say a model is "thinking," they don’t mean anything like human deliberation. They mean the model is generating a sequence of intermediate text — working through the problem the way you might on scratch paper — before it commits to a final answer.
This piece explains how reasoning models actually differ from standard models, where the extra step genuinely helps, and where it’s mostly overhead. It’s written for anyone trying to figure out when a slower, more expensive reasoning model is worth using over a faster standard one.
The distinction matters more than marketing copy usually lets on, since the two model types are genuinely suited to different kinds of tasks, not just different price points of the same thing.
How Reasoning Models Are Actually Trained
A standard language model is trained to predict the next word that looks like a good continuation of the text so far. A reasoning model adds a further training stage on top of that, using reinforcement learning, where the model is rewarded specifically for reaching a correct final answer after working through a problem.
Crucially, that reward is checked against the final answer, not the intermediate reasoning steps themselves. This means a reasoning model can arrive at a genuinely useful problem-solving pattern through trial and error, without a human ever writing out the "correct" way to think through that specific type of problem.
The training process rewards the model for exploring an idea, backtracking when it notices an error, and trying a different approach — all inside the same response, before committing to an answer. This looks superficially similar to a person working through a hard problem out loud, even though the underlying mechanism is entirely different.
Because the reward only checks the final answer, the model isn’t taught a single "correct" reasoning style the way a textbook might teach one. Two reasoning models trained on the same problems can end up with visibly different internal habits, some more verbose, some more terse, while still converging on similarly accurate final answers.
Where Extra Reasoning Steps Actually Help
The clearest gains show up on math, formal logic, and multi-step planning problems, where getting the final answer right genuinely depends on getting every intermediate step right. A reasoning model that catches its own arithmetic mistake mid-chain and corrects it will often land on the right answer where a standard model, committed to its first instinct, won’t.
Coding tasks that require holding multiple constraints in mind — matching a function signature, handling an edge case, and following a style guide simultaneously — also benefit, since the reasoning step gives the model room to check its own work before outputting final code. This overlaps with what we’ve seen covered in our report on AI coding agents, where the agents that reason through a change before applying it make noticeably fewer breaking errors.
We’ve tested this directly at Emergent Wire on multi-step word problems and found the gap between reasoning and standard models widest exactly where a person would also need to slow down and double-check their own work.
Where the Extra Step Is Mostly Wasted Cost
Simple factual lookups — "what year did X happen," "define this term" — see almost no benefit from reasoning, since there’s no multi-step chain to get right in the first place. Running a reasoning model on this kind of request mostly just adds latency and cost for an answer a standard model would have gotten right anyway.
Creative writing tasks are a similar mismatch. Reasoning training optimizes for a single correct final answer, which doesn’t map cleanly onto tasks where there isn’t one right output, and some reasoning models produce noticeably flatter creative writing as a side effect of that training.
This is why most serious products now route requests to different model types rather than using a reasoning model for everything by default — treating it as a specialized tool for the problems that actually need it, not a strict upgrade.
Why Longer Reasoning Chains Don’t Guarantee Better Answers
This is an active area of research, and benchmark work from the National Institute of Standards and Technology (NIST) has specifically flagged reasoning-chain length as a metric that can be gamed independently of actual answer quality, complicating how labs measure genuine progress on this front. Separate analysis from Epoch AI has tracked how reasoning-model inference costs scale with chain length, finding steeply diminishing accuracy returns well before the chains stop growing.
A longer chain of reasoning steps sounds like it should always help, but several failure modes push back on that assumption:
- A model can reason in circles, revisiting the same wrong idea repeatedly without new information.
- Longer chains cost more and take longer, with diminishing returns past a certain length.
- Some providers let a reasoning model reason indefinitely, which can produce a worse answer than stopping earlier would have.
- Evaluation benchmarks sometimes reward length itself rather than reasoning quality, skewing training incentives.
Choosing Between a Reasoning Model and a Standard One
The practical rule: if a task genuinely requires multiple dependent steps where an early mistake breaks the final answer, reach for a reasoning model. If the task is a direct lookup, a short creative request, or something a standard model already handles reliably, the reasoning model is added cost without added value.
This decision pairs closely with the tradeoffs in our piece on model distillation, since reasoning and distillation solve opposite problems — one adds compute for depth, the other removes it for speed — and a mature product often routes between both depending on the request. It also connects to how models handle long context windows, since a reasoning chain effectively consumes context space the same way retrieved documents do.
For anyone testing this themselves, the fastest way to know is to run a batch of real requests through both model types side by side and compare not just accuracy but cost per correct answer, which is usually the number that actually determines which one belongs in production, rather than accuracy alone.
The Bottom Line
Reasoning models aren’t a strictly better version of standard models — they’re a different tool, tuned for a different shape of problem, at a real cost in speed and price. The "thinking" framing is useful shorthand but shouldn’t be mistaken for anything resembling how a person actually reasons through a problem.
Knowing which category your task actually falls into is the single most useful thing to figure out before choosing between the two, and it’s worth testing both directly rather than assuming the more expensive option is automatically the better one.
Emergent Wire covers AI models, capabilities, and the industry building them for readers who want the real story behind the demos.