Ask a language model a multi-step arithmetic or logic question and demand only the answer, and it will often be wrong. Ask it to write out the steps first, and it is right far more often. That observation, published as chain-of-thought (CoT) prompting by Wei et al. in 2022, is one of the most useful and most misused techniques in prompt engineering.

This article explains why reasoning tokens help, the two main ways to elicit them, how to get a reliably parseable answer out of a reasoning-heavy output, where CoT makes things worse, what changes with models that reason by default, and how to measure whether it is worth its tokens on your task. By the end you should be able to add CoT to a pipeline, prove it helps, and keep it from leaking cost or confusing users.

Advertisement

Why writing the steps helps

A transformer produces each token with a fixed amount of computation: one pass through its layers. If a question needs several dependent steps, such as multiplying, then subtracting, then comparing, a direct answer forces all of them into the single pass that produces the first answer token. Writing the steps out changes that. Each intermediate result becomes text in the context, and every later token can attend to it. The model gets more passes, and a scratchpad to carry state between them.

The evidence is strong for tasks with this multi-step structure. In Wei et al.'s experiments, PaLM 540B went from 17.9% to 56.9% on the GSM8K grade-school maths benchmark when eight worked examples showed reasoning before each answer. Kojima et al. showed the same effect without examples: appending "Let's think step by step" raised text-davinci-002 from 17.7% to 78.7% on MultiArith and from 10.4% to 40.7% on GSM8K. Both papers also found that the gain appeared mainly in large models; small models produced fluent but wrong chains.

The corollary is the most important practical rule: CoT helps when the task has intermediate steps the model would otherwise have to skip. For lookups, classification by surface features, or single-fact questions it adds tokens and latency and rarely adds accuracy.

The two ways to elicit it

Zero-shot CoT is an instruction: think step by step, then answer. It is cheap to write and works across tasks. Few-shot CoT puts a handful of worked examples in the prompt, each showing reasoning and then an answer in exactly the format you want. It costs more prompt tokens but gives you control over the style and length of the reasoning and the format of the answer, which matters when a program reads the output.

DIRECT = (
    "Answer with a number only.\n"
    "Q: {question}\nA:"
)

ZERO_SHOT_COT = (
    "Solve the problem. Think through it step by step, then give the result on a "
    "final line in the form 'Answer: <number>'.\n"
    "Q: {question}"
)

FEW_SHOT_COT = (
    "Q: A shop has 3 boxes of 12 pens and sells 7 pens. How many pens are left?\n"
    "Reasoning: 3 boxes x 12 pens = 36 pens. 36 - 7 = 29.\n"
    "Answer: 29\n\n"
    "Q: A train leaves at 09:40 and the trip takes 2 h 35 min. When does it arrive?\n"
    "Reasoning: 09:40 + 2 h = 11:40. 11:40 + 35 min = 12:15.\n"
    "Answer: 12:15\n\n"
    "Q: {question}\n"
    "Reasoning:"
)

Three details make these templates work. The direct baseline says "number only", so you are comparing like with like. Both CoT variants put the answer on a fixed final line, so it can be parsed. And the few-shot prompt ends with Reasoning:, so the model starts in the right mode. Exemplars should be correct, short and representative of your real inputs; wrong reasoning in an example is copied faithfully.

Advertisement

A worked example

Take the question: "A storage plan charges $0.023 per GB per month; another charges a flat $1.50 per month. For 80 GB, which is cheaper and by how much?" Asked for a direct answer, a model has to multiply, compare and subtract in one step, and a common failure is to answer "the per-GB plan" because $0.023 looks far smaller than $1.50 at a glance, which skips the multiplication entirely.

With CoT the output reads: 80 x 0.023 = 1.84, so plan A costs $1.84; plan B costs $1.50; B is cheaper by $0.34. Answer: B, $0.34. Every step is checkable. If the model makes an arithmetic slip, you can see which step went wrong, which is also how you debug a prompt: read twenty failed chains and group the failures by step (misread the question, wrong operation, arithmetic error, correct reasoning but wrong format). Each group has a different fix.

Getting the answer out reliably

The reasoning is for the model; the answer is for your program. Treat them separately.

Chain-of-thought in production: elicit reasoning, extract the answer, check it, measure itPrompttask + exemplars or"think step by step"Modelreasoning tokens first,then the answerOutputReasoning: 3 x 12 = 36; 36 - 7 = 29Answer: 29parse final lineAnswer extractionstrict format, reject on failureOptional checksvote, verify, tool callFinal answerreasoning kept in logsEvaluation harnesssame labelled set, direct vs CoT: accuracy, tokens per answer, latency, parse-failure rate
The model writes reasoning then answer; a strict parser takes the final answer line; optional checks run before the answer is used; and a harness compares CoT with the direct baseline on the same labelled set.

With free text, require a fixed final line and parse the last match, because models sometimes state a tentative answer mid-chain and revise it. With few-shot prompts written as a Q, Reasoning, Answer pattern, first cut the output at any new Q: line (or pass it as a stop sequence), because models sometimes continue the pattern with an invented question and answer, and the last match would then be the invented one. Treat a missing answer line as a failure to count, not something to guess from the reasoning. With structured output, order matters: put the reasoning field before the answer field.

{
  "reasoning": "Plan A: 80 GB x $0.023 = $1.84. Plan B: flat $1.50. B is cheaper by $0.34.",
  "answer": "B"
}

Models generate fields in order. If answer comes first, it is produced before any reasoning exists and the reasoning becomes a justification written after the fact, which removes the benefit. Schema-constrained output and parsing are covered in structured output.

Do not show the chain to end users by default. It is long, may contain abandoned wrong turns, and can quote internal instructions from the prompt. Log it for debugging, show the answer, and if users need an explanation, generate a short one from the verified result.

Where it fails, and why the chain is not an explanation

CoT has well-documented failure modes, and knowing them keeps you from over-trusting it.

  • Unfaithful reasoning. Turpin et al. (2023) showed that when a prompt contained a biasing feature, such as always putting the correct answer in position A in the examples, models shifted their answers towards it while their written reasoning never mentioned it. The chain is generated text that often, but not always, reflects what drove the answer. Never use it as an audit trail on its own.
  • Error propagation. One wrong early step is carried forward confidently. Longer chains have more places to fail.
  • Hurting easy tasks. On simple classification or extraction, reasoning can talk the model out of an obvious correct answer, or bury the answer in text a parser misses.
  • Cost and latency. A reasoning chain of 200 to 600 tokens where a direct answer takes 5 multiplies output cost and time-to-answer accordingly, and output tokens are usually the expensive ones.
  • Format drift. Long reasoning makes models more likely to forget the answer format. Watch the parse-failure rate as a first-class metric.

Techniques built on CoT

Several techniques extend the basic idea, each owning a different failure mode. Self-consistency samples several chains at non-zero temperature and takes a majority vote on the final answers; Wang et al. reported it lifting PaLM 540B on GSM8K from 56.9% to 74.4%, at several times the cost. Least-to-most prompting first decomposes a hard problem into easier sub-questions and solves them in order, which helps when the problems at test time are harder than the examples. Tree of thoughts explores and scores several partial chains, for search-like problems. Step-back prompting asks for the governing principle before the specific steps.

A simpler and often stronger addition is to hand the fragile steps to tools. If the chain includes arithmetic, have the model write an expression and evaluate it in code; if it needs a fact, retrieve it. CoT then organises the work, and deterministic tools do the parts models get wrong.

Reasoning models change the advice

Many current models are trained to reason before answering, either always or when a reasoning mode is enabled through the provider's API. For these models an explicit "think step by step" instruction is usually redundant, and detailed prescriptions of how to reason can constrain a process the model already does well. Providers' own guidance for their reasoning models generally recommends stating the goal, constraints and output format clearly and leaving the method to the model. Some return the reasoning, some return a summary of it, and some hide it; do not build a parser that depends on reasoning text you might not receive. Check your provider's current documentation for how reasoning is enabled, budgeted and billed, because it changes between model generations.

Two things carry over unchanged: put the answer in a strict, parseable place, and measure. Reasoning tokens are still tokens, and on easy tasks a lower reasoning setting or a non-reasoning model is often as accurate and much cheaper.

Measure it: a small evaluation harness

The only way to know whether CoT is worth it for your task is to run the same labelled set with and without it and compare accuracy, cost and failures side by side.

import re
import statistics

ANSWER_RE = re.compile(r"^Answer:\s*(.+?)\s*$", re.MULTILINE)

def extract(text):
    # Few-shot completions may run on into an invented next "Q:"; cut there first.
    text = text.split("\nQ:")[0]
    # Take the LAST 'Answer:' line; the model may restate earlier guesses.
    found = ANSWER_RE.findall(text)
    return found[-1] if found else None

def evaluate(complete, template, dataset, max_tokens):
    # complete(prompt, max_tokens) -> (text, output_tokens); wraps your model client.
    correct, parse_fail, tokens = 0, 0, []
    for item in dataset:
        text, used = complete(template.format(question=item["q"]), max_tokens)
        tokens.append(used)
        answer = extract(text) if "Answer:" in template or "Reasoning:" in template else text.strip()
        if answer is None:
            parse_fail += 1
        elif answer == item["a"]:
            correct += 1
    n = len(dataset)
    return {"accuracy": correct / n, "parse_fail": parse_fail / n,
            "median_tokens": statistics.median(tokens)}

# for name, t, budget in [("direct", DIRECT, 16), ("zero-shot CoT", ZERO_SHOT_COT, 600),
#                         ("few-shot CoT", FEW_SHOT_COT, 600)]:
#     print(name, evaluate(complete, t, labelled_set, budget))

A hundred to a few hundred labelled examples drawn from real traffic are enough to see a large effect. Look at four numbers per variant: accuracy, parse-failure rate, median output tokens and latency. Then compute cost per correct answer, not cost per call; a variant that costs five times more per call but halves errors can be the cheaper one if errors are expensive. Building a proper evaluation set is covered in prompt evaluation.

A worked comparison makes this concrete. Suppose the direct prompt scores 70% at 8 output tokens and zero-shot CoT scores 88% at 320 output tokens. Output tokens per correct answer go from about 11 to about 364, a 32-fold increase. If each wrong answer costs a support ticket, the 18 points of accuracy almost certainly pay for themselves. If a wrong answer only costs a retry that a user barely notices, they may not. The harness gives you the numbers; the business cost of an error decides.

What to do next

  1. Collect 100 to 300 real inputs with correct answers for one task you run in production.
  2. Run the direct prompt as a baseline and record accuracy, parse failures, tokens and latency.
  3. Add zero-shot CoT with a fixed final answer line; compare. If the task is simple and accuracy does not move, stop and keep the direct prompt.
  4. If it helps, try three to five short, correct exemplars (few-shot CoT) and compare again.
  5. Read twenty failed chains, group failures by step, and fix the biggest group: better examples, a tool for arithmetic, or a clearer question.
  6. For structured output, move the reasoning field before the answer field.
  7. Keep chains out of the user interface and in the logs; alert on parse-failure rate.
  8. If you use a reasoning model, remove step-by-step instructions and test whether a lower reasoning setting keeps accuracy.
Key takeaway: Chain-of-thought prompting improves accuracy on multi-step problems by letting the model write intermediate results it can then build on. Elicit it with an instruction or a few worked examples, put the answer on a strict final line or after the reasoning field, never treat the chain as a faithful explanation, and prove the gain on a labelled set by comparing accuracy, parse failures and cost per correct answer against a direct baseline.