Chain-of-thought prompting asks a model to write intermediate reasoning before its answer. Why that helps, and how to elicit it with an instruction or worked examples, is covered in chain-of-thought prompting in depth; read that first if the idea is new. This article starts where a prototype ends: you have seen CoT raise accuracy on your task in a notebook, and now it must run thousands of times a day inside a product.
At that point new questions dominate. Where does the reasoning go in a structured response? What does it cost in tokens, latency and money? Should every request get it? Who is allowed to see the chain? What changes with models that reason natively? And can you keep the accuracy while paying for a smaller model? Each has a concrete engineering answer.
CoT as a pipeline component
In production, CoT is one stage in a pipeline, not a phrase in a prompt. A request is routed to a direct or a reasoning path; the reasoning call returns a structured object; a parser validates it; failures escalate; the user sees only the answer; and the reasoning is stored, redacted, for debugging and evaluation. Treating the chain as an internal artefact with a schema, a budget and a retention policy is what separates a robust deployment from a demo.
The output contract: reasoning before answer
Most products need structured output: a label, a JSON object, a function call. The naive approach asks for JSON with an answer field and hopes the model reasoned silently. That throws away the benefit, because a model generates tokens left to right and each token is conditioned only on the tokens before it. If the answer field comes first, it is produced before any reasoning exists in the context; a reasoning field written afterwards can only rationalise a decision already made.
So put the reasoning field first in the schema and the answer after it, and say so in the prompt. Constrain the reasoning's shape as well as its position: numbered steps, a maximum count, and a list of what each step must establish. Constrained steps are shorter, easier to audit, and less likely to wander into irrelevant detail.
import json
SCHEMA_HINT = {
"type": "object",
"properties": {
"reasoning": {"type": "string", "description": "numbered steps, at most 8"},
"answer": {"type": "string", "enum": ["eligible", "not_eligible", "needs_human"]},
"policy_clause": {"type": "string"}
},
"required": ["reasoning", "answer", "policy_clause"] # reasoning FIRST
}
PROMPT = '''You decide refund eligibility under the policy below.
Work through the request in short numbered steps inside "reasoning":
1) dates and amounts, 2) which clause applies, 3) any exception.
Then give "answer" and the clause number. Output JSON only.
POLICY:
{policy}
REQUEST:
{request}'''
def decide(policy, request, complete):
raw = complete(PROMPT.format(policy=policy, request=request),
max_tokens=600, json_schema=SCHEMA_HINT)
out = json.loads(raw) # raises on truncation
if out["answer"] not in SCHEMA_HINT["properties"]["answer"]["enum"]:
raise ValueError("bad label")
if out["policy_clause"] not in policy: # cheap grounding check
raise ValueError("cited clause not in policy")
return outThree details matter. The max_tokens limit is set from measured chain lengths plus headroom, because a truncated chain produces invalid JSON, and the parser must treat that as a failure, not a partial answer. The answer is an enum, so the parser can reject anything else. And the cited clause is checked against the policy text: a cheap grounding check that catches confident reasoning built on a clause that does not exist.
An alternative is the two-call pattern: one call produces free-form reasoning, a second, cheaper call extracts the structured answer from it. It costs an extra round trip but helps when strict JSON mode degrades reasoning quality on your model, which you should test rather than assume.
Budgeting tokens and latency
Reasoning tokens are output tokens, and output tokens are the expensive and slow ones. Output is generated one token per decoding step, so latency grows roughly linearly with chain length, and most providers price output tokens several times higher than input tokens.
Work through numbers for your own deployment. Suppose a direct answer averages 30 output tokens and a CoT answer averages 350, and your model decodes at about 60 tokens per second. The direct path adds about half a second of generation; the CoT path adds nearly six seconds. At one million requests a month, the CoT path generates 320 million extra output tokens. Multiply by your provider's output price and you have the monthly cost of reasoning on every request, before any retries.
That arithmetic leads to three levers. Shorten the chain with step limits and terse-step instructions, then measure whether accuracy holds. Stream the response so the user waits for the answer, not the reasoning, when the interface allows. And, most powerfully, do not reason when you do not need to.
Routing: reason only when it pays
CoT helps tasks with dependent intermediate steps and does little for lookups or surface classification. Within a single task, difficulty also varies: most refund requests are obviously in or out of policy, and a minority involve exceptions and date arithmetic. Spending reasoning tokens on the easy majority buys almost nothing.
Build an escalation ladder. Each rung is more expensive and more capable, and a request climbs only when the rung below fails validation or reports low confidence.
def answer(task, x, complete, classify_difficulty, vote):
# Rung 0: deterministic code for anything that is really arithmetic or lookup.
if task.has_tool and task.tool_applies(x):
return task.tool(x)
# Rung 1: direct answer for easy items (route learned from your eval set).
if classify_difficulty(task, x) == "easy":
out = task.direct(x, complete)
if task.validate(out):
return out
# Rung 2: one CoT call with the structured contract.
out = task.cot(x, complete)
if task.validate(out) and out.get("confidence", 1.0) >= task.threshold:
return out
# Rung 3: self-consistency over k samples, or a stronger model, or a human.
outs = [task.cot(x, complete, temperature=0.7) for _ in range(task.k)]
winner, share = vote([o["answer"] for o in outs if task.validate(o)])
return {"answer": winner} if share >= 0.6 else {"answer": "needs_human"}The difficulty classifier can be simple: rules on input length or features, a small model, or the direct answer's own validation result. Train and tune it on your labelled evaluation set, measuring accuracy and cost per rung. The top rung uses self-consistency, sampling several chains and voting, which multiplies cost by the sample count, so reserve it for the hard tail. Rung 0 matters too: anything that is truly arithmetic, date maths or a database lookup should be done by code or a tool call, as in ReAct, rather than reasoned about in text, because models make arithmetic slips that code never does.
Handling the chain: who sees it, where it goes
Treat the reasoning as internal by default. Raw chains are long, sometimes wrong in intermediate steps even when the answer is right, and occasionally contain things you would not show a user: quoted system instructions, speculation about the user, or text injected by a malicious document that the model echoed while reasoning. Showing them invites users to argue with step four.
If users need an explanation, generate a short, separate justification from the final answer and its cited evidence, or show the cited policy clause. Store full chains in a trace store with the same access controls and retention limits as the inputs, because they will contain any personal data the input contained. And remember that a chain is not a faithful account of how the model reached its answer; use it for debugging and evaluation, never as an audit record of the decision on its own.
Native reasoning models change the recipe
Many current models reason internally before answering, often controlled by a provider-specific effort or budget setting, and some return only a summary of their reasoning rather than the full chain. With these models, adding "think step by step" is usually redundant and can be counterproductive, because you are asking for a second, visible chain on top of the internal one and paying for both.
The production principles carry over unchanged. Route by difficulty, but now the lever is the effort setting or the choice between a reasoning and a non-reasoning model. Budget latency, because internal reasoning tokens are generated and usually billed even when you do not see them. Keep a structured answer contract, and validate outputs exactly as before. Re-run your evaluation when you switch, because prompts tuned to elicit visible reasoning from an older model may simply be noise to a reasoning model. Check your provider's documentation for the exact controls rather than assuming a parameter exists.
Distilling CoT into a smaller model
If one task runs at high volume, you can trade a one-off training cost for lower serving cost. Use a large model with CoT as a teacher to produce rationales and answers, keep only the items where the answer is verified, and fine-tune a smaller student on them. Hsieh et al.'s Distilling Step-by-Step (2023) trained small models on both labels and teacher rationales as separate tasks and reported that they matched or beat much larger few-shot models using less training data than standard fine-tuning, which matches what practitioners see on narrow tasks.
# Build a rationale-augmented training set from a teacher, keep only verified items.
records = []
for item in unlabeled_or_labeled_pool:
for _ in range(4): # several tries per item
out = teacher_cot(item.input, temperature=0.7)
if item.label is None or out["answer"] == item.label:
records.append({"input": item.input,
"rationale": out["reasoning"],
"answer": out["answer"]})
break
# Multi-task targets: the student learns to produce the answer and, separately,
# the rationale, so at inference time you can ask for the answer alone.
train = ([{"prompt": "[answer] " + r["input"], "target": r["answer"]} for r in records] +
[{"prompt": "[explain] " + r["input"], "target": r["rationale"]} for r in records])The multi-task framing lets the student answer without generating a rationale at inference time, so you keep most of the accuracy gain without paying for reasoning tokens. The filter is essential: training on rationales that led to wrong answers teaches the student to reason wrongly with confidence. Evaluate the student on held-out data and on the hard tail specifically, and keep the escalation ladder so its failures still reach the teacher. SLM distillation covers the training side in more depth.
Worked example: refund eligibility at scale
A support platform handles 400,000 refund requests a month. The policy has twelve clauses, three of which involve date windows and partial-refund arithmetic. A direct-answer prompt is fast but misreads the date-window cases; the CoT contract above fixes most of them but adds about 300 output tokens and five seconds to every request.
The team labels 1,000 historical requests and runs each rung on them. Their routing rule sends a request to CoT only when it mentions a date window, a partial refund or a previous refund, about a quarter of traffic. Date arithmetic moves to a small function the prompt receives as a precomputed fact, such as days since delivery equals 34, which removes the most common reasoning slip outright. Requests whose cited clause fails the grounding check, or whose three-sample vote splits, go to a human queue.
The resulting design reasons on roughly 100,000 requests a month instead of 400,000, cuts average added latency by about three quarters, and puts the human team's attention on the few hundred genuinely ambiguous cases. After three months of logged teacher outputs, the team distils a small model for the routing step itself. None of these numbers transfer to your task; the method does: label, measure each rung, route, verify, escalate.
Failure modes
- Answer before reasoning in the schema, so the chain rationalises instead of reasoning.
- Truncated chains from a tight
max_tokens, producing invalid JSON that a lenient parser half-accepts. - Runaway length: chains that grow over time as prompts accrete instructions, silently raising cost and latency.
- Arithmetic in prose where a tool or precomputed fact would be exact.
- Leaked chains shown to users or logged without redaction.
- Format contamination: reasoning text that bleeds into the answer field, breaking downstream parsers.
- Stale prompts on a new model, especially explicit CoT instructions sent to a model that already reasons internally.
What to do next
- Label a few hundred real inputs and measure direct and CoT accuracy, latency and output tokens on them.
- Put the reasoning field first in your schema, constrain its steps, and validate every response strictly.
- Move arithmetic, dates and lookups into code or tools, and pass results in as facts.
- Add a router and an escalation ladder, and track cost and accuracy per rung on a dashboard.
- Keep chains internal, redact them in a trace store, and give users a short, separately generated justification if needed.
- When you adopt a reasoning model, drop explicit CoT instructions and re-run the evaluation.
- For high-volume narrow tasks, collect verified teacher rationales and evaluate a distilled student.