Chain-of-thought prompting works best when you show the model a few worked examples. The trouble is that someone has to write those examples, they have to match the problem in front of you, and a fixed set rarely does. Analogical prompting removes the hand-written examples. Before solving, the model is told to recall a few related problems it already knows, write out how each was solved, and only then tackle the real one.
The technique comes from Yasunaga and colleagues at Google DeepMind and Stanford, in the paper Large Language Models as Analogical Reasoners (ICLR 2024). They report that it beats zero-shot chain of thought, and matches or beats manual few-shot chain of thought, on maths (GSM8K, MATH), code generation (Codeforces) and BIG-Bench reasoning tasks. This article explains why it works, gives templates and code you can run against any model, works an example end to end, and spends most of its time on where it breaks and how to tell whether it is helping your workload.
The idea from first principles
A language model solves a problem better when the context already contains a solved problem with the same structure. That is the whole mechanism behind few-shot prompting: the exemplar fixes the format, and more importantly it shows a method. The exemplar does not have to come from a human. A large model has seen thousands of textbook problems. If you ask it to bring the relevant ones into the context, it writes its own demonstrations, and they fit the target because they were chosen with the target in view.
Humans do the same. A student facing a new geometry problem thinks of a similar one they have solved and reuses the method. The paper's name for this is analogical reasoning, and the prompt simply makes that step explicit and visible.
There are two variants. In self-generated exemplars the instruction is, in the paper's words, "Recall three relevant and distinct problems. For each problem, describe it and explain the solution." In self-generated knowledge + exemplars the model first gets a step that says "Identify core concepts in the problem and provide a tutorial", and then recalls exemplars. The authors found the knowledge step helps most on harder tasks such as code generation, where the exemplars alone tend to copy surface details and miss the underlying algorithm.
Two details in that instruction matter. Relevant pulls the exemplars towards the same method. Distinct stops the model writing three near-copies of one problem, which adds tokens and no information. The paper reports that performance is stable once there are three or more exemplars, with the best results at three or five.
The prompt
Below is a template in the style of the paper, with two additions that pay off in production: a fixed answer marker so the output can be parsed, and an explicit instruction not to copy numbers from the exemplars. Keep the problem first, so the recall step is conditioned on it.
Your task is to tackle the problem below.
# Problem
{problem}
# Instructions
## Relevant problems
Recall {k} relevant and distinct problems. For each problem, describe it
and explain the solution. Label them "Exemplar 1", "Exemplar 2" and so on.
## Solve the initial problem
Now solve the initial problem step by step, using whatever method from the
exemplars actually applies. Do not copy numbers from the exemplars.
End with a single line: "ANSWER: <final answer>".For the knowledge variant, insert a section before the exemplars: "Identify core concepts in the problem and provide a tutorial." Put the tutorial first: the paper found that generating knowledge before exemplars works better, because naming the core concept first pulls the exemplars towards the same underlying method.
Worked example: a counting problem
Target: In how many arrangements of the letters of BANANA are no two A's adjacent? A zero-shot model often counts all arrangements (6!/(3!·2!) = 60) and then subtracts the adjacent cases by hand, which is slow and easy to get wrong.
With the analogical prompt, a capable model typically recalls problems like these:
- Exemplar 1. Arrangements of MISSISSIPPI. Method: multinomial coefficient, 11!/(4!·4!·2!) for repeated letters.
- Exemplar 2. Seat 4 boys and 3 girls in a row with no two girls together. Method: place the boys first (4! ways), which creates 5 gaps, then choose and order gaps for the girls (5·4·3).
- Exemplar 3. Choose 3 non-adjacent numbers from 1 to 10. Method: the gap argument again, giving C(8,3).
The second and third exemplars carry the method that matters: arrange the other items, then put the restricted items into the gaps. Applied to the target: arrange B, N, N in 3!/2! = 3 ways. That leaves 4 gaps (both ends and the two between letters). The three A's are identical, so choose 3 of the 4 gaps: C(4,3) = 4. Total 3 × 4 = 12. The first exemplar contributes the repeated-letter division that the 3!/2! step needs.
This shows both why the technique works and how it can fail. If the model had recalled three repeated-letter problems and no gap problem, every exemplar would be relevant on the surface and none would carry the method. That is why the instruction asks for distinct problems.
Implementation: one call, two calls, and voting
The paper's setup is a single call: exemplars and solution in one response. It is cheap and simple, but you cannot inspect the exemplars before they influence the answer. A two-call version first asks only for exemplars, filters them, then asks for the solution with the survivors in context. Filtering can be cheap, for example re-running any arithmetic in an exemplar, or executing exemplar code against its own stated test. Both versions combine with self-consistency, which samples several runs and takes a majority vote.
import re
from collections import Counter
def llm(prompt: str, temperature: float = 0.0, max_tokens: int = 2048) -> str:
"""Wrap whichever model client you use; nothing below depends on the vendor."""
raise NotImplementedError
ANSWER_RE = re.compile(r"^ANSWER:\s*(.+)$", re.MULTILINE)
def final_answer(out: str) -> dict:
answers = ANSWER_RE.findall(out)
if not answers:
return {"answer": None, "raw": out, "error": "no answer marker"}
# Use the LAST marker: exemplars sometimes end with their own "ANSWER:" line.
return {"answer": answers[-1].strip(), "raw": out}
def analogical(problem: str, k: int = 3, temperature: float = 0.0) -> dict:
return final_answer(llm(ANA_TEMPLATE.format(problem=problem, k=k), temperature))
def two_pass(problem: str, k: int = 3) -> dict:
"""Generate exemplars, filter them, then solve with only the survivors."""
ex = llm(RECALL_ONLY.format(problem=problem, k=k), temperature=0.7)
kept = [e for e in split_exemplars(ex) if passes_checks(e)] # your validators
if not kept:
return analogical(problem, k) # fall back to the one-call form
out = llm(SOLVE_WITH.format(problem=problem, exemplars="\n\n".join(kept)))
return {**final_answer(out), "kept": len(kept)}
def vote(problem: str, n: int = 5) -> str:
"""Self-consistency on top: sample n analogical runs, majority-vote the answers."""
answers = []
for _ in range(n):
r = analogical(problem, temperature=0.7) # voting needs diverse samples
if r["answer"]:
answers.append(normalise(r["answer"]))
return Counter(answers).most_common(1)[0][0] if answers else NoneTwo parsing rules prevent most production bugs. Take the last answer marker, because exemplars often end with an answer line of their own. And treat a missing marker as an error, not an empty answer, so it shows up in your metrics instead of being scored as wrong.
How it compares with the alternatives
| Technique | Exemplars come from | Strength | Weakness |
|---|---|---|---|
| Zero-shot chain of thought | None | Cheapest; no setup | No method shown; weak on unfamiliar structure |
| Manual few-shot chain of thought | Humans, fixed set | Reliable, auditable exemplars | Same exemplars for every input; costly to write per task |
| Retrieved (dynamic) few-shot | A labelled pool, by similarity | Real, verified exemplars fitted per input | Needs a pool and an index; fails outside its coverage |
| Analogical prompting | The model itself, per input | Fitted per input with no labelled data | Exemplars can be wrong; extra output tokens; weak on small models |
| Step-back prompting | The model, as an abstract principle | Good when a governing principle exists | Less help when the method is procedural, not conceptual |
The practical takeaway: if you have a good labelled pool, dynamic few-shot retrieval usually beats self-generation, because its exemplars are real and checked. Analogical prompting earns its place when there is no pool, the input distribution is broad, or you need a strong baseline quickly. A hybrid works well: retrieve when similarity is high, and fall back to self-generated exemplars when the nearest neighbour is too far away. Step-back prompting is a close relative. It asks for the principle, while analogical prompting asks for worked cases, and on some tasks the two can share one prompt.
Failure modes
- Wrong exemplars. The model writes a confident, incorrect solution for a recalled problem, and the target inherits the error. This is the defining risk. The two-call form with cheap validators is the main defence.
- Surface analogies. The exemplars share vocabulary but not structure, so the model applies the wrong method. Watch for exemplars that repeat the target's nouns. Asking for distinct problems, or adding the knowledge step, helps.
- Answer leakage. The model copies a number from an exemplar into the final answer. The "do not copy numbers" instruction and the last-marker parse reduce this, and an evaluation set with distractor numbers exposes it.
- Small models. The paper reports that with smaller models, few-shot chain of thought using labelled exemplars beat the analogical method, and notes that generation can fail when the model has not learned related problems. Test on your own model before you adopt it.
- Cost and latency. Three exemplars with solutions can triple the output tokens of a plain chain-of-thought answer. Output tokens are usually the expensive and slow ones, so measure this before routing all traffic through it.
- Evaluation contamination. If the model recalls the benchmark problem itself as an exemplar, scores rise for the wrong reason. Look for exemplars that match test items word for word.
- Leaking to users. Raw exemplars are scaffolding. Show only the parsed solution unless the product is a tutor, where the exemplars may actually be the most useful part.
Evaluating it on your workload
Run a small, fixed comparison before you adopt it. Pick 100 to 300 representative problems with checkable answers. Run zero-shot chain of thought, your current few-shot prompt, and the analogical prompt at K = 3 at temperature zero, then the analogical prompt with five-way voting. For each, record accuracy, the rate of missing answer markers, mean output tokens and p95 latency. Then do three cheap ablations: K = 1 against K = 3, with and without the word distinct, and with and without the knowledge step.
Also hand-check exemplars on a sample of 30 runs. Count how many exemplars are correct and how many share a method with the target. If fewer than about two in three are correct, the gains you see are fragile and will move when the model version changes. Re-run the whole comparison on every model upgrade, because self-generated exemplars depend on what the model knows more than any other prompting technique does.
For broader context on reasoning prompts, see chain-of-thought prompting and self-consistency. The voting code above is a direct application of the latter.
Operational guidance
- Cap output tokens with headroom. A truncated response loses the answer marker, which is the last thing written.
- Cache by normalised problem text. Hosted models are not perfectly deterministic even at temperature zero, but repeated questions still do not need fresh exemplar generation.
- Log the exemplars along with the answer. When an answer is wrong, the exemplars usually explain why, and they show you which validators to add.
- Route by difficulty. Easy inputs rarely need exemplars. A cheap classifier, or a first attempt with a confidence check, can send only hard inputs to the analogical prompt.
- Keep the template under version control together with the evaluation set, and treat prompt changes like code changes.
What to do next
- Build a 100 to 300 item evaluation set with checkable answers for one real task.
- Measure zero-shot chain of thought and your current few-shot prompt as baselines, including token cost and latency.
- Add the analogical template above with K = 3, the answer marker and the last-marker parse, and compare.
- Hand-audit 30 runs of exemplars for correctness and method relevance before trusting the score.
- If exemplar errors are common, move to the two-call form and add cheap validators.
- If you have labelled data, test a retrieval-first hybrid that falls back to self-generated exemplars.
- Re-run the comparison on every model change and keep the winner behind a flag.