An emergent ability, in the sense used by Wei and colleagues in their 2022 paper of that name, is one that is absent in smaller models and present in larger ones, so that you could not have predicted it by extrapolating the smaller models' performance. Few-shot arithmetic, multi-step word problems and the benefit of chain-of-thought prompting were among their examples. In 2023 Schaeffer, Miranda and Koyejo argued that many such jumps are produced by the choice of metric rather than by the model, and the debate is not settled.
For a prompt engineer, the debate matters less than its practical consequence: a prompt is not portable across models. A technique that transforms results on a large model can do nothing, or harm, on a small one, and the threshold depends on the model family, its training data and its tuning, not on a parameter count you can memorise. The theory and the mathematics are covered in the transformer-math series, starting with emergence in depth. This article is about what to do with it: how to measure whether a capability is present in the models you can afford, and how to rewrite prompts when it is not.
Two kinds of emergence a prompt engineer meets
Emergent tasks are tasks a model either can or cannot do with a plain prompt: multi-digit arithmetic, following a long chain of instructions, keeping a strict output schema across a long document. Emergent prompting techniques are methods whose benefit appears only above some capability level. The chain-of-thought paper by Wei and colleagues (2022) reported that asking for intermediate reasoning gave large gains on arithmetic benchmarks for the largest models they tested, around 100 billion parameters and above, and little or negative benefit for smaller ones, which tended to produce fluent but wrong chains. Instruction following and in-context learning from few examples show similar patterns in the literature.
The second kind is the one that bites in production. Teams develop a prompt against a frontier model, then move to a cheaper or self-hosted model for cost or privacy, and the technique that carried the prompt stops working. The model is not broken; the prompt depended on a capability the smaller model lacks.
Scale here means more than parameters. Training tokens, data quality, instruction tuning and preference tuning all move the threshold, which is why a recent 8-billion-parameter model can outperform a much larger model from a few years earlier on tasks that once looked emergent. Treat any published threshold as a statement about the models in that paper.
Why the curve can look like a cliff
The measurement critique is easy to see with one line of arithmetic. Suppose a task needs eight correct steps and the model gets each step right independently with probability p. Exact-match accuracy is p^8. At p = 0.5 that is 0.004; at p = 0.8 it is 0.17; at p = 0.95 it is 0.66. A steady improvement in per-step skill looks flat for a long time and then shoots up. Schaeffer and colleagues showed that switching such tasks to continuous metrics, like per-token or per-step accuracy, often turns the cliff into a smooth curve.
Two practical lessons follow. First, your production metric is usually all-or-nothing: the JSON parses or it does not, the invoice total is right or it is not. So even if the underlying skill improves smoothly, what your users experience can genuinely jump between model sizes. Second, a continuous metric is a much better instrument for predicting where a model will land, because it moves before the exact-match number does. Measure both. The full derivation, including thresholds and the sigmoid in log-compute, is in the mathematics of emergent abilities.
Building a capability probe
Do not decide from benchmarks whether a model can handle your prompt; measure it on your task. A capability probe is a small evaluation harness that runs a grid of prompt variants across a ladder of candidate models and scores each output with two rulers.
import random, statistics
MODELS = ["small-8b", "mid-30b", "large-hosted"] # your candidate ladder
VARIANTS = {"zero_shot": zero_shot, "few_shot": few_shot, "cot": chain_of_thought}
def score(pred, gold):
# Two rulers on the same output.
exact = float(pred == gold) # all-or-nothing
steps = [a == b for a, b in zip(pred.steps, gold.steps)]
partial = sum(steps) / max(len(gold.steps), 1) # per-step credit
return exact, partial
def bootstrap_ci(xs, n=2000, seed=0):
rng = random.Random(seed)
means = sorted(statistics.fmean(rng.choices(xs, k=len(xs))) for _ in range(n))
return means[int(0.025 * n)], means[int(0.975 * n)]
results = {}
for model in MODELS:
for name, build_prompt in VARIANTS.items():
exact, partial = [], []
for case in EVAL_SET: # 200+ held-out cases
out = parse(call(model, build_prompt(case), temperature=0))
e, p = score(out, case.gold)
exact.append(e); partial.append(p)
results[(model, name)] = (statistics.fmean(exact), bootstrap_ci(exact),
statistics.fmean(partial))
for (model, name), (em, ci, part) in sorted(results.items()):
print(f"{model:13} {name:9} exact={em:.2f} ci=({ci[0]:.2f},{ci[1]:.2f}) step={part:.2f}")Design notes. Use a held-out set of at least a couple of hundred cases drawn from real traffic, including the awkward ones; ten hand-picked examples will tell you what you want to hear. Fix temperature at zero for comparability, then run a second pass at your production temperature. Report confidence intervals, because a difference of three points on 100 cases is noise. Keep the parse step strict and count parse failures as failures, since a smaller model's first symptom of missing capability is usually malformed output. A general framework for prompt evaluation is in prompt evals.
Worked example: extraction with arithmetic
A finance team extracts line items from supplier invoices and must return the items, a computed subtotal, tax and total as JSON. The prompt was developed on a large hosted model using chain-of-thought: list the items, compute each line, sum, apply tax, then emit JSON. The team wants to move to a self-hosted 8-billion-parameter model.
They run the probe on 300 historical invoices. The numbers below are illustrative of the pattern teams commonly see, not measurements of any named model.
| Model | Prompt | Exact match | Per-step accuracy | Parse failures |
|---|---|---|---|---|
| Large hosted | Chain of thought | 0.94 | 0.99 | 0% |
| Large hosted | Zero-shot JSON | 0.88 | 0.97 | 0% |
| 8B self-hosted | Chain of thought | 0.41 | 0.86 | 9% |
| 8B self-hosted | Zero-shot JSON | 0.47 | 0.88 | 3% |
| 8B self-hosted | Decomposed + calculator tool | 0.90 | 0.97 | 1% |
Read the table carefully. The small model's per-step accuracy, 0.86 to 0.88, is not far behind, but exact match collapses because invoices have many lines and every one must be right, the p^n effect in practice. Chain of thought hurts the small model: its long reasoning adds steps to get wrong and pushes it out of the JSON format. The fix was not a better wording but a different structure: extract line items only, compute the arithmetic in code, and validate the schema. That moved the arithmetic, which the small model lacked, out of the model entirely.
Reading probe results: a decision procedure
A probe produces a grid of numbers; turn it into a decision with a fixed procedure rather than intuition.
- Start from the production threshold. Decide in advance what exact-match rate the business needs, for example 0.95 for automated posting or 0.80 when a human reviews every output. Any model and prompt below it is out, however good the partial-credit number looks.
- Compare the two rulers. High per-step accuracy with low exact match means the model has most of the skill and fails on accumulation; decomposition and tools usually close that gap. Low per-step accuracy means the skill itself is missing, and prompt changes rarely rescue it.
- Check overlap of intervals. If the confidence intervals of two variants overlap, treat them as equal and pick the cheaper one.
- Look at the failures, not only the rates. Read thirty failed outputs per cell. Format breaks, arithmetic slips and misread inputs need different fixes.
- Price it. Multiply tokens per call, including any reasoning and retries, by volume. A small model that needs three calls and a vote can cost more than one large-model call.
Porting prompts to smaller models
When a probe shows a capability gap, the techniques that close it all reduce how much the model must do in one pass.
- Decompose. Split one demanding prompt into a chain of simpler ones, each checkable. Fewer steps per call means a smaller exponent in
p^n. Least-to-most prompting is a structured version of this. - Offload to tools. Arithmetic, date handling, lookups and sorting belong in code. Ask the model to extract operands, not to compute.
- Show, do not tell. Smaller models follow abstract instructions less reliably but imitate examples well. Two or three well-chosen examples often beat a paragraph of rules; see few-shot prompting.
- Constrain output. Use the serving stack's structured or grammar-constrained decoding where available, so format is guaranteed rather than requested.
- Shorten or drop reasoning. If chain of thought hurts, remove it, or ask for a brief fixed-format scratchpad rather than free prose.
- Sample and vote, carefully. Self-consistency helps when per-step accuracy is decent, but multiplies cost, and can erase the savings that motivated the smaller model.
- Route. Send easy cases to the small model and hard ones, detected by length, validation failure or low confidence, to the large one.
Going the other way: upgrades and inverse scaling
Emergence also affects upgrades. A new, more capable model can make old workarounds counterproductive: heavy few-shot examples can anchor a strong model to the examples' quirks, and over-specified step lists can stop it from using a better approach. The Inverse Scaling Prize work (McKenzie and colleagues, 2023) collected tasks where larger models did worse, often because they followed a misleading pattern in the prompt more faithfully. The lesson is symmetric: re-run the probe on every model change, in both directions, and delete scaffolding that no longer earns its tokens.
Failure modes
| Failure | What it looks like | Prevention |
|---|---|---|
| Assuming portability | Quality drops silently after a model switch | Probe every candidate before migration |
| Benchmark substitution | Choosing a model by leaderboard rank for a task unlike the benchmark | Use your own held-out set |
| Single-metric blindness | Exact match at zero hides a nearly capable model | Track per-step or partial credit too |
| Tiny eval sets | Decisions flip on rerun | Hundreds of cases and confidence intervals |
| Technique cargo cult | Chain of thought added to every prompt, hurting small models | Ablate each technique per model |
| Stale scaffolding | Workarounds for old models degrade new ones | Re-probe and simplify on upgrade |
Trade-offs
Bigger models buy capability margin: more tasks work with simple prompts, and prompts survive edits. Smaller models buy cost, latency and control, but demand more engineering in decomposition, tools and validation, and that engineering is code you must maintain. The right choice is usually a mix: a small model on the high-volume, well-structured path, a large model as fallback, and a probe that tells you when the boundary between them moves. Prompting techniques are not free either; chain of thought and self-consistency trade tokens and latency for accuracy, and on some models they trade them for nothing. Chain-of-thought prompting covers when the trade pays.
What to do next
- Build a held-out set of 200 or more real cases for each important prompt, with gold outputs.
- Write a scorer with two rulers: the all-or-nothing production metric and a partial-credit metric.
- Run the probe across at least three model sizes and each prompting technique you use, with confidence intervals.
- Ablate chain of thought, few-shot examples and self-consistency per model; keep only what helps.
- For any gap, move deterministic work into code and decompose the prompt before trying rewording.
- Add the probe to your release process so every model change, up or down, is measured before rollout.