Few-shot prompting means putting a handful of worked examples, called demonstrations, into the prompt before the real input, so the model infers the task from them. It was the headline result of the GPT-3 paper (Brown et al., 2020), and it is still the cheapest way to make a model follow a precise output format or an idiosyncratic labelling policy without training anything.

It is also widely misunderstood. People assume the model learns the input-to-label mapping from the examples, then are surprised when shuffling the examples changes accuracy, when a model starts over-predicting whichever label came last, or when the output copies an entity from an example. This article explains what demonstrations really do, the biases they introduce, how to correct them with calibration, and how to measure whether each example earns its tokens. For designing a fixed example set (formatting, ordering, caching) see the few-shot architecture guide; for choosing examples per request see dynamic example selection.

Advertisement

From first principles: conditioning, not learning

A language model computes a probability distribution over the next token given everything before it. Few-shot prompting changes the everything before it. No weights move; the demonstrations simply make some continuations far more likely than others. If the prompt contains three lines of the form Review: ... / Sentiment: positive, then after Sentiment: the tokens positive and negative become overwhelmingly probable and free-form prose becomes unlikely.

This is called in-context learning, and its practical consequence is that a demonstration can teach several different things at once. Separating them is the key to using few-shot prompting well:

  1. Format: the shape of the answer, its delimiters, field order and length.
  2. Label space: which outputs are allowed at all.
  3. Input distribution: what typical inputs look like, so the query looks in-distribution.
  4. Mapping: which input gets which label, the part people usually assume is being taught.
One few-shot promptInstructiontask + label setDemo 1: input -> labelDemo 2: input -> labelDemo k: input -> labelQuery input-> model completesWhat demos carryformat, label space,input distribution, mappingBiases they addmajority, recency,common-tokenCalibrationcontent-free proberescale label probsMeasurement harnesszero-shot baselinek-shot, P orderingsshuffled-label controlleave-one-out per demokeep only demosthat earn tokensDemonstrations change the conditional distribution the model samples from; they do not change weights.Treat every demo as a hypothesis about what it teaches, and test it like one.
Anatomy of a few-shot prompt, the four things a demonstration can carry, the biases it adds, and the harness that tests each example.

Min et al. (2022), in a paper titled Rethinking the Role of Demonstrations, ran the decisive ablation: on a range of classification and multiple-choice tasks, replacing the gold labels in demonstrations with random labels barely reduced accuracy, while removing the format, the label space or realistic inputs hurt much more. For those models and tasks, the first three items did most of the work. Later work (Wei et al., 2023) found that larger models do follow the mapping more, to the point of adopting deliberately flipped labels. The safe conclusion is not a rule about models but a habit: test which of the four your examples are actually providing, because it varies by model and task.

The biases demonstrations introduce

Zhao et al. (2021), Calibrate Before Use, documented three systematic biases in few-shot classification:

  • Majority-label bias. The model over-predicts whichever label appears most often among the demonstrations. Three positives and one negative pushes borderline inputs to positive.
  • Recency bias. Labels near the end of the prompt are favoured. The same four examples in a different order can produce a different answer for the same query.
  • Common-token bias. Labels that are frequent tokens in pretraining text are preferred over rare ones, independent of the demos, so a label called other can win against billing_dispute for reasons unrelated to the input.

Lu et al. (2022), Fantastically Ordered Prompts, showed the order effect can be large: across permutations of the same examples, accuracy ranged from near state-of-the-art to near chance on some tasks and models. Order sensitivity is weaker in newer, instruction-tuned models, but it has not vanished, and the only way to know its size for your setup is to measure it. A prompt whose accuracy swings by ten points between orderings is not a prompt you have finished.

Advertisement

Correcting bias: contextual calibration

The fix in Zhao et al. is simple enough to implement in an afternoon, provided your API returns log-probabilities for the label tokens (many do, some do not, and some restrict it; if yours does not, skip to the measurement harness, which works with sampled outputs). The idea: feed the model your few-shot prompt with a content-free query such as N/A, an empty string, or [MASK]. A well-calibrated model should be indifferent between labels on such an input; whatever preference it shows is bias from the prompt itself. Divide it out.

import numpy as np

LABELS = ["billing", "bug", "how_to", "other"]

def label_probs(prompt, query, client):
    """Probability of each label's first token after the prompt. Returns a vector over LABELS."""
    lp = client.next_token_logprobs(prompt + query + "\nLabel:")   # your API wrapper
    raw = np.array([np.exp(lp.get(" " + l, -1e9)) for l in LABELS])
    if raw.sum() == 0:              # no label in the returned top-k: treat as uniform
        return np.full(len(LABELS), 1.0 / len(LABELS))
    return raw / raw.sum()          # renormalise over the label set only

def calibrate(prompt, client, probes=("N/A", "", "[MASK]")):
    p_cf = np.mean([label_probs(prompt, q, client) for q in probes], axis=0)
    p_cf = np.maximum(p_cf, 1e-6)   # top-k APIs can omit a label; avoid divide-by-zero
    W = np.diag(1.0 / p_cf)         # Zhao et al.: W = diag(p_cf)^-1, b = 0
    return W

def predict(prompt, query, client, W):
    p = W @ label_probs(prompt, query, client)
    return LABELS[int(np.argmax(p))]

Two caveats keep this honest. Labels must be distinguishable by their first token, or you must sum log-probabilities over each label's full token sequence. And calibration fixes a constant bias, not a wrong mapping: if the examples teach the wrong policy, calibration makes the model confidently wrong in a more balanced way. Recompute W whenever the prompt or the model version changes, because the bias belongs to that exact pair.

Worked example: is each example earning its tokens?

A support team classifies tickets into billing, bug, how_to and other. Their prompt has eight demonstrations, about 900 tokens, prepended to a million tickets a month. They have 400 hand-labelled tickets for evaluation. The question is not whether few-shot works but which of the eight examples matter, and whether any hurt.

The harness runs five measurements on the same evaluation set:

  1. Zero-shot baseline with the instruction and label definitions only. If this is within a point of the few-shot score, the examples are mostly paying for format, and a structured-output schema may do that job more cheaply.
  2. k-shot over several orderings (say five random permutations). Report mean and spread, not the best run, or you are selecting on noise.
  3. Shuffled-label control: same inputs, labels permuted. A small drop means the model leans on format and label space; a large drop means it is reading the mapping, so example correctness matters a lot.
  4. Leave-one-out: drop each example in turn. Examples whose removal does not lower accuracy are candidates to cut; examples whose removal raises accuracy are actively harmful.
  5. Calibrated vs uncalibrated, if logprobs are available.
import itertools, random, statistics

def accuracy(demos, evalset, run):            # run(prompt, query) -> label
    prompt = build_prompt(INSTRUCTION, demos)
    return sum(run(prompt, x) == y for x, y in evalset) / len(evalset)

def order_spread(demos, evalset, run, n=5, seed=0):
    rng = random.Random(seed)
    scores = []
    for _ in range(n):
        d = demos[:]; rng.shuffle(d)
        scores.append(accuracy(d, evalset, run))
    return statistics.mean(scores), max(scores) - min(scores)

def leave_one_out(demos, evalset, run):
    full, _ = order_spread(demos, evalset, run)
    return {i: full - order_spread(demos[:i] + demos[i+1:], evalset, run)[0]
            for i in range(len(demos))}       # positive = demo helps

def shuffled_labels(demos, seed=0):
    labels = [y for _, y in demos]; random.Random(seed).shuffle(labels)
    return [(x, y) for (x, _), y in zip(demos, labels)]

Suppose the results are: zero-shot 81%, eight-shot 88% with a 4-point spread across orderings, shuffled labels 84%, and leave-one-out shows two examples contributing nothing and one costing two points (it was mislabelled: a refund question labelled how_to). Fixing the label and dropping the two idle examples gives six examples, about 680 tokens, at 89% with a 2-point spread. At a million calls a month, 220 fewer input tokens per call is 220 million tokens saved monthly, while accuracy rose. With 400 evaluation items, a difference of one or two points is within noise; confirm small gains on a second sample before believing them.

Keep this harness in the repository and rerun it on every model upgrade; the prompt evaluation guide covers building the labelled set and tracking scores over time.

Generation tasks: few-shot for format and style

For extraction, rewriting and code generation the label space is open, and demonstrations mostly teach format and style. Two failure modes dominate. Content leakage: the model copies entities, numbers or phrases from an example into an unrelated answer, especially when the query lacks the corresponding field. Use obviously synthetic values in examples (ACME-0001, Jane Example) so leaks are easy to grep for, and include at least one example whose correct output is an explicit empty or null field. Length anchoring: outputs mimic the examples' length, so three long examples produce long answers for short inputs. Vary example length deliberately.

In chat APIs you can encode examples either inside one user message, wrapped in tags such as <example> and </example>, or as fabricated user and assistant turns. Fabricated turns are strong conditioning for format but can be confused with real conversation history in multi-turn apps; tagged blocks inside the system or user message are easier to keep separate from live context. When the output must be machine-parsed, combine a small number of examples with a schema, as described in structured output prompting; the schema enforces shape and the examples teach judgement about edge cases.

For reasoning tasks, demonstrations that include worked intermediate steps (chain-of-thought exemplars, Wei et al., 2022) teach the model to show its working before answering. The examples' reasoning must be correct and in the style you want, because the model copies its structure; the chain-of-thought guide covers when that helps and what it costs.

How many examples: from few-shot to many-shot

Classic few-shot means a handful, constrained historically by short context windows. Long-context models changed the economics. Agarwal et al. (2024), Many-Shot In-Context Learning, reported that moving from a few to hundreds or thousands of demonstrations kept improving results on many tasks, and that enough examples could overcome some pretraining biases that a handful could not. The costs are latency and input tokens on every call, which prompt caching can reduce when the example block is a stable prefix.

Rules of thumb that survive measurement: start with one example per label (or per output pattern) plus one for each known hard case; grow until the leave-one-out curve flattens; cover every label so majority bias has nothing to amplify; and prefer a few diverse, correct examples over many near-duplicates. When the example count you need climbs into the hundreds and the task is stable, compare the per-call cost against fine-tuning a smaller model, which bakes the mapping into weights and leaves the prompt short.

Failure modes and trade-offs

  • Silent mislabels. One wrong example teaches a wrong policy for the inputs that resemble it. Leave-one-out finds it; a second reviewer prevents it.
  • Label imbalance. Demo label counts leak into predictions through majority bias. Balance them or calibrate.
  • Distribution drift. Examples drawn from last year's tickets stop resembling today's inputs. Refresh from recent traffic and rerun the harness.
  • Test contamination. An evaluation item that also appears as a demonstration inflates scores. Deduplicate examples against the evaluation set.
  • Model upgrades. A new model version can respond differently to the same demonstrations; biases and calibration weights do not carry over.
  • Instruction-example conflict. When the instruction says one thing and the examples show another, models often follow the examples; check with the harness. Make them agree.
ApproachBest whenCosts
Zero-shot + clear instructionCommon task, format enforced by schemaMay miss house-style judgement
Few-shot, fixed setStable task, a few edge cases to teachTokens per call; order and label biases
Few-shot, retrieved per requestDiverse inputs, large labelled poolRetrieval latency; selection must be evaluated
Many-shotHard mappings, long-context model, cacheable prefixLatency and input tokens
Fine-tuningHigh volume, stable task, labelled dataTraining pipeline, retraining on change

What to do next

  1. Write down what each demonstration is meant to teach: format, label space, input distribution or mapping.
  2. Build a labelled evaluation set of at least a few hundred items, deduplicated against your examples.
  3. Measure zero-shot, k-shot over five orderings, a shuffled-label control and leave-one-out; cut idle and harmful examples.
  4. Balance labels across demonstrations, and add contextual calibration if your API exposes label log-probabilities.
  5. For generation tasks, use synthetic values in examples, include an empty-field case and vary example length.
  6. Rerun the harness on every model or prompt change, and compare token cost against a schema, retrieval or fine-tuning.
Key takeaway: Few-shot demonstrations condition a model rather than train it, and they often teach format, label space and input style more than the input-to-label mapping. They also add majority, recency and common-token biases that ordering and calibration can expose and reduce. Treat every example as a claim to test: measure zero-shot, orderings, shuffled labels and leave-one-out, keep only the examples that raise accuracy, and rerun the measurement whenever the model changes.