A fixed few-shot prompt shows every request the same handful of examples. That works when inputs look alike. It breaks when they do not: a text-to-SQL assistant that answers questions about twenty tables cannot show a worked example for each in every prompt, and the three it does show are irrelevant to most questions. Dynamic few-shot fixes this by choosing the examples at request time, from a pool, based on the input.

The idea is simple. The details decide whether it helps. Which examples count as similar, how to avoid five near-copies, how to keep labels balanced, where to put the examples so the prompt cache still works, and how to prove the selector beats random choice: those are algorithm questions, and they are this article's subject. The general craft of few-shot prompting is covered in the few-shot prompting guide, and the surrounding system (exemplar store, feedback loop) in the dynamic few-shot architecture article. Here we write the selector itself.

Advertisement

What selection can change, and what it cannot

Research on in-context learning gives a useful mental model. Liu and colleagues showed in 2021 that picking examples semantically close to the test input, using a sentence encoder and nearest-neighbour search, beat random examples on several tasks. Min and colleagues found in 2022 that, for classification tasks, much of the benefit of demonstrations comes from showing the input distribution, the label space and the output format, and that even randomly assigned labels hurt less than one would expect. Lu and colleagues showed the same examples in a different order could swing accuracy widely, and Zhao and colleagues documented a majority-label bias and a recency bias: models lean toward labels that appear often and labels that appear last.

Put together: examples teach format and scope reliably, and they teach the specific mapping only partly. Selection pays most when the task has many sub-patterns that look different on the surface, such as SQL over many schemas, extraction across document types, or code in several frameworks. It pays least on a single-pattern task where any three clean examples already teach the format. If your fixed prompt already does well, a selector adds latency and moving parts for little gain. Measure before you build.

The request-time pipeline

A selector runs in four steps. Recall pulls a generous candidate set, say fifty, from the pool using cheap similarity. Set selection picks k of those, trading relevance against redundancy, label balance and a token budget. Ordering arranges them. Assembly places them between the static instructions and the user input. The pool sits outside the request path: it is curated, versioned and evaluated offline.

Per request: retrieve candidates, select a set, order it, then assemble the promptUser inputthe queryEmbed + BM25query vector, termsCandidate recalltop 50 from poolSet selectionMMR, labels, budgetExample poolvetted, versionedANN indexOrdermost similar lastPromptstatic instructions (cacheable prefix) | k examples | the queryquery goes lastOffline evalrandom vs kNN vs MMR on a frozen test setPool curationadd fixes, retire examples that misleadnew pool version
Recall is cheap and wide; set selection is where the quality comes from. The pool and its evaluation are an offline loop, and every pool change ships as a new version.
Advertisement

Building the pool

The selector can only choose what the pool contains, so the pool deserves more care than the algorithm. Each record should hold the input, the ideal output, a label or task type where one exists, the token cost of the rendered example, a source, and a version. Store the rendered text, not just raw fields, so the token count you budget against is the count you send.

Four rules keep a pool honest. Every example must be correct, because the model treats demonstrations as authoritative; one wrong SQL join in the pool will be copied faithfully whenever it is retrieved. Deduplicate near-identical inputs, or recall will return clusters. Keep the pool disjoint from your evaluation set, including paraphrases, or your offline numbers will flatter you. And cover the edges deliberately: refusals, empty results, ambiguous inputs that need a clarifying question. Su and colleagues showed that when you must choose which examples to annotate in the first place, picking a diverse and representative set (their vote-k method) beats random annotation, which is a good argument for curating rather than dumping production logs into the pool.

What to embed

Embed the input side of each example only, with the same encoder and the same preprocessing you apply to live queries. Embedding the output too drags the vector toward answer text the live query does not have. For structured tasks, normalise surface noise first: for text-to-SQL, embedding the question with literal values masked (dates, names, numbers replaced by placeholders) makes questions match on structure rather than on the customer name they happen to mention.

Pure dense similarity misses exact identifiers. A question naming the refunds table should retrieve examples that touch refunds, whatever the embedding thinks. Hybrid recall, unioning dense neighbours with BM25 or a simple keyword filter, fixes this cheaply. For pools under a few thousand items, brute-force cosine over a matrix is fast enough. Beyond that, an approximate index such as HNSW keeps recall in single-digit milliseconds.

Algorithm 1: top-k nearest neighbours

The baseline: rank candidates by cosine similarity and take the first k. It is easy to explain and often a large step up from random selection. Its flaw shows as soon as the pool has clusters. If the pool holds twelve examples about monthly revenue, a revenue question retrieves five of them, and the model sees one pattern five times instead of five patterns once. Top-k also ignores labels, so a classification query near a dense region of one class gets a prompt that is all that class, which feeds straight into majority-label bias.

Algorithm 2: maximal marginal relevance

Maximal marginal relevance (Carbonell and Goldstein, 1998) builds the set greedily. At each step it picks the candidate that maximises lambda times its similarity to the query, minus (1 - lambda) times its highest similarity to anything already chosen. With lambda at 1 it is top-k; lower values buy diversity. Values between 0.5 and 0.8 are a reasonable starting range, but tune it on your evaluation set rather than trusting a default.

import numpy as np

def mmr(query_vec, cand_vecs, k, lam=0.7):
    """Greedy maximal marginal relevance. Vectors must be L2-normalised."""
    rel = cand_vecs @ query_vec                  # similarity to the query
    pair = cand_vecs @ cand_vecs.T               # similarity between candidates
    chosen, remaining = [], list(range(len(cand_vecs)))
    while remaining and len(chosen) < k:
        if chosen:
            redundancy = pair[np.ix_(remaining, chosen)].max(axis=1)
        else:
            redundancy = np.zeros(len(remaining))
        scores = lam * rel[remaining] - (1 - lam) * redundancy
        best = remaining[int(np.argmax(scores))]
        chosen.append(best)
        remaining.remove(best)
    return chosen

Run MMR over the recalled fifty, not over the whole pool: the pairwise matrix is quadratic in candidate count, and fifty squared is trivial.

Algorithm 3: label-stratified and budget-aware selection

For classification-shaped tasks, enforce a spread of labels. One practical rule: take the best candidate from each of the top m labels by aggregate similarity, then fill the remaining slots by MMR. This keeps the decision boundary visible and blunts majority-label bias without showing irrelevant classes.

Then respect the budget. Examples vary in length, and a single long example can cost as much as four short ones. Treat selection as filling a token budget rather than a count: skip candidates that would overflow it and keep going down the ranked list. Cap k anyway, because past a point extra examples add cost and dilute attention without improving accuracy.

Algorithm 4: learned retrievers and set-level selection

Similarity to the query is a proxy for what you actually want: examples that make the model answer correctly. Rubin, Herzig and Berant (2022) trained a retriever directly on that signal. Their method, EPR, scores candidate examples by how much each raises the language model's likelihood of the correct output on training inputs, then trains a dense retriever with contrastive learning to find high-scoring examples. Later work selects sets jointly, since examples interact.

A learned retriever is worth it when you have thousands of labelled pairs, a stable model, and a measured gap between kNN and an oracle selector. It has a cost: the scores are specific to the model that produced them, so a model upgrade can quietly invalidate the retriever. For most teams, hybrid recall plus MMR plus label stratification captures most of the gain.

Ordering and placement

Order matters because of recency bias. A common heuristic is to put the most similar example last, nearest the query, so the closest pattern is freshest. Whatever rule you pick, make it deterministic, because random ordering turns into random variance in production and makes regressions impossible to bisect.

Placement interacts with prompt caching. Providers cache a stable prefix, and dynamic examples are by definition not stable. Put everything static first (role, instructions, schema notes, output contract), then the selected examples, then the query. The static block stays cacheable; only the tail is recomputed. Putting examples ahead of the instructions throws the cache away on every request. The same budgeting logic applies when examples compete with retrieved documents for space, which the context-packing article treats as one constrained problem.

Worked example: a text-to-SQL selector

A finance team's assistant answers questions over twelve tables. The pool holds 400 vetted question-and-SQL pairs, each tagged with the tables it touches. The selector recalls by dense similarity on masked questions and by table-name overlap, runs MMR, guarantees at least one example per table the question names, and stops at a 1,500-token budget.

import re

MASKS = [(re.compile(r"\b\d{4}-\d{2}-\d{2}\b"), "<date>"),
         (re.compile(r"\b\d+(\.\d+)?\b"), "<num>"),
         (re.compile(r"'[^']*'"), "<str>")]

def mask(q):
    for pat, tok in MASKS:
        q = pat.sub(tok, q)
    return q

def select_examples(question, pool, index, embed, tables, k_max=6, budget=1500):
    qv = embed(mask(question))
    named = {t for t in tables if t in question.lower()}
    dense = index.search(qv, top=50)                          # ids by cosine
    lexical = [e.id for e in pool if named & e.tables][:50]   # table overlap
    cands = list(dict.fromkeys(dense + lexical))              # union, keep order
    order = mmr(qv, index.vectors(cands), k=len(cands), lam=0.7)
    picked, used, covered = [], 0, set()
    # pass 1: one example per named table; pass 2: fill by MMR rank
    for must_cover in (True, False):
        for i in order:
            ex = pool[cands[i]]
            if ex in picked or used + ex.tokens > budget or len(picked) >= k_max:
                continue
            if must_cover and not (ex.tables & (named - covered)):
                continue
            picked.append(ex)
            used += ex.tokens
            covered |= ex.tables
    picked.sort(key=lambda e: float(index.vector(e.id) @ qv))  # most similar last
    return picked

The prompt renders the static block (role, the twelve table schemas, the rule to output a single SELECT statement), then each picked example as a question and SQL pair, then the live question. Log the selected example IDs and the pool version with every request; when an answer is wrong, the first question is which examples the model saw.

Evaluating a selector

Never ship a selector on intuition. Freeze a test set disjoint from the pool, fix the model version and temperature, and run an ablation: no examples, fixed examples, random-k, top-k, MMR, and MMR with stratification, with several random seeds for the random arm. For SQL, score by executing against a fixture database and comparing result sets, not by string match. Report accuracy, mean prompt tokens and p95 latency per arm. A selector earns its place only if it beats fixed examples by more than seed-to-seed noise at acceptable cost.

Also slice the results. Dynamic selection usually helps rare sub-patterns most and common ones barely at all; an aggregate number hides both the win and any regression in a slice.

Failure modes

  • Copying. A near-identical example makes the model reuse its answer, literals and all. Masking and a similarity ceiling (skip candidates above, say, 0.98) reduce it.
  • Wrong examples retrieved precisely when relevant. A bad pool entry hurts exactly the queries it resembles. Audit examples that appear in failing traces.
  • Leakage. Test items in the pool produce excellent offline numbers and ordinary production ones.
  • Encoder drift. Changing the embedding model without re-embedding the pool silently returns nonsense neighbours.
  • Prompt injection via the pool. If examples come from user-generated content, they are untrusted input placed in a position of authority. Curate them.
  • Cache loss. Examples placed before static instructions turn every request into a full-price prefill.

Trade-offs

ApproachStrengthCost or risk
Fixed examplesCacheable, predictable, trivialIrrelevant to diverse inputs
Top-k kNNSimple, big gain on clustered tasksRedundant sets, label skew
MMR plus stratificationRelevant and varied; boundary visibleOne tuned parameter, more code
Learned retrieverOptimises the real objectiveTraining data, model-specific, retrain on upgrade
Fine-tuning insteadNo examples in the prompt at allTraining pipeline, slower to change

What to do next

  1. Run a fixed-versus-random-versus-kNN ablation on a frozen test set before writing any selector code.
  2. Build the pool as versioned records with rendered text, token counts, labels and sources, disjoint from the test set.
  3. Embed masked inputs only, and add lexical recall for identifiers such as table or API names.
  4. Implement MMR over a recalled set of about fifty, add label or coverage constraints, and fill to a token budget.
  5. Place static instructions first, examples next, the query last, and confirm the cache hit rate.
  6. Log example IDs and pool version per request, and review examples that recur in failing traces every week.
Key takeaway: Dynamic few-shot is a small retrieval system whose output is a set, not a list. Recall widely with hybrid similarity on masked inputs, then pick a set that is relevant, varied, label-balanced and within a token budget; maximal marginal relevance with simple constraints gets most of the value. Order deterministically, keep the static prefix cacheable, and prove the selector against fixed and random examples on a frozen, leak-free test set before you trust it.