Zero-shot prompting means asking a model to do a task from an instruction alone, with no worked examples in the prompt. It is the first thing everyone tries and often the last thing they need. It is also the style where the gap between 'it looked right on five inputs' and 'it is right on the traffic we serve' is widest, because nothing in the prompt anchors the model to your label set, your format or your idea of the task.
How to write a zero-shot prompt as a specification, and when to move to few-shot or chain of thought, is covered in Zero-shot prompting as a specification. This article covers the measurement side: why zero-shot works at all, how to read a probability for every candidate answer instead of parsing free text, how to correct the model's built-in preference for some labels, and how to measure how much your accuracy depends on wording you chose by accident.
What zero-shot means, precisely
The term has two lineages, and mixing them up causes confusion in design reviews. In classic machine learning, zero-shot meant recognising classes never seen in training. In prompting, it means no demonstrations in the context window: the model has certainly seen tasks like yours during training, but your prompt contains only the instruction and the input. A third usage survives in libraries: the Hugging Face zero-shot-classification pipeline uses a natural language inference model and asks, for each label, whether the input entails a hypothesis such as 'This text is about billing.' That is a cheap baseline, but a different mechanism.
In this article zero-shot means the second sense: an instruction-tuned generative model, a prompt with an instruction and an input, and no examples.
Why a bare instruction works
A base model trained only on next-token prediction is a document completer. Given 'Classify this ticket: ...' it may continue with another ticket, a forum reply or a list of categories. Zero-shot ability became dependable when models were fine-tuned on many tasks phrased as instructions. The 2021 paper 'Finetuned Language Models Are Zero-Shot Learners' (FLAN) showed that instruction tuning a 137-billion-parameter model on dozens of NLP datasets described with natural language templates made it beat zero-shot GPT-3 175B on most of the unseen tasks evaluated, and the T0 work the same year found a similar effect with multitask prompted training.
Two practical consequences follow. First, the model is matching your instruction against the many task phrasings it was tuned on, so naming the task in its conventional form ('classify', 'extract', 'summarise in one sentence', 'translate to German') beats an inventive description. Second, the model brings priors: preferences for certain label words, for the first or last option listed and for answers that look common. Few-shot examples partly overwrite those priors; zero-shot leaves them in place, which is why the calibration step later in this article matters more here than anywhere else.
Two ways to read the answer
Path A is generate and parse: let the model write, then extract the label with a regular expression or a JSON parser. It works with any API and is what most teams ship. Its weaknesses: one answer with no confidence, hedged replies that need ever-growing parsing rules, and answers that change with sampling temperature. Constrained decoding and schema-enforced outputs, covered in structured output architecture, remove most parsing failures but still return only one choice.
Path B is label scoring: for a closed label set, compute the log-probability the model assigns to each label as the continuation of the prompt, then pick the highest. You get a full distribution over labels from one forward pass per label (or one batched pass), with no parsing and no sampling noise. It needs access to token log-probabilities, which local models always provide and many hosted APIs do not, so check yours before designing around it.
A label scorer that is actually correct
Label scoring is short to write and easy to get subtly wrong. The code below scores every label as the assistant's reply under the model's own chat template. Three details carry the correctness: the chat template is applied, because an instruct model scored on raw text sees a distribution it was not tuned on; the prompt and the label are tokenised together, because tokenising the label alone splits it differently at the boundary (a leading space or newline often merges into the first label token); and each label token is scored with the logits from the position before it, since a causal model's output at position i predicts token i+1.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL = "Qwen/Qwen2.5-1.5B-Instruct" # any causal LM that ships a chat template
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype=torch.bfloat16).eval()
LABELS = ["billing", "bug", "account", "feature request"]
INSTRUCTION = ("Classify the support ticket into exactly one category: "
"billing, bug, account, feature request. Answer with the category name only.")
def prompt_for(ticket, instruction=INSTRUCTION, fmt="Ticket: {t}"):
msgs = [{"role": "user", "content": instruction + "\n\n" + fmt.format(t=ticket)}]
return tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
@torch.no_grad()
def label_logprobs(ticket, **kw):
prefix = prompt_for(ticket, **kw)
prefix_ids = tok(prefix, add_special_tokens=False).input_ids
n = len(prefix_ids)
scores = {}
for label in LABELS:
ids = tok(prefix + label, add_special_tokens=False, return_tensors="pt").input_ids
# tokenise prompt and label TOGETHER, then check the prompt part did not merge
assert ids[0, :n].tolist() == prefix_ids, "tokens merged across the boundary"
logp = torch.log_softmax(model(ids).logits[0].float(), dim=-1)
targets = ids[0, n:] # the label's tokens
positions = torch.arange(n - 1, ids.shape[1] - 1) # token i is scored at i-1
scores[label] = logp[positions, targets].sum().item()
return scoresThe assertion catches the boundary problem instead of silently scoring the wrong tokens. Multi-token labels need a decision: summing log-probabilities penalises long labels such as 'feature request' (two or more tokens, each with its own probability below one). Options are to rename labels so each is one distinctive token, to present options with letters and score only the letter, or to divide by the token count. Pick one, apply it consistently and measure it, because each changes which label wins on borderline inputs.
Label bias and contextual calibration
Even with a perfectly neutral input, an instruction-tuned model does not split probability evenly across labels. The 2021 paper 'Calibrate Before Use: Improving Few-Shot Performance of Language Models' named three sources: majority-label bias from the examples shown, recency bias toward whatever appeared last, and common-token bias toward label words that are frequent in pretraining text. Zero-shot prompts have no examples, but the option order in the instruction and the label words themselves still bias the output.
The paper's fix, contextual calibration, is cheap. Feed the prompt with content-free inputs such as 'N/A', an empty string and '[MASK]', average the label distributions to estimate the prior the prompt alone induces, then divide every real prediction by that prior and renormalise.
import math
def to_probs(scores):
m = max(scores.values())
z = sum(math.exp(s - m) for s in scores.values())
return {k: math.exp(s - m) / z for k, s in scores.items()}
CONTENT_FREE = ["N/A", "", "[MASK]"]
def content_free_prior(**kw):
ps = [to_probs(label_logprobs(x, **kw)) for x in CONTENT_FREE]
return {k: sum(p[k] for p in ps) / len(ps) for k in LABELS}
def calibrated(ticket, prior, **kw):
p = to_probs(label_logprobs(ticket, **kw))
q = {k: p[k] / prior[k] for k in LABELS}
z = sum(q.values())
return {k: v / z for k, v in q.items()}A worked example shows why it matters. Suppose the content-free prior over (billing, bug, account, feature request) comes out as (0.50, 0.20, 0.20, 0.10): the model leans toward 'billing' before it reads anything. A real ticket scores (0.40, 0.35, 0.15, 0.10). Uncalibrated, 'billing' wins. Dividing by the prior gives (0.80, 1.75, 0.75, 1.00), which sums to 4.30, so the calibrated distribution is about (0.186, 0.407, 0.174, 0.233) and 'bug' wins clearly. The raw scores were never saying 'billing'; they were saying 'less billing than this prompt usually produces, much more bug'.
Calibration is not free accuracy. It can hurt when the real class distribution genuinely favours the label the model prefers. Measure both variants on a labelled development set and keep the winner, per prompt.
Measuring prompt sensitivity
Instructions that mean the same thing to a person can produce very different accuracy. The 2023 paper 'Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design' found that meaning-preserving formatting changes, such as separators, casing and spacing, moved accuracy by large margins on several open models, and that the best format for one model often was not the best for another. A single accuracy number for a single wording therefore tells you about that wording, not about the approach.
The fix is to treat wording as a random variable. Write several paraphrases of the instruction and several input formats, evaluate every combination on the same development set, and report the spread:
import itertools, statistics
INSTRUCTIONS = [
INSTRUCTION,
"Which team should handle this ticket? Options: billing, bug, account, feature request. "
"Reply with one option.",
"Read the customer message and output its category (billing, bug, account or "
"feature request) and nothing else.",
]
FORMATS = ["Ticket: {t}", "<ticket>\n{t}\n</ticket>", "Customer message:\n\"\"\"{t}\"\"\""]
def accuracy(dev, calibrate, **kw):
prior = content_free_prior(**kw) if calibrate else None
hits = 0
for ticket, gold in dev:
p = calibrated(ticket, prior, **kw) if calibrate else to_probs(label_logprobs(ticket, **kw))
hits += max(p, key=p.get) == gold
return hits / len(dev)
def sensitivity(dev, calibrate):
accs = [accuracy(dev, calibrate, instruction=i, fmt=f)
for i, f in itertools.product(INSTRUCTIONS, FORMATS)]
return min(accs), statistics.median(accs), max(accs)Read the result as follows. A narrow spread with a good median means the task is well within the model's zero-shot ability and wording barely matters. A wide spread means the model is near the edge of its ability and you are about to pick a winner by chance; the best cell on 200 examples often regresses on fresh traffic. That is the signal to add demonstrations (few-shot prompting in depth), split the task or move to a stronger model, not to keep polishing words. Wire the harness into your evaluation suite so a model upgrade re-runs it; see prompt evaluation architecture.
Worked example: routing support tickets
A team routes tickets into four queues with a 1.5-billion-parameter instruct model running locally. They label 300 recent tickets by hand, stratified so each queue has at least 50, and hold 100 of them back as a final test set. On the 200 development tickets they run the harness above with nine prompt variants, with and without calibration. The table shows the shape of results to expect; the numbers are illustrative, not a benchmark.
| Setting | Min accuracy | Median | Max | What it tells you |
|---|---|---|---|---|
| Generate + parse, temperature 0 | 0.71 | 0.78 | 0.84 | parsing failures and verbose replies cost points |
| Label scoring, uncalibrated | 0.74 | 0.81 | 0.86 | no parse failures; 'billing' over-predicted |
| Label scoring, calibrated | 0.80 | 0.84 | 0.87 | spread narrows; bias removed |
Three decisions follow. Calibrated scoring is adopted because its worst case beats the others' median. The prompt is chosen from the middle of the calibrated cells, not the single best one, and confirmed once on the held-out tickets. And because the scorer returns a probability, tickets whose top label scores below 0.5 go to human triage; the team tracks that abstention rate and the accuracy on the rest.
Failure modes
- Scoring without the chat template. The model is evaluated on a format it never saw in tuning; accuracy is lower and calibration estimates are wrong.
- Off-by-one log-probabilities. Using the logits at the label token's own position scores the next token instead. Results look plausible and are noise. Test with a label the prompt makes obvious.
- Label words that collide with the input. If the ticket text contains 'account', copying favours that label. Prefer label names that rarely appear in inputs, or letters.
- Option-order bias. Listing labels in the same order in every prompt hides a positional preference. Shuffle order across paraphrases in the harness.
- Tuning on the test set. Picking the best of nine prompts on the same examples you report is selection bias. Keep a held-out split untouched until the end.
- Silent model changes. A hosted model updated behind the same name moves accuracy and calibration. Pin versions and re-run the harness.
- Zero-shot chain of thought on a scored task. Adding 'think step by step' means the label is no longer the immediate continuation, so label scoring stops making sense. Score the label after the reasoning, or use path A for that variant; see chain-of-thought prompting.
Trade-offs
Zero-shot costs the fewest tokens and has no example set to curate or let drift. In exchange it leaves the model's priors in charge, so it is fragile on unusual labels and house-specific definitions. Label scoring buys confidence values and calibration at the price of log-probability access and closed label sets; generate-and-parse is universal but noisier. Calibration costs a few forward passes but adds a moving part to re-measure on every model change.
What to do next
- Label 200 to 300 real inputs, stratified by class, and lock away a third as a test set.
- Check whether your model or API exposes token log-probabilities; if it does, implement the scorer above and test it on an obvious case.
- Write three paraphrases and three input formats, and run the sensitivity harness with and without calibration.
- If the spread is wide, stop editing words: add demonstrations, split the task or try a stronger model.
- Choose a prompt from the middle of the good cells, confirm it once on the held-out set, and set an abstention threshold on the top label's probability.
- Add the harness to your evaluation suite so every model or prompt change re-runs it.