Most teams that try LLM data labeling start the same way: a loop that sends each item to a chat endpoint with a long instruction, parses a JSON answer, and writes it to a table. It works on a thousand rows. At two million rows it takes days, costs more than the human labeling it was meant to replace, and nobody can say how accurate it is.

This article treats labeling as what it is: an offline batch inference workload with an unusual shape, plus a measurement problem. The shape is a long shared prompt prefix, a short variable item, and an answer that should be one of a handful of labels. That shape decides how to use the GPU: cache the prefix, constrain the output, read the label probabilities instead of generating text, and batch aggressively. The measurement problem decides whether the labels are worth anything: a human-labeled gold set, agreement statistics, and calibrated thresholds that send uncertain items to a stronger model or a person. It does not cover generating new training data from scratch, which is a different workload with its own article.

The workload shape

A transformer request has two phases. Prefill processes every prompt token in parallel; it is compute-bound, and the GPU runs large matrix multiplications at high utilization. Decode generates one token per step per sequence; it is memory-bandwidth-bound, because each step reads all the weights and the KV cache to produce a single token per sequence, and only large batches amortize that read.

A classification-style labeling prompt is almost all prefill. A typical one has a 1,000 to 2,000 token prefix (task definition, label descriptions, a few worked examples) and a 200 to 500 token item. If the model answers with a label code, decode is one to three tokens. If you let it write a rationale and then a JSON object, decode can be 50 to 150 tokens, and the job becomes decode-bound for no gain in the label itself. The two biggest levers are therefore: do not recompute the shared prefix, and do not generate tokens you will throw away.

An LLM labeling job on GPUs: one shared prefix, many short items, a cheap labelRaw itemstickets, docs, rowsPrompt builderfixed prefix + itemvLLM engineprefix cache, batchingLabel + scoreconstrained, logprobsConfidentwrite labelUncertainbigger model or humanGold sethuman labels, stratifiedCalibrationthreshold per classp highp lowGPU cost is mostly prefilldecode is 1-3 tokens per item when the label is constrainedThe gold set sets the thresholds; the thresholds decide how much reaches humans.
The labeling pipeline. Everything left of the engine is CPU work; the engine, and sometimes a second larger model for the uncertain slice, are the GPU work.

Prefix caching: pay for the instructions once

vLLM's automatic prefix caching stores KV-cache blocks keyed by a hash of the tokens they contain and everything before them. When a new request starts with the same tokens, those blocks are reused and prefill runs only on the new suffix. It is an engine argument, enable_prefix_caching=True, and recent versions enable it by default for many models; check what yours does rather than assuming.

The cache only helps if the prefix is byte-identical and comes first. That has consequences for how you build prompts:

  • Put the instructions and few-shot examples first and the item last. A ticket id, timestamp or per-item hint at the top of the prompt makes every prefix unique and disables the cache.
  • Freeze the few-shot examples for the whole job. Rotating examples per item is a reasonable bias control, but it multiplies prefill cost; if you want it, use a small fixed set of prefixes and group items by prefix.
  • Apply the chat template once with the tokenizer and check that the rendered prefix is identical across items. Templates that inject the current date break caching silently.
  • Watch the engine's prefix cache hit rate metric. On a well-built labeling job it should be close to the prefix share of total prompt tokens; if it is near zero, something in the prefix varies.

Constrained labels and logprobs

The label should come from a closed set, and the model should not be able to answer outside it. vLLM's structured outputs support exactly that. In current releases you pass a StructuredOutputsParams object, with a choice list, a regex or a JSON schema, through the structured_outputs field of SamplingParams; older releases called the same feature GuidedDecodingParams and guided_decoding. The decoder masks any token that would leave the allowed set, so every output parses.

The second step is to stop treating the label as text. Ask for the top logprobs of the generated label token and you get a probability for each candidate label from the same forward pass. That number is what makes routing possible: it lets you auto-accept the confident items and escalate the rest. Use single-token label codes (A, B, C, or short words you have checked tokenize to one token) so the first token identifies the label. Labels like billing and billing_dispute share a first token, and their first-token probabilities are meaningless. Also check whether your engine version reports logprobs before or after the constraint mask is applied, since that changes how you normalize them.

import math
from vllm import LLM, SamplingParams
from vllm.sampling_params import StructuredOutputsParams

CODES = ["A", "B", "C", "D", "E", "F"]          # each verified to be one token
NAMES = dict(zip(CODES, ["billing", "outage", "account access",
                         "feature request", "cancellation", "other"]))

llm = LLM(model="Qwen/Qwen2.5-7B-Instruct", enable_prefix_caching=True,
          max_model_len=4096, gpu_memory_utilization=0.90)
tok = llm.get_tokenizer()

PREFIX = open("label_prompt.txt").read()        # task, label defs, 8 frozen examples

def render(item):
    msgs = [{"role": "system", "content": PREFIX},
            {"role": "user", "content": "Ticket:\n" + item + "\nAnswer with one letter."}]
    return tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)

params = SamplingParams(temperature=0.0, max_tokens=1, logprobs=20,
                        structured_outputs=StructuredOutputsParams(choice=CODES))

def label_batch(items):
    outs = llm.generate([render(x) for x in items], params)
    rows = []
    for o in outs:
        top = o.outputs[0].logprobs[0]            # {token_id: Logprob}
        p = {lp.decoded_token.strip(): math.exp(lp.logprob) for lp in top.values()}
        z = sum(p.get(c_, 0.0) for c_ in CODES) or 1.0
        probs = {c_: p.get(c_, 0.0) / z for c_ in CODES}   # renormalize over labels
        best = max(probs, key=probs.get)
        rows.append((NAMES[best], probs[best], probs))
    return rows

Pass large lists to generate (tens of thousands of items per call) and let the engine's continuous batching schedule them; a Python loop of single requests leaves the GPU mostly idle. Write results in chunks so a crash loses minutes, not hours, and key every output row by item id and prompt version so reruns are idempotent.

Worked example: budgeting 2 million items

Estimate cost before launching, with rates you measured on your own hardware and model. Run a benchmark of a few thousand real items and record prefill tokens per second and decode tokens per second at the batch sizes the engine actually reaches. The figures below assume 20,000 prefill tokens per second and 2,500 decode tokens per second per GPU for a 7B to 8B model. They are placeholders for the arithmetic, not a claim about any specific GPU.

The job: 2,000,000 support tickets, a 1,200-token prefix, items averaging 350 tokens.

  • Prefill without caching: 2,000,000 x 1,550 = 3.1 billion tokens; at 20,000 per second that is 155,000 seconds, about 43 GPU-hours.
  • Prefill with caching: only the 350-token suffix is computed, 0.7 billion tokens, about 9.7 GPU-hours.
  • Decode with an 80-token rationale: 160 million tokens at 2,500 per second, about 17.8 GPU-hours.
  • Decode with a constrained label: about 3 tokens including end-of-sequence, 6 million tokens, well under one GPU-hour.

Prefill and decode overlap inside the engine, so the real total is somewhat lower than the sum, but the ratio is the point: caching plus constrained labels turns a 60-hour job into roughly a 10-hour one, which spread over eight GPUs is a lunch break. Data-parallel replicas (one engine per GPU, items sharded by hash) scale almost linearly for a 7B model, because there is no communication between replicas. Reach for tensor parallelism only when the model does not fit on one GPU.

Where the GPU time goes for 2,000,000 items (illustrative rates, see text)Prefill, no prefix cache43.1 GPU-hoursPrefill, prefix cached9.7 GPU-hoursDecode, 80-token rationale17.8 GPU-hoursDecode, 3-token label0.7 GPU-hoursAssumed per-GPU rates: 20,000 prefill tokens/s and 2,500 decode tokens/s. Measure yours.
GPU-hours by phase for the worked example. Red bars are the naive pipeline; green bars are the same job with prefix caching and a one-token constrained label.

Measuring quality: gold sets, kappa, calibration and cascades

Throughput is the easy half. Before any LLM label goes into a training set, measure it against humans. Build a gold set of 1,000 to 2,000 items, stratified so rare classes have at least 50 examples each, labeled by two people with disagreements adjudicated. Then report, per class, precision, recall and the confusion matrix, and overall Cohen's kappa: (p_o - p_e) / (1 - p_e), where p_o is observed agreement and p_e is the agreement expected by chance from each rater's label frequencies. Compare the model's kappa with the human-human kappa on the same items. If people only agree at 0.7, asking the model for 0.95 is asking it to be consistent with noise.

Next, calibrate. Bucket the gold items by the model's top-label probability and check accuracy per bucket. Raw probabilities from instruction-tuned models are usually overconfident, so pick thresholds from the measured curve, per class if classes behave differently, rather than trusting 0.9 to mean 90 percent. Suppose the curve shows that items above 0.85 are 96 percent accurate and make up 82 percent of traffic. Then 82 percent of the job is labeled automatically at known quality, and 18 percent, 360,000 tickets here, goes to the next stage.

The next stage should be a cascade, not a crowd. Run the uncertain slice through a larger model, or the same model with self-consistency (several samples at non-zero temperature, majority vote). Its cost applies only to the slice, so a model ten times more expensive per item, run on 18 percent of items, adds about 1.8 times the first pass, not ten times. Send only what remains uncertain after that to people, and feed their labels back into the gold set.

Failure modes

Labeling jobs fail quietly. These are the patterns to watch for:

  • Position and label-order bias. Models favour labels listed first or used in the last few-shot example. Shuffle label order across a few prompt variants on the gold set and check whether predictions move.
  • Template drift. A tokenizer or chat-template update changes the rendered prompt. Labels shift and the cache hit rate collapses at the same moment. Pin the model revision and the template, and hash the rendered prefix into every output row.
  • Silent truncation. Items longer than max_model_len minus the prefix are cut or rejected. Count them, and label long items by a separate path rather than dropping them.
  • Distribution shift between gold and production. A gold set sampled last quarter misses new products. Refresh a slice of it every run and track per-class accuracy over time.
  • Feedback loops. Training a classifier on LLM labels, then using that classifier to choose which items humans review, compounds the LLM's systematic errors. Keep a random audit sample that humans label regardless of confidence.
  • Data handling. Self-hosting keeps sensitive text off third-party APIs, but output files and logs still contain it. Apply the same retention rules as the source data.

Trade-offs

ChoiceGainCost
Constrained one-token labelMinimal decode, always parses, gives probabilitiesNo rationale to audit
Rationale then labelAuditable, sometimes more accurate on hard itemsDecode-bound, 10 to 30 times the decode cost
Small model plus cascadeCheap bulk, quality where neededTwo models to operate and calibrate
Large model for everythingSimplest pipelineSeveral times the GPU cost
Fine-tuned small classifierFastest at inference, no promptNeeds labels first; often trained on LLM labels

What to do next

To put this into practice:

  1. Write the label definitions, then build and double-label a stratified gold set before touching a GPU.
  2. Benchmark prefill and decode rates for your model on a few thousand real items, and recompute the budget above with your numbers.
  3. Restructure the prompt so the prefix is fixed and first; confirm a high prefix cache hit rate.
  4. Switch to single-token label codes with structured outputs and logprobs; plot accuracy against confidence on the gold set and choose thresholds.
  5. Add a cascade for the uncertain slice and a random human audit sample, and version every row by model, template hash and prompt.
  6. Read more on running vLLM on GPUs, continuous batching, synthetic data generation for the generative counterpart of this workload, and alignment data formats for what happens to the labels next.
Key takeaway: LLM labeling is a prefill-bound batch job with a measurement problem attached. Put a frozen prefix first so the KV cache is reused, constrain the answer to single-token label codes and read their probabilities, batch large and shard across replicas, and budget with rates you measured. Then earn trust in the labels with a stratified gold set, kappa against human agreement, calibrated thresholds, a cascade for the uncertain slice and a permanent random audit.