Active learning chooses which unlabeled examples to send to annotators, so that each label teaches the model as much as possible. For an LLM task, such as classifying support tickets, judging answers or rating outputs for preference tuning, the labels come from domain experts and cost real money. The selection step has a cost of its own: before every round, a model has to run over a large pool of candidates to decide which are worth a human's time. At pool sizes in the millions, that scoring pass is a GPU workload that needs a budget like any other.
This page treats active learning for LLMs as a GPU engineering problem. It covers acquisition signals that work with LLM outputs, including the bias of entropy computed from truncated logprobs, cheap ensembles built from LoRA adapters, a two-stage uncertainty-then-diversity selector that runs on GPU, the arithmetic of rescoring the pool each round, scorer staleness, and the failure modes that make an active learner worse than random sampling. The general loop and classic query strategies are covered in the companion page on active learning architecture. The inference mechanics of scoring are covered in the page on LLM data labeling.
The round, and where GPU time goes
A round has five steps. Draw candidates from the pool. Score them with the current model. Select a batch of B examples. Get labels. Retrain. Three of those steps use GPUs: scoring, embedding and fine-tuning. The one that grows with the pool is scoring. Fine-tuning grows with the labeled set, which is small by design. Embeddings are computed once per prompt and cached, because they come from a frozen encoder and do not change when the task model does.
That asymmetry decides the architecture. Keep embeddings in a vector store keyed by prompt hash. Treat scores as a cache keyed by (prompt hash, scorer version), which is invalidated wholesale every time the adapters change. Then decide each round how much of the pool can be rescored.
Acquisition signals that work with LLM outputs
Classic uncertainty measures assume a classifier that outputs a full probability vector. An LLM gives you token logprobs, which is close but not the same. Four signals work in practice.
- Label margin for constrained outputs. When the answer is one of a few labels, prompt for a single label token, read the logprobs of each label's first token, renormalize over the labels and take the gap between the top two. A small margin means the model is torn. Check that every label's first token is distinct after tokenization, including the leading space, or two labels will share a probability.
- Entropy from top-k logprobs. Inference engines return only the top k alternatives per position. vLLM's
SamplingParams(logprobs=k)returns up to k + 1 entries, because the sampled token is always included, and the engine caps k with itsmax_logprobssetting. Entropy over a truncated distribution is biased low, because the tail mass is missing. Either renormalize over the top k and accept the bias consistently, or add a tail term that treats the leftover probability 1 - Σp as a single outcome. The tail-term version ranks heavy-tailed, genuinely uncertain positions higher, which is usually what you want. - Length-normalized sequence likelihood. For free-form outputs, score the model's greedy answer by its mean token logprob. Do not use the raw sum: it penalizes long answers, so selection fills up with long prompts.
- Disagreement between adapters. The closest LLM analogue of BALD-style committee disagreement is an ensemble. Full-model ensembles are unaffordable, but K LoRA adapters trained from different seeds or data orders on the same base are cheap. A multi-adapter server can run them against one copy of the base weights. The score is the disagreement between their label distributions: the entropy of the mean minus the mean of the entropies. That measures what the models disagree about, rather than what is inherently ambiguous.
Monte Carlo dropout, the usual cheap Bayesian trick, works poorly for LLMs. Most modern LLMs are trained with little or no dropout, so turning it on at inference changes the model instead of sampling its uncertainty. Self-consistency, which samples the answer several times and measures disagreement, works but multiplies decode cost by the sample count. It suits small candidate sets and reasoning tasks, not million-prompt pools.
Scoring the candidates
Scoring is prefill-dominated: hundreds of prompt tokens in, one or a few tokens out. Everything that speeds up offline batch inference applies, including a shared instruction prefix that the prefix cache reuses, length-sorted batches and large batch sizes. Those mechanics are covered in detail in the labeling page, so this sketch only shows the scoring function. It uses vLLM's offline API with multi-LoRA. Check the field names against your installed version, because the output structure has changed between releases.
import math
from vllm import LLM, SamplingParams
from vllm.lora.request import LoRARequest
LABELS = [" billing", " outage", " account", " other"] # leading spaces matter
llm = LLM(model=BASE, enable_lora=True, max_loras=4, max_logprobs=20)
tok = llm.get_tokenizer()
label_ids = [tok.encode(l, add_special_tokens=False)[0] for l in LABELS]
assert len(set(label_ids)) == len(LABELS), "labels share a first token"
params = SamplingParams(max_tokens=1, temperature=0.0, logprobs=20)
def label_dists(prompts, adapter_name, adapter_id, adapter_path):
outs = llm.generate(prompts, params,
lora_request=LoRARequest(adapter_name, adapter_id, adapter_path))
dists = []
for o in outs:
top = o.outputs[0].logprobs[0] # {token_id: Logprob} at position 0
p = [math.exp(top[t].logprob) if t in top else 0.0 for t in label_ids]
mass = sum(p)
dists.append(([x / mass for x in p] if mass > 0 else None, mass))
return dists # mass < 0.5 means the model wants to answer off-label: flag itThe returned label mass is a signal on its own. When most of the probability goes to tokens that are not labels, the prompt is malformed, out of domain, or the model is ignoring the format. Route those to a separate review bucket instead of ranking them as uncertain. Otherwise they dominate every batch.
Selecting a diverse batch on the GPU
Pure top-B uncertainty picks near-duplicates: one confusing ticket template appearing 400 times yields 400 nearly identical labels. The standard remedy combines uncertainty with diversity. Core-set selection (Sener and Savarese, 2018) and BADGE (Ash et al., 2020) are the references. A simple, robust version that runs comfortably on one GPU has two stages: keep the top M candidates by uncertainty, with M between 5 and 10 times B, then choose B of them with greedy k-center over their embeddings, seeded with the already-labeled set so the batch also covers what earlier rounds missed.
import torch
def select_batch(emb, unc, labeled_emb, B, oversample=8):
"""emb: [N, d] candidate embeddings (normalized), unc: [N] uncertainty scores,
labeled_emb: [L, d]. Returns indices of B candidates."""
M = min(len(unc), oversample * B)
top = torch.topk(unc, M).indices
X = emb[top] # [M, d] on GPU
if len(labeled_emb):
# distance from each candidate to its nearest labeled example, chunked
mind = torch.cat([torch.cdist(X[i:i + 8192], labeled_emb).min(1).values
for i in range(0, M, 8192)])
else:
mind = torch.full((M,), float("inf"), device=X.device)
picked = []
for _ in range(B):
j = int(torch.argmax(mind)) # farthest from everything chosen
picked.append(j)
mind = torch.minimum(mind, torch.cdist(X, X[j:j + 1]).squeeze(1))
mind[j] = -1.0
return top[torch.tensor(picked, device=top.device)]The cost is O(M·B·d) for the greedy loop plus O(M·L·d) for the seed distances. With M = 40,000, B = 5,000, d = 768 and L = 50,000, that is a few TFLOPs, or seconds on a single GPU. Scoring is the expensive part. The per-step Python loop could be batched further, but at these sizes kernel launch overhead is not the bottleneck worth chasing.
Worked example: budgeting the rescoring pass
Make the scoring budget explicit. The numbers below are an illustrative assumption, not a benchmark. Substitute your own measured prefill throughput.
| Quantity | Value |
|---|---|
| Pool size N | 1,200,000 prompts |
| Mean prompt length, after the cached shared prefix | 350 tokens |
| Adapters in the ensemble K | 4 |
| Tokens per full rescoring | 1.2M × 350 × 4 ≈ 1.68B |
| Assumed measured throughput | 20,000 prefill tokens/s per GPU |
| GPU time per full rescoring | ≈ 84,000 s ≈ 23 GPU-hours |
At 23 GPU-hours per round, rescoring everything each round is usually wasteful. Three techniques cut it by an order of magnitude.
- Score a random candidate sample, not the pool. Rescoring a random 15% (180,000 prompts) drops the round to about 3.5 GPU-hours. You lose the global top of the uncertainty ranking. Because uncertain regions are rarely tiny, a random sample that is 30 to 40 times the batch size usually finds them.
- Refresh only what can have changed. Keep last round's scores. Rescore the candidates whose old score was in the top fraction, plus a fresh random slice. Examples the old model was confident about rarely become the most uncertain after one round of training. Prove that on your own data by rescoring a full sample once and measuring rank correlation.
- Score with fewer adapters first. Run one adapter over the sample, then all K only over its top 20%. The disagreement signal is needed only where the single-model score is already high.
Stale scores and retraining cadence
Every retrain makes cached scores stale. The question is how often to retrain. Retraining after every small batch keeps selection sharp but spends GPU time on fine-tuning and rescoring and slows annotators down while the batch builds. Retraining rarely turns later picks into near-random selections from an outdated model. A workable default is to make the batch large enough that one round of labeling takes about as long as one retrain plus rescore, so neither the GPUs nor the annotators wait. Use warm-start adapters, initialized from last round's weights, between periodic from-scratch retrains. Warm starts are cheaper, but from-scratch runs detect whether early, biased selections have locked the model in.
Keep the scorer and the training prompt template identical. A scorer that reads a slightly different system prompt from the one the deployed model uses measures uncertainty about the wrong task.
Failure modes
- Selecting noise. The most uncertain examples are often unlabelable: garbled text, empty tickets, prompts in a language the task does not cover. Filter by label mass and language ID, and let annotators mark items as skip, without counting skips as labels.
- Biased evaluation. An actively selected set over-represents hard cases, so accuracy on it understates real accuracy and drifts as selection changes. Keep a fixed random test set, labeled once, and always report against it.
- Sampling bias in training. The labeled set's class balance no longer matches production. Add a random slice, 10 to 20% of each batch, so calibration and priors stay anchored. That slice also measures whether active selection is beating random at all.
- Cold start. With few labels, uncertainty is noise. Label the first one or two batches by random or diversity-only selection.
- Tokenization traps. Labels whose first tokens collide, or labels that tokenize differently with and without a leading space, silently merge classes in the margin computation.
- Stale score caches. Scores reused across a scorer version change look valid and are not. Key the cache by adapter hash and fail closed on a mismatch.
Trade-offs
| Choice | Cheaper option | Better option | Switch when |
|---|---|---|---|
| Uncertainty signal | Single-model margin | K-adapter disagreement | Labels are noisy or ambiguous |
| Pool coverage | Random 10-20% sample | Full rescoring | Rare classes hide in small regions |
| Diversity | Top-B only | Top-M then k-center | Pool has many near-duplicates (almost always) |
| Retraining | Warm-start each round | Periodic from scratch | Every 3-5 rounds, or on drift |
| Stopping | Fixed label budget | Stop when the audit slice stops improving | Gains on the random test set flatten |
What to do next
- Before building the loop, run one experiment: label two equal batches, one random and one chosen by single-model margin, and compare test-set gains. If active selection does not win, stop.
- Cache embeddings once. Key scores by (prompt hash, adapter hash), and measure your real prefill throughput so the budget table uses your numbers.
- Implement label mass filtering and the top-M then k-center selector above, and set aside a fixed random test set before the first round.
- Add a random slice to every batch and plot active-versus-random gains per round.
- Introduce a K-adapter ensemble only once single-model margin plateaus, and score with it only over the single model's top candidates.
- Keep learning: the active learning loop and query strategies, running LLM labeling as batch inference, GPU data curation pipelines and how vLLM spends GPU memory.