Labels are usually the most expensive input to a supervised model. Active learning is the idea that the model should help choose which examples a human labels next, so that each label teaches it as much as possible. Instead of labeling a random 20,000 support tickets, you label 2,000 at random, train, let the model point at the tickets it is most confused about, label those, and repeat. When it works, the same accuracy arrives with a fraction of the labels. When it is built carelessly, it produces a biased training set, an evaluation number that cannot be trusted and a model that is worse than random sampling would have given you.

This article treats active learning as a system rather than a formula. It walks through the components of the loop, derives the common query strategies from first principles, works a small example by hand, gives a runnable loop, and then spends most of its time on what goes wrong in production and how to decide when to stop. For the wider labeling operation (tooling, guidelines, workforce, quality control) read data labeling; for the cousin technique that uses unlabeled data without asking a human, read semi-supervised learning.

Advertisement

The loop and its components

Unlabeled pool100k to 10M itemsEmbed + scorecurrent modelBatch selectoruncertainty + diversityLabeling queuehumans, guidelines, QALabel storeversioned, with provenanceRetrainseed + acquired labelsModel registryround N checkpointrescoreRandom eval setnever actively chosenStop / continuelearning curve, budgetSeed setrandom, stratifiedOne round: score the pool, pick a diverse batch of informative items, label them, retrain, measure on a random held-out set, decide whether to continue.The evaluation set and the seed set are drawn at random; only the training additions are chosen by the model.
Pool-based active learning: the model scores the unlabeled pool, a selector picks a diverse batch, humans label it, the model is retrained and measured on a random held-out set.

Almost every practical system is pool-based: you already hold a large unlabeled pool (logs, documents, images) and choose from it. Two other settings exist. In stream-based selection, items arrive one at a time and you decide immediately whether to request a label, which suits moderation queues and sensors. In membership query synthesis, the learner generates its own inputs to be labeled, which rarely works for natural data because humans cannot label nonsense. The rest of this article assumes the pool setting.

Seed set. A random, ideally stratified sample labeled before any model exists. It gives the first model something to learn from and anchors the training distribution. A few hundred to a few thousand items is typical; too small and the first scores are noise.

Scorer. The current model scores every pool item with an acquisition function. For large pools, cache embeddings from a frozen encoder so rescoring is a cheap pass over vectors rather than a full forward pass over raw inputs.

Batch selector. Humans label in batches, so the system picks a batch of items, not one. Picking the top-k by score alone tends to pick k near-duplicates, which is why the selector combines informativeness with diversity.

Labeling queue and label store. Labels arrive asynchronously, with disagreement and errors. The store records who labeled what, under which guideline version, and in which round the item was acquired, so you can later separate actively chosen data from random data. Version it like code; see data versioning.

Random evaluation set. A fixed held-out set drawn uniformly from the pool and labeled once. It is the only honest measure of progress, because every actively chosen item is, by construction, unrepresentative.

Query strategies from first principles

An acquisition function answers one question: if I learned the label of x, how much would the model improve? Nobody can compute that directly, so strategies use proxies. Write p(y | x) for the model's predicted class probabilities and y1, y2 for the top two classes.

Least confidence scores 1 - p(y1 | x). Margin scores p(y1 | x) - p(y2 | x) and prefers the smallest gap, so it looks at items near a decision boundary. Entropy scores the sum over classes of -p log p, which uses the whole distribution and favours items where many classes are plausible. With two classes all three rank items identically; with many classes margin is usually the most robust starting point, because entropy can be dominated by a long tail of tiny probabilities.

Query by committee trains several models (different seeds, bootstraps or architectures) and picks items where they disagree, measured by vote entropy or average KL divergence from the consensus. BALD (Bayesian active learning by disagreement) formalises this: it scores the mutual information between the label and the model parameters, approximated with Monte Carlo dropout or a deep ensemble. Both target epistemic uncertainty, what the model does not know, rather than aleatoric uncertainty, items that are genuinely ambiguous. That distinction matters: a plain entropy score cannot tell a hard-but-learnable item from a hopelessly ambiguous one and will keep asking about the latter.

Diversity and representativeness. Core-set selection picks items that cover the embedding space, using greedy k-center: repeatedly add the pool point farthest from everything already labeled. It ignores the model's confusion and is strong in the earliest rounds. Hybrids combine both: shortlist the most uncertain items and cluster the shortlist, or use BADGE, which runs k-means++ seeding over gradient embeddings so the batch is both uncertain and spread out.

StrategyNeedsGood atWeak at
Least confidence / margin / entropyOne model with probabilitiesCheap, strong baselineRedundant batches, outliers, ambiguous items
Committee / BALDEnsemble or MC dropoutTargets what the model does not knowSeveral times the scoring cost
Core-set (k-center)EmbeddingsCold start, coverageIgnores the decision boundary
Uncertainty + clustering, BADGEEmbeddings and probabilitiesLarge batchesMore moving parts to tune
RandomNothingUnbiased, the baseline to beatWastes labels on easy items

Whatever you choose, the scores are only as meaningful as the probabilities behind them. Deep networks are often overconfident, and an overconfident model reports small uncertainty on exactly the items it gets wrong. Temperature scaling on the seed set is cheap insurance; model calibration covers how to measure and fix it.

Advertisement

A worked example by hand

Suppose an intent classifier for support tickets has three classes: billing, bug and account. The numbers here are illustrative. The model returns these probabilities for four pool items:

Itemp(billing)p(bug)p(account)MarginEntropy (nats)
A0.960.030.010.930.19
B0.480.450.070.030.90
C0.400.310.290.091.09
D0.500.490.010.010.74

Margin ranks D, B, C, A: D and B sit right on the billing-versus-bug boundary. Entropy ranks C first, because C spreads mass over all three classes. Which is more useful depends on why C is spread out. If C is a ticket that genuinely mixes topics, labeling it teaches little and the guideline should say how to handle mixed tickets. If C is a new product area the model has never seen, it is precisely the item you want, and a committee would show high disagreement on it. A is useless to label under any strategy.

Now scale up. With a pool of 200,000 tickets, a 2,000-item random seed, a budget of 6,000 further labels and batches of 500, there are twelve rounds. If B and D are representative of a thousand nearly identical "refund for duplicate charge" tickets, top-500-by-margin will spend much of a round on one confusion. Clustering the 5,000 most uncertain items into 500 groups and taking one per group spends the round on 500 different confusions. That single change is often the difference between active learning beating random sampling and losing to it.

A runnable loop

The loop below uses fixed feature vectors (for text, embeddings from a frozen sentence encoder) and a logistic regression, which retrains in seconds and gives usable probabilities. The selector is the hybrid just described: margin shortlist, then k-means for diversity, one item per cluster. The oracle stands in for your labeling queue.

import numpy as np
from sklearn.cluster import KMeans
from sklearn.linear_model import LogisticRegression

def margin_scores(probs):
    """Small margin between the top two classes = the model is torn = informative."""
    top2 = np.sort(probs, axis=1)[:, -2:]
    return top2[:, 1] - top2[:, 0]            # lower is more uncertain

def select_batch(model, X_pool, pool_idx, batch_size, oversample=10, seed=0):
    """Uncertainty pre-filter, then k-means for diversity inside the shortlist."""
    probs = model.predict_proba(X_pool[pool_idx])
    m = margin_scores(probs)
    k = min(len(pool_idx), batch_size * oversample)
    short = pool_idx[np.argsort(m)[:k]]       # the k most uncertain items
    km = KMeans(n_clusters=batch_size, n_init=3, random_state=seed).fit(X_pool[short])
    chosen = []
    for c in range(batch_size):               # one item per cluster: closest to centroid
        members = short[km.labels_ == c]
        if len(members) == 0:
            continue
        d = np.linalg.norm(X_pool[members] - km.cluster_centers_[c], axis=1)
        chosen.append(members[np.argmin(d)])
    return np.array(chosen)

def active_learning(X_pool, oracle, X_eval, y_eval, seed_idx, rounds, batch_size):
    labeled = list(seed_idx)
    y = {i: oracle(i) for i in labeled}       # oracle = your labeling queue
    history = []
    for r in range(rounds):
        model = LogisticRegression(max_iter=2000, class_weight="balanced")
        model.fit(X_pool[labeled], [y[i] for i in labeled])
        acc = model.score(X_eval, y_eval)     # random, fixed, never acquired
        history.append((len(labeled), acc))
        pool_idx = np.setdiff1d(np.arange(len(X_pool)), labeled)
        batch = select_batch(model, X_pool, pool_idx, batch_size, seed=r)
        for i in batch:
            y[i] = oracle(i)                  # in production: async, with QA
        labeled.extend(batch.tolist())
    return model, history

Run it twice on the same seed and eval set: once with select_batch and once with a random batch. Plot accuracy against labels for both. That comparison, not the absolute accuracy, tells you whether active learning is paying for itself on your data. Repeat with a few seeds; differences between strategies are often smaller than the run-to-run noise in the first rounds.

Engineering the loop at production scale

Scoring cost. Rescoring ten million items every round with a large model is the dominant compute cost. Freeze an encoder, store embeddings once, and train only a light head between rounds; or rescore a random subsample of the pool each round. Approximate nearest-neighbour indexes make k-center and clustering tractable at that scale.

Asynchrony. Labelers never finish a batch at once. Keep a small buffer of selected-but-unlabeled items, do not reselect items already in the queue, and allow retraining on a partial batch. A selector that blocks on the slowest labeler idles your team.

Deduplication. Near-duplicate removal before the loop starts (hashing, embedding similarity) saves budget no matter which strategy you use.

Label quality. Uncertain items are also the items humans disagree on. Route them to two labelers, measure agreement per round, and expect it to fall as the loop progresses; a sharp drop usually means the guideline has a gap.

Provenance. Store the acquisition round, strategy and score with each label. Without that you cannot later reweight, audit or rebuild the training set for a different model.

Failure modes

  • Sampling bias. The labeled set over-represents boundary regions and under-represents easy, common cases. A model trained only on it can be miscalibrated on the real distribution. Keep the random seed in the training set, mix a fraction of random items into every batch (10 to 20 percent is a common choice), and consider importance weighting.
  • Cold start. With few labels, uncertainty scores are noise and the loop can lock onto one region. Use diversity-only selection or random sampling for the first rounds.
  • Outlier and noise seeking. Corrupted, off-topic or unanswerable items look maximally uncertain forever. Filter them, add an "invalid" label, and monitor what fraction of each batch labelers reject.
  • Missed classes. Uncertainty only looks between classes the model already knows. A class absent from the seed set may never be queried. Diversity sampling and periodic random batches are the guard.
  • Model coupling. A set chosen by one model is not necessarily good for another. If you plan to switch architectures, expect the actively chosen data to transfer worse than its size suggests.
  • Evaluation leakage. Reporting accuracy on a held-out slice of the actively acquired data measures performance on hard items and tells you nothing about deployment. Only the random eval set counts.
  • Drift. If the pool is a snapshot and production moves, the loop optimises for the past. Refresh the pool and watch drift signals to trigger new rounds.

When to stop

Stop when one of three things happens. The learning curve on the random eval set flattens: two or three consecutive rounds each add less than your minimum useful gain. Predictions stabilise: the fraction of a fixed reference set whose predicted label changes between rounds drops near zero, which needs no extra labels to measure. Or the budget runs out. Write the stopping rule down before round one, together with the random-sampling baseline you will compare against; otherwise the loop continues because it exists.

Trade-offs and when not to bother

Active learning has a fixed cost: a scoring pipeline, a queue that can take model-chosen items, provenance tracking and a random eval set. It pays off when labels are expensive (experts, long documents, medical images), the pool is large and redundant, and the model can be retrained quickly. It pays off poorly when labels are cheap, when the pool is small enough to label outright, or when a strong pretrained model with a few hundred random labels already meets the target. Large pretrained models have shifted the economics: with good embeddings, random sampling is a stronger baseline than it used to be, so measure before assuming a gain. Combining active learning with pseudo-labeling the confident items is often stronger than either alone.

What to do next

  1. Label a random, stratified seed set and a separate random evaluation set; freeze the evaluation set.
  2. Cache embeddings for the whole pool and deduplicate near-identical items.
  3. Implement margin-shortlist plus clustering selection, and mix 10 to 20 percent random items into each batch.
  4. Calibrate the model on the seed set before trusting its uncertainty.
  5. Record round, strategy and score with every label in a versioned store.
  6. Run the loop side by side with random sampling for a few rounds and plot both learning curves.
  7. Write the stopping rule and budget down before starting, and review label agreement each round.
Key takeaway: Active learning is a loop, not a formula: score the pool, pick a diverse batch of informative items, label, retrain and measure on a random held-out set. Margin plus clustering is a strong default; calibration, random mixing and provenance keep the training set honest. Prove the gain against random sampling on your own data, and stop by a rule written before you started.