A large share of production LLM traffic is classification in disguise: route this ticket, tag this message's intent, decide whether this comment breaks policy, pick which tool should handle this request. A frontier model does these well from a prompt, and that is the right way to start. It is rarely the right way to stay, because classification has a fixed output space, a stable definition and high volume, which is exactly the shape where a model a hundred times smaller, trained on the task, matches the large one at a fraction of the cost and latency.

This article walks through the decision: which small architecture to use, how to get training labels from the LLM you already have, how to train, how to calibrate scores so you can abstain on uncertain cases, and how to serve the result as the first stage of a cascade. The running example is routing support tickets into twelve queues.

Advertisement

Three ways to classify with a small model

Fine-tuned encoder. A bidirectional encoder reads the whole input and a linear head maps a pooled representation to class logits. This is the classic BERT recipe, and modern encoders have made it better: ModernBERT, released by Answer.AI and LightOn, comes in base (149M parameters) and large (395M) sizes with an 8,192-token context, long enough for most tickets and emails. Encoders are the cheapest option per request and train in minutes to hours on one GPU.

Decoder SLM with a classification head. A small causal model of a few hundred million to a few billion parameters, with its language-modelling head replaced by a classification head on the last token's hidden state, fine-tuned fully or with LoRA. It brings more world knowledge and better handling of instructions embedded in inputs, at several times the encoder's cost.

Prompted SLM scored by label likelihood. No training: put the label definitions in a prompt and compare the model's log-probability of each label string. Useful for cold starts and for label sets that change weekly, but less accurate than a fine-tuned model once you have a few thousand examples.

ApproachData neededCost per requestBest when
Fine-tuned encoderA few thousand labelled examplesLowest; CPU is often enoughStable labels, high volume, latency matters
Decoder SLM + headThousands, LoRA is fineMedium; GPU preferredInputs need reasoning or world knowledge
Prompted SLM scoringLabel definitions onlyMedium, grows with prompt lengthCold start, frequently changing labels

Pipeline overview

Classification with a small model: LLM-labelled data, a calibrated student, and a cascade for the uncertain tailRaw traffictickets, messagesLLM labellerrubric + label setHuman auditgold set, disagreementsTraining setdeduped, balancedFine-tuneencoder or SLMCalibratetemperature, thresholdsSmall modelserves every requestConfidentact on labelUncertainLLM or humanp above thresholdp below thresholdnew labelscorrectionsfeed retraining
An LLM labels traffic against a rubric, humans audit a gold set, the small model is trained and calibrated, and requests below a confidence threshold escalate to the LLM or a person, whose answers feed retraining.

The key idea is the cascade at the bottom. The small model does not have to be right about everything; it has to be right about the cases it is confident about and honest about the rest. If 85% of tickets clear the threshold at accuracy comparable to the LLM, the LLM bill falls by 85% and its quality is preserved on the hard tail.

Advertisement

Worked example: the ticket router

Suppose a support platform receives 400,000 tickets a month and an LLM prompt assigns each to one of twelve queues: billing, refunds, shipping, returns, account access, security, product defect, feature request, enterprise sales, partner, legal and general. The prompt carries the queue definitions and a dozen examples, so each call sends a couple of thousand input tokens to produce one short label. The numbers below are illustrative; substitute your own.

The team labels 600 tickets by hand as a gold set, 50 per queue, oversampling the rare legal and security queues. The LLM agrees with the humans on 91% of them; most disagreements sit between refunds and returns, where the definitions overlap. They rewrite those two definitions, the agreement rises, and the LLM labels 40,000 deduplicated historical tickets. A ModernBERT-base classifier trained on those labels reaches roughly the teacher's accuracy on the gold set after three epochs.

Calibrated, the student keeps about 85% of traffic above its threshold with accuracy on that slice slightly above the teacher's overall figure, because the confident cases are the easy ones. The remaining 15% go to the original LLM prompt. LLM calls fall from 400,000 to about 60,000 a month, median routing latency drops from the LLM's response time to a few milliseconds for most tickets, and a single small inference server handles the student. The hard tail still gets the LLM's judgement, and every escalated ticket becomes a candidate training example for the next retrain.

Getting labels from the LLM you already have

Most teams have no labelled data, but they have an LLM prompt that classifies acceptably. Use it as a teacher. Run it over a few tens of thousands of real, deduplicated inputs with a rubric that defines each label, gives a boundary example between confusable pairs, and allows an unclear answer. Ask for the label only, through structured output, so parsing never fails.

Then build a gold set by hand: a few hundred examples labelled by people who know the domain, stratified so rare classes are present. The gold set is your source of truth. Measure the teacher against it first: if the LLM is 88% accurate, a student trained on its labels will inherit its mistakes, and you need to know that ceiling before blaming the student. Where teacher and humans disagree systematically, fix the rubric and relabel. More on teacher-student recipes is in distilling from an LLM.

Training an encoder

With Hugging Face Transformers, an encoder classifier is a short script. Class weights in the loss counter imbalance, which matters when the billing queue is ten times the size of the security queue.

import numpy as np, torch
from datasets import load_dataset
from sklearn.metrics import f1_score
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
                          DataCollatorWithPadding, Trainer, TrainingArguments)

MODEL = "answerdotai/ModernBERT-base"
ds = load_dataset("json", data_files={"train": "train.jsonl", "val": "val.jsonl"})
labels = sorted(set(ds["train"]["label"]))
l2i = {l: i for i, l in enumerate(labels)}

tok = AutoTokenizer.from_pretrained(MODEL)
def enc(b):
    out = tok(b["text"], truncation=True, max_length=512)
    out["labels"] = [l2i[l] for l in b["label"]]
    return out
ds = ds.map(enc, batched=True, remove_columns=["text", "label"])

counts = np.bincount(ds["train"]["labels"], minlength=len(labels))
weights = torch.tensor(counts.sum() / (len(counts) * np.maximum(counts, 1)), dtype=torch.float)

class WeightedTrainer(Trainer):
    def compute_loss(self, model, inputs, return_outputs=False, **kw):
        y = inputs.pop("labels")
        out = model(**inputs)
        loss = torch.nn.functional.cross_entropy(out.logits, y, weight=weights.to(out.logits.device))
        return (loss, out) if return_outputs else loss

model = AutoModelForSequenceClassification.from_pretrained(
    MODEL, num_labels=len(labels), id2label=dict(enumerate(labels)), label2id=l2i)

args = TrainingArguments("out", learning_rate=3e-5, num_train_epochs=3,
                         per_device_train_batch_size=32, eval_strategy="epoch",
                         save_strategy="epoch", load_best_model_at_end=True,
                         metric_for_best_model="macro_f1", bf16=True)

def metrics(p):
    return {"macro_f1": f1_score(p.label_ids, p.predictions.argmax(-1), average="macro")}

WeightedTrainer(model=model, args=args, train_dataset=ds["train"], eval_dataset=ds["val"],
                data_collator=DataCollatorWithPadding(tok), compute_metrics=metrics).train()

Select the checkpoint on macro-F1, not accuracy: with imbalanced classes, accuracy rewards a model that ignores the small queues. Keep the gold set out of training and validation entirely; it is for the final report. For decoder models, the same script works with a causal checkpoint loaded through the sequence-classification class, usually with LoRA adapters to keep memory down; see LoRA for small models and fine-tuning small models.

Prompted scoring done correctly

If you score a prompted decoder, compare the probability of each complete label, not the first token. Labels that share a first token, such as billing_refund and billing_invoice, are indistinguishable at the first token, and multi-token labels are penalised for length unless you account for it.

import torch

@torch.no_grad()
def label_logprobs(model, tok, prompt, labels):
    """Sum of token log-probabilities of each full label after the prompt."""
    p_ids = tok(prompt, return_tensors="pt").input_ids
    scores = {}
    for lab in labels:
        l_ids = tok(lab, add_special_tokens=False, return_tensors="pt").input_ids
        ids = torch.cat([p_ids, l_ids], dim=1).to(model.device)
        logp = model(ids).logits.log_softmax(-1)[0]
        # logits at position t predict token t+1
        tgt = ids[0, p_ids.shape[1]:]
        pos = torch.arange(p_ids.shape[1] - 1, ids.shape[1] - 1, device=ids.device)
        scores[lab] = logp[pos, tgt].sum().item()
    return scores

The function sums log-probabilities, so it favours shorter labels; either keep labels at similar token lengths or compare the mean per-token log-probability. Tokenise each label with the leading space the model would naturally produce after the prompt (" billing"), or tokenise prompt and label together and split at the prompt length, because tokenising them separately can change how the boundary splits.

Two refinements: give labels short, distinct surface forms (or single-letter codes with definitions in the prompt), and reuse the prompt's KV cache across labels instead of re-running it, which turns twelve forward passes into one plus twelve short ones. Constraining generation to the label set achieves similar results; see guided decoding.

Calibration and the abstain threshold

Fine-tuned networks are often overconfident: a softmax of 0.97 does not mean 97% of such predictions are right. The cascade depends on the score meaning something, so calibrate. Temperature scaling fits a single scalar on validation logits and leaves the ranking unchanged:

import torch

def fit_temperature(logits, labels):
    t = torch.nn.Parameter(torch.ones(1))
    opt = torch.optim.LBFGS([t], lr=0.1, max_iter=100)
    def step():
        opt.zero_grad()
        loss = torch.nn.functional.cross_entropy(logits / t, labels)
        loss.backward()
        return loss
    opt.step(step)
    return t.item()

# at serving time
probs = (logits / T).softmax(-1)
conf, pred = probs.max(-1)
route = "small_model" if conf >= THRESHOLD else "escalate"

Choose the threshold from a coverage-accuracy curve on held-out data: for each candidate threshold, plot the fraction of traffic the small model keeps against its accuracy on that fraction. Pick the point where accuracy on kept traffic matches the teacher, and price the remainder at LLM cost. Thresholds can be per class: a mistaken route to the security queue costs more than a mistaken route to general enquiries.

Evaluation that survives production

  • Macro-F1 and a confusion matrix on the human gold set. The matrix shows which pairs of classes the model confuses, which usually points to a rubric problem rather than a model problem.
  • Slices. Score by input length, language, channel and customer tier. A model that is excellent on English emails and poor on short chat messages averages out to a misleading number.
  • Teacher agreement on fresh traffic. Weekly, run a sample through both models; falling agreement is the earliest drift signal you will get.
  • Escalation rate. If the share below threshold rises, inputs have shifted or a new category has appeared.

The general method is in evaluating small models.

Serving

An encoder classifier is a single forward pass with no generation, so throughput comes from batching and truncation. Set a maximum input length from the length distribution of real inputs rather than the model's maximum; attention cost grows with length. Many teams export to ONNX or another optimised runtime and serve on CPU, which is often adequate for base-size encoders at moderate rates; measure on your hardware before buying GPUs. Log the input hash, model version, calibrated confidence and routing decision for every request so you can rebuild any decision later and harvest escalations as new training data.

Treat the classifier as a versioned artefact. Retrain on a schedule, monthly is common, from the previous training set plus audited escalations; evaluate the candidate on the frozen gold set and on a recent slice; and roll it out in shadow mode first, logging its decisions beside the current model's without acting on them. Promote only if macro-F1 holds and per-class recall on the critical queues does not fall. Refit the temperature and recheck thresholds with every new model, because calibration does not transfer between checkpoints.

Failure modes

FailureSymptomFix
Teacher ceilingStudent matches teacher, both wrong on the same casesMeasure the teacher on the gold set; fix the rubric
Rare class collapseHigh accuracy, near-zero recall on small queuesClass weights, oversampling, macro-F1 selection
OverconfidenceCascade keeps wrong predictionsTemperature scaling; per-class thresholds
Label driftNew topics forced into old classesTrack escalation rate; add an other class; retrain on schedule
Truncation lossLong inputs misclassifiedRaise max length or classify a summary
First-token scoringPrompted SLM always picks one of two similar labelsScore full label sequences

What to do next

  1. Find the LLM calls in your system that return one of a fixed set of labels, and rank them by volume.
  2. Label a stratified gold set of a few hundred examples by hand and measure the current LLM against it.
  3. Have the LLM label tens of thousands of real inputs with a rubric, and audit disagreements.
  4. Fine-tune a base-size encoder, select on macro-F1, and compare with the teacher on the gold set.
  5. Calibrate with temperature scaling and choose thresholds from the coverage-accuracy curve.
  6. Deploy as the first stage of a cascade, log every decision, and feed escalations back into training.
Key takeaway: Classification is the easiest place to replace an LLM with a small model, because the output space is fixed and the volume is high. Use the LLM you have as a labeller, keep a human gold set as ground truth, fine-tune an encoder first and a small decoder only if inputs need more knowledge, calibrate so confidence means something, and deploy the small model as the first stage of a cascade that escalates the uncertain tail.