A small language model, or SLM, is a language model small enough to run where a large one cannot or should not: on a phone, a laptop, a single inexpensive GPU, or at a volume where a frontier model's price per token would sink the product. In the last few years such models have gone from toys to practical tools. In 2024 a 3.8B-parameter model, Phi-3-mini, reported about 69 percent on MMLU, a level that required models roughly an order of magnitude larger only a short time before.

This overview explains what makes a model small in a useful sense, how modern SLMs are built, the hardware arithmetic that determines where they run, and a clear account of the tasks on which they beat large models and the ones on which they do not. For the current model families see SLM Landscape 2026; for the cost case, The Economics of SLMs; and for training one from scratch, the SLM architecture and pipeline guide.

Advertisement

What counts as small

There is no standard cutoff. Various writers draw the line at 1B, 7B, 10B or 13B parameters, and the line moves as hardware improves. A more useful definition is operational: a model is small when it fits the memory and latency budget of a constrained target, typically a single consumer or data-center GPU, a laptop, or a phone, with room for its context. In current practice that means roughly a few hundred million to around ten to fifteen billion parameters, with phones at the low end and a single 24 GB GPU near the upper end.

Memory is the binding constraint, and it is easy to compute. Weights take parameters times bits divided by eight: a 3B model is about 6 GB in 16-bit and about 1.5 GB at 4-bit. The KV cache adds 2 x layers x KV heads x head dimension x tokens x bytes per value, which is why grouped-query attention, fewer KV heads than query heads, matters so much at small scale. Runtime buffers and the operating system take the rest. A phone with 8 to 12 GB of shared RAM can realistically host a 1B to 4B model at 4-bit alongside everything else it runs.

def weights_gb(params_b, bits):              # memory for weights alone
    return params_b * bits / 8

def kv_cache_gb(layers, kv_heads, head_dim, tokens, bytes_per=2):
    return 2 * layers * kv_heads * head_dim * tokens * bytes_per / 1e9

def decode_tokens_per_s_upper(params_b, bits, bandwidth_gb_s):
    # Batch-1 decoding reads every weight once per token, so bandwidth bounds speed.
    return bandwidth_gb_s / weights_gb(params_b, bits)

print(weights_gb(3, 4))                       # 1.5 GB for a 3B model at 4-bit
print(kv_cache_gb(28, 8, 128, 8192))          # ~0.94 GB for 8k tokens in FP16
print(decode_tokens_per_s_upper(3, 4, 100))   # ~67 tok/s ceiling at 100 GB/s
print(decode_tokens_per_s_upper(70, 4, 100))  # ~2.9 tok/s ceiling, same device

Why small models became good

Three changes, more than any architectural trick, explain the improvement.

Over-training. The Chinchilla scaling result suggests about 20 training tokens per parameter for the lowest training loss per unit of training compute. That optimizes the training bill, not the serving bill. When a model will answer billions of queries, it pays to train a smaller model for far longer. Gemma 2's 2.6B and 9B models were trained on more than 50 times the compute-optimal token count, the 9B on 8 trillion tokens; Llama 3.2's 1B and 3B saw up to 9 trillion. The loss keeps falling well past the Chinchilla point, just more slowly; see Chinchilla scaling.

Data quality. Microsoft's Phi series showed that heavily filtered web data plus synthetic textbook-style data lets a small model learn far more per token. Phi-3-mini, 3.8B parameters trained on 3.3 trillion tokens of such data, reported 69 percent on MMLU. The pretraining side is covered in filtering, dedup and mixture.

Teacher signal. A large model's full next-token distribution carries much more information than the single observed token. Gemma 2's 2.6B and 9B were trained by distillation from the 27B model, and Llama 3.2's 1B and 3B were derived by pruning a larger Llama and then trained with logits from Llama 3.1 8B and 70B as token-level targets. See distillation from an LLM.

import torch.nn.functional as F

def distill_loss(student_logits, teacher_logits, labels, T=1.0, alpha=0.5):
    # logits: [batch, seq, vocab]; average both terms per real token.
    mask = (labels != -100).flatten().float()
    s = F.log_softmax(student_logits.flatten(0, 1) / T, dim=-1)
    t = F.log_softmax(teacher_logits.flatten(0, 1) / T, dim=-1)
    # Soft targets: KL(teacher || student) per token, over the vocabulary.
    kl = (t.exp() * (t - s)).sum(-1)
    kd = (kl * mask).sum() / mask.sum() * T * T
    # Hard targets: ordinary next-token cross-entropy on the data.
    ce = F.cross_entropy(student_logits.flatten(0, 1), labels.flatten(), ignore_index=-100)
    return alpha * kd + (1 - alpha) * ce
Advertisement

Design levers inside a small model

At small scale, choices that are negligible in a 70B model become significant. The embedding table is the clearest: a 128,000-token vocabulary with hidden size 2,048 is about 262 million parameters, a quarter of a 1B model. Tying the input embedding and output projection saves that again, and several small model families do it. Vocabulary size is a trade: a larger vocabulary compresses text into fewer tokens, which speeds generation, but spends parameters that could have gone into layers.

Shape matters too. Meta's MobileLLM work on sub-billion models reported that deeper, thinner networks beat wide, shallow ones at the same parameter count, and combined that with embedding sharing and grouped-query attention. Grouped-query attention shrinks the KV cache several-fold, which matters most on devices where memory is shared with everything else. Context length is a budget choice: long context is cheap in parameters but costs KV-cache memory and prefill time at inference.

After pretraining, the same post-training steps used for large models apply: supervised fine-tuning, preference tuning such as DPO, and often task-specific fine-tuning. Small models respond strongly to fine-tuning because they have less general capacity to fall back on; a small model tuned on a few thousand examples of one task frequently outperforms a much larger general model prompted for it.

How a small language model is made and deployedCurated datafiltered web + syntheticOver-trainedfar past compute-optimalTeacher signaldistill / prunePost-trainingSFT, preference tuningQuantize8-bit or 4-bit weightsPhone / laptopNPU, GPU, CPUSingle server GPUhigh batch, low costCascadesmall first, large on doubtWhere it wins: narrow tasks, latency, privacy, volume, drafting for speculative decodingWhere it loses: long-tail knowledge, long multi-step reasoning, broad robustness
Curated data and long training produce a capable base; a teacher adds signal; post-training and quantization make it deployable on a device, a single GPU, or as the first stage of a cascade.

The hardware arithmetic of running small

Generating one token at batch size one reads every weight once, so decoding speed is bounded by memory bandwidth divided by model size in bytes. A laptop with about 100 GB/s of memory bandwidth has a ceiling of roughly 67 tokens per second for a 3B model at 4-bit and under 3 tokens per second for a 70B model at 4-bit, which would not fit in most laptops anyway. Real speeds are lower because of KV-cache reads and kernel overheads, but the ratio holds. On a data-center GPU, the same logic means a small model leaves most memory free for KV cache, so a server can batch many more concurrent requests, which is where the per-token cost advantage comes from.

Quantization is almost always part of the deployment. 8-bit weights are close to lossless for most models, and 4-bit weight-only methods such as AWQ and GPTQ usually lose a little quality, more so for the smallest models, which have less redundancy. Measure on your task rather than trusting average benchmark deltas; see SLM quantization and on-device inference.

Where small models beat large ones

  • Narrow, well-defined tasks. Classification, extraction, routing, redaction, and tool calling against a fixed schema. After fine-tuning on a few thousand labelled examples, a 1B to 8B model often matches or beats a prompted frontier model on the target metric; see SLMs for tool calling.
  • Latency. Time to first token and inter-token latency are lower, and there is no network round trip when the model runs locally, which matters for autocomplete, voice and interactive agents.
  • Privacy and offline operation. On-device inference keeps data on the device and works without connectivity.
  • Volume and cost. At millions of requests per day on a narrow task, a small model on your own GPUs is often much cheaper per request than a large one.
  • Supporting roles. As the draft model in speculative decoding, as a guard or classifier in front of a large model, and as the first stage of a cascade or router.

The cascade pattern captures most of the benefit without betting everything on the small model: serve each request with the small model, check cheap confidence signals, and escalate only uncertain requests to the large one.

def answer(request):
    out = small_model.generate(request.prompt, schema=request.schema, max_tokens=256)
    confident = (
        out.valid_json                              # parsed against the schema
        and out.mean_token_logprob > -0.3           # threshold tuned on a labelled set
        and not request.needs_long_reasoning        # routed by a cheap classifier
    )
    if confident:
        return out, "small"
    return large_model.generate(request.prompt, schema=request.schema), "large"
# Log which path served each request; if escalations climb, the traffic has shifted.

Where they fall short

Parameters store knowledge, so small models know less about rare facts and answer long-tail questions with fluent confabulation more often. They are weaker at long multi-step reasoning and at synthesizing information across very long contexts, even when their context window nominally allows it. They are more sensitive to prompt wording and format, and their multilingual quality falls off faster outside the main training languages. Retrieval helps with the knowledge gap by supplying facts in the prompt, but it does not fix reasoning depth. When a task needs broad world knowledge, open-ended reasoning, or robustness to unpredictable inputs, a large model is still the safer default.

These weaknesses are easy to miss in a demo and obvious in production. A small model that handles the twenty example prompts a team wrote by hand can fail on the long tail of real inputs: unusual phrasings, mixed languages, malformed documents, questions that combine two topics. The reliable way to find the boundary is to sample real traffic, have both a small and a large model answer it, and grade the disagreements. The share of requests where the small model is wrong, and how costly those errors are, decides whether it can serve alone, sit behind a confidence check, or should not be used for the task.

Choosing and evaluating a small model

Start from the constraint, not the leaderboard: the target device's memory and bandwidth, the latency budget, and the request volume define the largest model you can serve. Shortlist two or three models under that size with licenses you can use, then build an evaluation set from your real traffic, a few hundred labelled examples scored with the metric you care about. Public benchmark scores are a coarse filter at best; contamination and format differences make small gaps meaningless. Test quantized versions on the target runtime, because a model that wins at 16-bit may not win at 4-bit. If no candidate is good enough prompted, fine-tune the best one before moving up a size; SLM evaluation covers the details.

Failure modes

  • Confident confabulation. The model invents facts it does not store; ground answers with retrieval and validate outputs.
  • Quantization cliffs. A 4-bit build quietly loses accuracy on the one task you care about; always evaluate the exact deployed artifact.
  • Distribution shift. A model fine-tuned on last quarter's data degrades as traffic changes; monitor escalation rates and task metrics in production.
  • Format brittleness. Small changes in the prompt template break structured output; use constrained decoding and the model's own chat template.
  • Context overreach. Filling a long context window degrades answers well before the limit; retrieve less, but more relevant, text.
  • Device memory pressure. On phones the model competes with the OS and other apps, and gets evicted or throttled; budget for KV cache and runtime, not just weights.

Trade-offs

A small model trades breadth for efficiency. You gain latency, privacy, cost and control, including the ability to fine-tune and host it yourself, and you give up knowledge, reasoning depth and robustness. The practical architecture for many products is therefore not small or large but both: a small, tuned model on the common path, a large model behind it for hard cases, and measurements that show how traffic splits between them. As training recipes keep improving, the boundary of what a small model can handle keeps moving upward, so re-run the evaluation when a new generation of models appears.

Key takeaway: A small language model is best defined by where it must run: a phone, a laptop or a single GPU, which in practice means a few hundred million to around ten to fifteen billion parameters. Small models became capable through over-training far past compute-optimal, curated and synthetic data, and distillation or pruning from larger teachers. They beat large models on narrow fine-tuned tasks, latency, privacy and volume, and they lose on long-tail knowledge and deep reasoning, so evaluate on your own traffic with the quantized artifact you will deploy and consider a cascade that escalates hard requests to a large model.