Two years ago the small-model landscape was a short list: a few dense transformers between 1 and 8 billion parameters, several under custom licences, all running in the same few runtimes. By October 2026 it is wider and less uniform. Most leading families now ship under MIT or Apache 2.0, several have stopped being plain transformers, parameter counts have split into effective, active and loaded figures that differ by a factor of two or more, and many models carry a reasoning mode you must decide whether to use. A table of benchmark scores does not help you navigate that; the scores go stale in months and say little about whether a model fits your device, licence and runtime.
This article is a map you can act on. It defines the size tiers, places each major open family in them with facts taken from the model cards, explains the four architecture changes that alter deployment, covers licences and runtime risk, gives a script that classifies any checkpoint from its configuration, and works through a shortlist for one concrete product. For a deeper three-way comparison of Phi, Qwen and Gemma, including KV-cache arithmetic, read Phi, Qwen, Gemma: small models compared.
What counts as small, by tier
Small is defined by where the model runs, not by a parameter threshold. Four tiers capture the decisions that matter.
- Under 1B. Fits in a few hundred megabytes at 4-bit. Good for classification, routing, extraction into a fixed schema and on-device autocomplete. Weak at open-ended reasoning and broad knowledge.
- 1B to 4B. The phone and laptop tier: roughly 0.7 to 2.5 GB of weights at 4-bit. Good for summarising, structured output, tool calling with a small tool set and chat over provided context.
- 7B to 14B. One consumer GPU or a workstation. Noticeably better at multi-step instructions, code and multilingual work; usually the smallest tier worth trying for a general assistant.
- Small active, large total. Mixture-of-experts models that compute like a 4B model per token but must hold every expert in memory. Fast on a server with spare memory, poor on a phone.
Pick the tier from the deployment target first. A model that does not fit the target is not a candidate, however good its scores.
The families on one map
| Family | Small checkpoints | Licence | Context | Notes from the card |
|---|---|---|---|---|
| Phi (Microsoft) | Phi-4-mini-instruct 3.8B; phi-4 14B | MIT | 128K (mini) | dense; 200K-token vocabulary; released February 2025 |
| Qwen3.5 (Alibaba) | 0.8B, 2B, 4B, 9B | Apache 2.0 | 262,144 native | hybrid Gated DeltaNet and gated attention; text, image and video in; thinking on by default |
| Gemma 4 (Google) | E2B, E4B; also 12B, 26B A4B, 31B | Apache 2.0 | 128K (E2B, E4B) | per-layer embeddings; audio input on small sizes; first released April 2026 |
| Ministral 3 (Mistral) | 3B, 8B, 14B | Apache 2.0 | 256K | base, instruct and reasoning variants; vision encoder; December 2025 |
| SmolLM3 (Hugging Face) | 3B | Apache 2.0 | 64K native, 128K with YaRN | fully open: data mixture, configs and code published; 11.2T training tokens |
| Granite 4.0 (IBM) | Nano sizes and larger | Apache 2.0 | varies | hybrid Mamba-2 and transformer variants, plus plain transformer variants |
| Llama 3.2 (Meta) | 1B, 3B | Llama 3.2 Community License | 128K | September 2024; text only at these sizes; the widest runtime support |
Two things stand out. Licences have converged on permissive terms, which removes most legal friction; Llama is the main exception. And the newest families are the least conventional architecturally, which shifts risk from the licence to the runtime.
Four architecture shifts that change deployment
Hybrid attention. Qwen3.5's small models repeat a block of three Gated DeltaNet layers followed by one gated full-attention layer; the 4B card lists eight such blocks. A DeltaNet layer carries a constant-size state from token to token rather than appending keys and values for every token, so only one layer in four pays memory that grows with context. Granite 4.0's hybrid models do something similar with Mamba-2 layers. The catch is that your inference engine needs kernels for those layers, on your hardware, and quantization recipes for them are younger than those for attention.
Per-layer embeddings. The Gemma 4 card gives two figures for each small model: E2B is 2.3B effective but 5.1B counting embeddings, and E4B is 4.5B effective but 8B in total. The extra parameters are lookup tables consulted per layer. They cost little compute, so the effective number describes speed, but whether they cost fast memory depends on whether the runtime streams them from storage. Plan memory from measurement, not from the label.
Mixture of experts. Gemma 4 26B A4B activates about 3.8 billion parameters per token from about 25 billion in total. Per-token compute resembles a 4B dense model; memory resembles a 25B one. That is a good trade on a server and a bad one on a phone. Small MoE architecture explains routing and why active counts mislead.
Reasoning modes. Qwen3.5 thinks by default and is switched off with enable_thinking; SmolLM3 accepts /think and /no_think; Ministral 3 ships separate reasoning checkpoints; Gemma 4 uses a control token. A thinking model can emit hundreds of tokens before its answer, which multiplies latency and cost per request. Decide per task, and evaluate in the mode you will deploy.
Licences and openness are not the same thing
MIT and Apache 2.0 allow commercial use, modification and redistribution with notice requirements, and Apache 2.0 adds an explicit patent grant. The Llama 3.2 Community License allows commercial use but adds an acceptable-use policy, attribution terms and a separate licence requirement for companies above 700 million monthly active users; read it rather than assuming it is open source.
Openness is a second axis. Most families publish weights only. SmolLM3 publishes its training data mixture, configurations, code and intermediate checkpoints, which matters if you need to audit what the model saw, reproduce training or continue pretraining from a known point. Read the licence file shipped inside the specific repository you pull weights from, because a third-party quantization or fine-tune may be published under terms that differ from its base model's.
Runtime support is the real compatibility matrix
A checkpoint is only useful if your engine runs it well: llama.cpp and GGUF on CPUs and laptops, MLX on Apple silicon, vLLM or SGLang on servers, ONNX Runtime, ExecuTorch or a vendor NPU stack on phones. Dense transformers in the Llama mould work almost everywhere. Hybrid layers, per-layer embeddings and new multimodal encoders arrive in each engine at different times and sometimes first without quantized or accelerated kernels. The cheap way to see what you are dealing with is to read the configuration before downloading weights:
import json, sys
from huggingface_hub import hf_hub_download # pip install huggingface_hub
def flags(cfg):
# Multimodal checkpoints often nest the language model config.
cfg = cfg.get("text_config", cfg)
keys = " ".join(cfg.keys()).lower()
layer_types = [str(t).lower() for t in cfg.get("layer_types", []) or []]
found = []
if any("expert" in k and cfg[k] for k in cfg):
found.append("MoE: load every expert, compute only the active ones")
if any("linear" in t or "mamba" in t or "deltanet" in t for t in layer_types) \
or "mamba" in keys or "ssm" in keys:
found.append("hybrid: recurrent layers need explicit runtime support")
if "per_layer" in keys:
found.append("per-layer embeddings: loaded bytes exceed 'effective' params")
window = cfg.get("sliding_window")
if window and cfg.get("use_sliding_window", True) and window < cfg.get("max_position_embeddings", float("inf")):
found.append("sliding-window attention on some layers")
return found or ["plain dense transformer (by these heuristics)"]
for repo in sys.argv[1:]:
# Gated repos such as Llama need `huggingface-cli login` first.
path = hf_hub_download(repo_id=repo, filename="config.json")
cfg = json.load(open(path))
print(repo, cfg.get("architectures"), cfg.get("model_type"))
for f in flags(cfg):
print(" -", f)Run it as python classify.py Qwen/Qwen3.5-4B microsoft/Phi-4-mini-instruct. The heuristics are deliberately simple, and configuration key names vary between architectures, so treat an unexpected result as a prompt to open the config and look. Anything flagged hybrid or per-layer goes on a list to test in your exact runtime build before it goes on a shortlist. For the GGUF path specifically, see GGUF runtime architecture.
Worked example: shortlisting for an offline field assistant
A company wants an assistant on technicians' Windows laptops: 16 GB of RAM, no discrete GPU, offline, English, German and French. It must answer questions over a retrieved 4,000-token chunk of a service manual and fill a JSON repair report. Typical prompts are 5,000 tokens; answers are short. The model must be usable commercially.
- Tier. CPU-only with other applications running leaves about 4 to 5 GB for the model, which means the 1B to 4B tier at 4-bit. The 7B to 14B tier and the MoE tier are out.
- Licence. All candidates in the tier pass. Llama 3.2 3B passes too but adds terms for legal review.
- Languages. SmolLM3 lists English, French and German among its six main languages; Phi-4-mini lists all three; Qwen3.5 and Gemma 4 advertise broad multilingual coverage. Keep all, but weight German in the evaluation.
- Runtime. The team uses llama.cpp. Phi-4-mini, SmolLM3, Ministral 3 3B and Llama 3.2 3B are dense and low-risk. Qwen3.5-4B and Gemma 4 E4B are flagged by the classifier and are tested in the team's pinned llama.cpp build before inclusion.
- Memory. A 3.8B model at roughly 0.55 to 0.6 bytes per parameter in a 4-bit format is about 2.1 to 2.3 GB of weights. A 5,000-token prompt on Phi-4-mini adds about 0.6 GB of 16-bit KV cache by the per-token arithmetic in the comparison article linked above. Both fit; measure peak resident memory to confirm.
- Shortlist. Phi-4-mini, SmolLM3-3B and Ministral 3 3B, plus Qwen3.5-4B if it passes the runtime test, with thinking switched off for latency.
- Bake-off. 300 real manual questions with reference answers, scored for groundedness and valid JSON, plus time to first token and tokens per second on the slowest laptop in the fleet.
Notice what decided the shortlist: memory, runtime and licence. Benchmark tables were never consulted. The evaluation method itself is covered in evaluating small models: common pitfalls.
Failure modes
- Choosing by leaderboard. A model ranked higher on public benchmarks loses on your prompts and format; the bake-off was skipped.
- Effective-parameter surprise. Memory planned from an effective or active count; the process is killed on the smallest devices.
- Runtime lag. A hybrid model runs, but on a slow fallback path without quantized kernels, at a fraction of the expected speed.
- Thinking left on. Default reasoning mode adds hundreds of hidden tokens per request and blows the latency budget.
- Licence drift. A community quantization or fine-tune carries different terms from the base model.
- Chat template mismatch. The runtime applies the wrong template, quality drops sharply and the model is wrongly judged weak.
- Stale map. A choice made on last year's families when a newer checkpoint in the same family fits better.
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| Dense, older family | runs everywhere, mature quantization | weaker at a given size |
| Hybrid attention | cheap long context | runtime and quantization risk |
| Per-layer embeddings | fast for its quality | memory depends on runtime |
| MoE small-active | fast per token | memory of the full model |
| Fully open data | auditable, reproducible | smaller choice of families |
| Reasoning mode on | better multi-step answers | latency and token cost |
What to do next
- Write down the deployment target, its free memory, the runtime and version, the languages and the licence constraint before naming any model.
- Pick the tier those constraints allow, and list every family with a checkpoint in it, re-checking each model card on the day.
- Run the classifier on each candidate and test anything flagged hybrid, MoE or per-layer in your exact runtime build.
- Measure peak resident memory for weights plus KV cache at your real prompt length on the weakest target device.
- Decide reasoning mode per task and evaluate in that mode with the model's own chat template.
- Run a bake-off of 200 to 500 real examples on two to four survivors, and pick on quality, valid-output rate and p95 latency together.
- Re-run the map every quarter; quantization guidance is in SLM quantization.