Phi from Microsoft, Qwen from Alibaba and Gemma from Google are the three open-weight small model families most teams shortlist when they want a model that runs on one GPU, a laptop or a phone. Comparisons of them usually arrive as leaderboard tables, and those tables age within weeks, mix up model versions and say nothing about whether a model fits your memory, your licence review, your runtime or your task.
This article compares the families on the properties that decide real deployments: what is actually published and under which licence, how many bytes a checkpoint really occupies, how each architecture spends memory as the context grows, how the prompt formats and thinking modes differ, and which runtime risks come with each. It deliberately gives no benchmark scores. Instead it ends with a bake-off harness that ranks the shortlist on your own data, which is the only ranking that matters. Model facts were checked against the official model cards on 2026-10-01; check them again before you decide, because these families release often.
What compared should mean
A useful comparison separates hard constraints from preferences. Hard constraints are pass or fail: the licence must be acceptable to your legal team, the model must accept the input modalities you have, your inference runtime must support the architecture, and weights plus KV cache at your context length must fit the memory you have. Only the checkpoints that pass all four are worth evaluating, and the evaluation must use your prompts, your output format and your latency targets.
Public benchmarks are a weak signal for this purpose. They use fixed question sets newer models may have seen in training, vendor-chosen prompts, and say nothing about valid JSON rate or latency.
The three families as published today
| Phi (Microsoft) | Qwen3.5 (Alibaba) | Gemma 4 (Google) | |
|---|---|---|---|
| Small checkpoints named here | Phi-4-mini-instruct, 3.8B dense; phi-4, 14B dense | Qwen3.5-4B and Qwen3.5-9B, plus smaller sizes | E2B and E4B; also 12B, 26B A4B (mixture of experts) and 31B |
| Licence on the card | MIT | Apache 2.0 | Apache 2.0 (Gemma 3 and earlier used Google terms of use) |
| Context | 128K for Phi-4-mini; 16K for phi-4 | 262,144 tokens native | 128K for E2B and E4B; 256K for larger sizes |
| Input modalities | Text for Phi-4-mini; separate multimodal Phi variants exist | Text, image and video | Text and image on all sizes; audio on E2B, E4B and 12B |
| Architecture note | Dense decoder, grouped-query attention, tied input and output embeddings, 200,064-token vocabulary | Hybrid: Gated DeltaNet linear-attention layers plus gated full-attention layers | E2B and E4B use per-layer embeddings; parameters quoted as effective |
| Reasoning mode | Separate reasoning checkpoints | Thinking on by default; switch off with enable_thinking | Configurable; enabled by a <|think|> token at the start of the system prompt |
Two things in this table change decisions. Gemma moved to Apache 2.0 with version 4, which removes the custom-terms review that earlier Gemma versions needed; Phi has been MIT for some time. And the Qwen3.5 small models are not plain transformers, which matters for both memory and runtime support, as the next sections show. Whatever the card says, the licence that binds you is the LICENSE file in the exact repository you download, including community quantizations, which can carry different terms from the original.
Effective parameters versus loaded bytes
The weight memory of a model is roughly its parameter count times bytes per parameter: about 2 bytes at 16-bit, about 1 byte at 8-bit and a little over 0.5 bytes at 4-bit once quantization scales are included. Phi-4-mini at 3.8 billion parameters is therefore around 7.6 GB at 16-bit and roughly 2 to 2.3 GB at 4-bit. Measure the file you actually use.
Gemma 4 small models need care here. The card describes E2B as 2.3 billion effective parameters and 5.1 billion with embeddings, and E4B as 4.5 billion effective and 8 billion with embeddings. The difference is per-layer embeddings: large lookup tables consulted by each layer rather than matrices multiplied on every token. Effective parameters describe compute per token. Loaded bytes depend on whether your runtime keeps those tables in fast memory or can stream them from slower storage, so the honest planning number for E2B at 4-bit is anywhere from about 1.2 GB to about 2.6 GB of weights depending on the runtime. Measure peak resident memory rather than trusting either label. The same caution applies to the 26B A4B model: only about 4 billion parameters are active per token, but all 26 billion must be loaded. The quantization trade-offs themselves are covered in SLM quantization.
How each architecture spends memory as context grows
Weights are a fixed cost. The KV cache grows with every token in the context, and at long contexts it dominates. For a standard transformer the cache costs, per token, 2 (keys and values) times layers times KV heads times head dimension times bytes per element. Phi-4-mini has 32 layers, 8 KV heads and a head dimension of 128, so at 16-bit each token costs 2 x 32 x 8 x 128 x 2 = 131,072 bytes, which is 128 KiB. A 32K-token context needs 4 GiB of cache and the advertised 128K needs 16 GiB, several times the 4-bit weights. The context a model supports and the context you can afford are different numbers.
Hybrid architectures change the equation. In Qwen3.5 the Gated DeltaNet layers keep a fixed-size recurrent state per sequence instead of a cache that grows, and only the gated full-attention layers keep a conventional KV cache. Per the 4B card, those layers use 4 KV heads of dimension 256, which is 4 KiB per token per attention layer at 16-bit. Long contexts are far cheaper, if your runtime implements that state efficiently. Read the layer layout from the checkpoint config rather than guessing; the script below does the arithmetic and is how you should compare candidates at your context length. The mechanics of the cache are in the KV cache article.
import json, sys
def kv_bytes_per_token(config_path, kv_bytes=2):
cfg = json.load(open(config_path))
cfg = cfg.get("text_config", cfg) # multimodal checkpoints nest the LLM config
layers = cfg["num_hidden_layers"]
kv_heads = cfg.get("num_key_value_heads", cfg["num_attention_heads"])
head_dim = cfg.get("head_dim") or cfg["hidden_size"] // cfg["num_attention_heads"]
# Hybrid models list a type per layer; only full-attention layers keep a cache that grows
# with context. Field names vary by family, so print the config and check before trusting.
types = cfg.get("layer_types")
if types:
layers = sum(1 for t in types if "full" in str(t))
return 2 * layers * kv_heads * head_dim * kv_bytes # K and V, 16-bit by default
if __name__ == "__main__":
per_tok = kv_bytes_per_token(sys.argv[1])
for ctx in (4096, 32768, 131072):
print(f"{ctx:>7} tokens: {per_tok * ctx / 2**30:6.2f} GiB of KV cache")Run it against each candidate's config.json. For Gemma, inspect the config for local sliding-window layers: earlier Gemma generations interleaved sliding-window and global attention layers, and a sliding-window layer caps its cache at the window size. If your runtime ignores such layouts and allocates a full cache for every layer, your measured memory will be far higher than the arithmetic.
Prompt formats and thinking modes
Each family was trained on its own chat format, and a model prompted in the wrong one degrades quietly rather than failing. Phi-4-mini uses <|system|>, <|user|> and <|assistant|> markers with <|end|> closing each turn. Qwen uses a ChatML-style format with <|im_start|> and <|im_end|>. Gemma has its own turn markers. Never build these strings by hand; use the tokenizer's chat template, and inspect its output once so you know what the model sees.
from transformers import AutoTokenizer
messages = [
{"role": "system", "content": "Classify the ticket. Reply with JSON only."},
{"role": "user", "content": "My invoice shows the wrong VAT number."},
]
for model_id in ["microsoft/Phi-4-mini-instruct", "Qwen/Qwen3.5-4B"]:
tok = AutoTokenizer.from_pretrained(model_id)
prompt = tok.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True,
enable_thinking=False, # extra kwargs reach the template; Qwen reads this one
)
print(model_id, repr(prompt[-200:])) # inspect the tail: the turn markers differ per familyThinking modes need a decision, not a default. A Qwen3.5 model with thinking on writes a <think> block before the answer, which can multiply output tokens and therefore latency, and breaks a parser that expects JSON as the first character. For extraction and classification, turn thinking off and measure; for multi-step reasoning, turn it on and measure again, counting the thinking tokens in your latency and cost. With Gemma 4, thinking is controlled by the <|think|> token in the system prompt. Phi ships reasoning as separate checkpoints, so the choice happens at download time.
Runtime support is a real risk
A model you cannot run efficiently in your serving stack is not a candidate. Dense decoders like Phi-4-mini are supported almost everywhere on day one. New architectural pieces take time: at release, the Qwen3.5 card recommended the main branches of transformers, vLLM and SGLang, and per-layer embeddings and audio towers in Gemma small models need explicit runtime support to get their memory benefits. Before shortlisting, check that your exact runtime version loads the checkpoint in your quantization format and matches the reference implementation on a few prompts.
For on-device targets the bar is higher, because mobile runtimes lag server ones and memory budgets are tight; SLMs on mobile in 2026 covers that budget.
Running the bake-off
Build a held-out set of 200 to 500 real examples from your task, labelled by people, never used in prompt development. Serve each shortlisted checkpoint at the quantization you would ship, behind an OpenAI-compatible endpoint so one client drives them all, with identical prompts and temperature 0. Measure task accuracy, valid output rate, latency at the median and 95th percentile, output tokens per call and peak memory on the target hardware.
# Bake-off against OpenAI-compatible local servers (vLLM, llama.cpp server and others expose one).
import json, statistics, time, httpx
CANDIDATES = { # name -> base URL; vLLM rejects a "model" that differs from --served-model-name
"phi-4-mini-q4": "http://localhost:8001/v1",
"qwen3.5-4b-q4": "http://localhost:8002/v1",
"gemma-4-e4b-q4": "http://localhost:8003/v1",
}
SYSTEM = "Classify the support ticket into one label from the list. Reply with JSON {\"label\": ...}."
def run(name, base, examples):
ok = valid = 0; lat = []; out_tokens = 0
for ex in examples: # ex = {"text": ..., "label": ...}
t0 = time.perf_counter()
r = httpx.post(f"{base}/chat/completions", timeout=120, json={
"model": name, "temperature": 0, "max_tokens": 64,
"messages": [{"role": "system", "content": SYSTEM},
{"role": "user", "content": ex["text"]}]}).json()
lat.append(time.perf_counter() - t0)
out_tokens += r["usage"]["completion_tokens"]
try:
pred = json.loads(r["choices"][0]["message"]["content"])["label"]
valid += 1; ok += pred == ex["label"]
except (json.JSONDecodeError, KeyError, TypeError):
pass # invalid output counts as wrong
n = len(examples); lat.sort()
return {"accuracy": ok / n, "valid_json": valid / n, "p50_s": statistics.median(lat),
"p95_s": lat[int(0.95 * (n - 1))], "tokens_out_per_call": out_tokens / n}
examples = [json.loads(l) for l in open("tickets_holdout.jsonl")]
for name, base in CANDIDATES.items():
print(name, run(name, base, examples))A worked example of the decision: a team classifying support tickets into twelve labels on 8 GB laptops filters by licence (all three pass), modality (text only, all pass), runtime (their llama.cpp build must load the checkpoint) and memory (4-bit weights plus a 4K context must stay under about 4 GB). That leaves Phi-4-mini, Qwen3.5-4B and Gemma 4 E4B, all at 4-bit. They then run the harness above with thinking off. The rule they write down before seeing results: highest accuracy wins unless its valid JSON rate is under 99 percent or its p95 latency is over the product budget, in which case the next model up wins. The broader evaluation discipline is in SLM evaluation.
Failure modes
- Wrong chat template. A hand-built prompt in another family's format produces fluent but worse answers. Always apply the tokenizer template and diff its output when you change runtimes.
- Thinking leaks into structured output. A reasoning block before the JSON breaks parsers. Disable thinking for extraction or strip the block before parsing, and count it in latency either way.
- Context on paper, not in memory. A 128K context window is meaningless if the KV cache at 128K exceeds your RAM. Plan with the arithmetic and cap the context in the server.
- Effective parameter confusion. Budgeting Gemma E2B as a 2B model and then finding roughly 10 GB resident at 16-bit. Budget loaded bytes, measured.
- Quantization damage. Aggressive 4-bit quantization tends to hurt less common languages and long reasoning most. Evaluate at the precision you ship.
- Tool-calling format mismatch. Each family formats function calls differently; SLMs for tool calling covers the patterns that survive the differences.
Trade-offs summarized
Phi-4-mini is the most conventional of the three: a dense transformer with a permissive MIT licence and universal runtime support, with a large vocabulary that helps multilingual tokenization but costs logit memory, and a KV cache that grows quickly with context. Qwen3.5 small models bring long native context with much cheaper memory per token and built-in vision, at the price of a newer architecture that needs recent runtimes and a thinking mode you must manage. Gemma 4 small models bring audio and image input and Apache 2.0 terms, with parameter counts that require care when budgeting memory. None of this is a ranking; the bake-off settles that. If you plan to fine-tune Phi afterwards, Fine-tuning Phi-4 covers its specific traps.
What to do next
- Write down your hard constraints: licence policy, input modalities, runtime and version, target hardware, maximum context and latency budget.
- Open the current model cards and LICENSE files for the candidate sizes; record the exact repository and revision you evaluate.
- Run the KV arithmetic script on each config.json and add the weight estimate; drop any candidate that does not fit at your context length.
- Load each survivor in your actual runtime and check outputs against the reference implementation on ten prompts.
- Build a held-out set of at least 200 labelled examples and write the decision rule before running anything.
- Run the bake-off at the precision you will ship, with thinking settings chosen deliberately, and measure peak memory on the target device.
- Keep the harness and data; rerun it whenever a family releases a new version, because the answer will change.