Phi is Microsoft's family of small language models, and it started as an argument rather than a product. The argument was that the quality of training data matters more than its quantity: a model trained on carefully filtered and synthetic textbook-style text can reason far better than its size suggests. Three years and many releases later, Phi models are used on laptops, phones and edge servers, and the family includes general instruction models, multimodal models and dedicated reasoning models.
This article explains the idea behind the family, walks through the generations with figures taken from the model cards, maps each model to the jobs it suits and covers the practical details of running one: the reasoning output format, memory sizing and how to evaluate a Phi model without being misled. Fine-tuning and chat templates are covered in the Phi-4 fine-tuning guide, and the comparison with other small families in the Phi, Qwen and Gemma comparison.
The thesis: data quality over scale
The first model, phi-1, was a 1.3-billion-parameter model for Python code, introduced in June 2023 in a paper titled "Textbooks Are All You Need". Its training data combined web code filtered for educational value with synthetic textbooks and exercises written by a larger model. phi-1.5, in September 2023, kept the size and extended the recipe to general natural language: its card describes a 30-billion-token dataset, seen for 150 billion training tokens, drawn from the phi-1 sources plus synthetic NLP texts, and deliberately excluding generic web crawls such as Common Crawl. phi-2, in December 2023, grew to 2.7 billion parameters and 1.4 trillion training tokens, mixing synthetic data with web data filtered with the help of a larger model. phi-1.5 and phi-2 were released as base models without instruction tuning or RLHF.
The idea has a predictable consequence that is worth understanding before choosing a Phi model. Synthetic, textbook-style data teaches patterns of explanation and reasoning very efficiently, but a small model trained on curated text has limited room for broad world knowledge; this is a known trade-off of the approach. In practice, Phi models tend to do well on structured reasoning, mathematics and code relative to their size, and less well at recalling obscure facts. Design for that: pair them with retrieval when the application needs knowledge, and use them alone when it needs reasoning over supplied context.
The generations, with figures from the cards
| Model | Released | Parameters | Context | Notes |
|---|---|---|---|---|
| phi-1 | June 2023 | 1.3B | Python code; research | |
| phi-1.5 | Sept 2023 | 1.3B | Base model; no web crawl | |
| phi-2 | Dec 2023 | 2.7B | 2,048 | Base model; 1.4T tokens |
| Phi-3-mini | 2024 | 3.8B | 4K and 128K variants | First instruction-tuned generation; small 7B and medium 14B also released |
| Phi-3.5-MoE | Aug 2024 | 16 x 3.8B, about 42B total | 128K | 6.6B active with 2 experts |
| phi-4 | Dec 2024 | 14B | 16K | 9.8T training tokens |
| Phi-4-mini-instruct | Feb 2025 | 3.8B | 128K | General model; function calling |
| Phi-4-multimodal-instruct | Feb 2025 | 5.6B | 128K | Text, image and speech input |
| Phi-4-reasoning and -plus | Apr 30, 2025 | 14B | 32K | Fine-tuned from phi-4 on reasoning traces; plus adds RL |
| Phi-4-mini-reasoning | Apr 2025 | 3.8B | 128K | Math only, synthetic math data |
| Phi-4-mini-flash-reasoning | June 2025 | 3.8B | 64K | Hybrid SambaY architecture; math only |
| Phi-4-reasoning-vision-15B | Mar 4, 2026 | 15B | 16,384 | Vision reasoning; thinking and non-thinking modes |
Two notes on reading this table. The mini-reasoning and mini-flash-reasoning cards both say the models were designed and tested for mathematical reasoning only; they are not general assistants. And for Phi-4-mini-flash-reasoning, Microsoft reports up to ten times higher throughput than its predecessor in one specific setting, a 2K-token prompt with 32K tokens generated on vLLM; treat that as a long-generation result and measure your own workload.
Which Phi for which job
| Job | Start with | Why |
|---|---|---|
| Classification, extraction, routing on a laptop or edge box | Phi-4-mini-instruct | 3.8B fits in a few GB at 4-bit; long context |
| Higher-quality general assistant on one GPU | phi-4 | 14B, stronger reasoning; 16K context |
| Hard multi-step reasoning, math, science | Phi-4-reasoning or -plus | Trained to think first; needs long output budgets |
| Math tutoring on constrained hardware | Phi-4-mini-reasoning or mini-flash-reasoning | Math-specialised at 3.8B |
| Speech or image input, small footprint | Phi-4-multimodal-instruct | One 5.6B model for text, image and audio |
| Screenshots, GUIs, charts with reasoning | Phi-4-reasoning-vision-15B | Vision plus switchable thinking |
If two rows seem to fit, start with the smaller model and move up only when your evaluation shows the gap. The step from 3.8B to 14B roughly quadruples weight memory and lowers throughput accordingly, and for tasks like classification a fine-tuned mini model often matches the larger one.
The reasoning output contract
Phi-4-reasoning does not answer immediately. Its card describes a response with two parts: a chain of thought inside <think> tags, then a solution section. The card also specifies a system prompt, to be used with its chat template, and sampling settings: temperature 0.8, top-k 50, top-p 0.95, with up to 32,768 new tokens for complex queries. Those settings are part of the model's contract, not tuning suggestions: greedy decoding and short budgets are common reasons for reports of poor quality.
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL = "microsoft/Phi-4-reasoning"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype="auto", device_map="auto")
SYSTEM = open("phi4_reasoning_system_prompt.txt").read() # copy verbatim from the model card
def solve(question, max_new_tokens=32768):
messages = [{"role": "system", "content": SYSTEM},
{"role": "user", "content": question}]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=max_new_tokens, do_sample=True,
temperature=0.8, top_k=50, top_p=0.95) # card's recommended settings
new_tokens = out[0, ids.shape[1]:]
text = tok.decode(new_tokens, skip_special_tokens=True)
thinking, sep, answer = text.partition("</think>") # split on the closing tag only
if not sep:
# Budget ran out inside the thinking section: there is no answer to show.
return {"status": "truncated", "tokens": len(new_tokens), "answer": None}
return {"status": "ok", "tokens": len(new_tokens),
"thinking": thinking.replace("<think>", "").strip(), "answer": answer.strip()}The truncation branch matters in production. If the token budget runs out while the model is still thinking, there is no answer, and showing the partial reasoning to a user is worse than an honest failure. Log the token count of every response, alert when the truncation rate rises, and decide per route whether to raise the budget, retry once, or fall back to phi-4 for a direct answer. Strip the thinking section before storing conversation history, or each turn will carry thousands of tokens of old reasoning into the next prompt.
Phi-4-reasoning-vision-15B adds a choice: it can think inside <think> blocks for mathematical and scientific problems, or answer directly in a <nothink> mode for perception tasks such as captioning and locating interface elements. Use the direct mode for perception; reasoning about a caption only adds latency.
Sizing memory: weights and KV cache
Weight memory is parameters times bytes per parameter: phi-4 at 14 billion parameters is about 28 GB in bf16 and roughly 8 to 9 GB at 4-bit including quantization scales; Phi-4-mini at 3.8 billion is about 7.6 GB in bf16 and around 2 to 2.3 GB at 4-bit. The quantization guide explains the format choices.
The part people forget is the KV cache, which grows with every token of context, and reasoning models generate a lot of tokens. It can be computed from the model configuration:
def kv_cache_bytes(layers, kv_heads, head_dim, tokens, bytes_per_value=2, batch=1):
# 2 = one key and one value vector per layer per KV head per token
return 2 * layers * kv_heads * head_dim * bytes_per_value * tokens * batch
# phi-4 / Phi-4-reasoning: 40 layers, 10 KV heads, head_dim 5120/40 = 128
print(kv_cache_bytes(40, 10, 128, 32_768) / 2**30) # ~6.25 GiB at 32K tokens, bf16
# Phi-4-mini: 32 layers, 8 KV heads, head_dim 3072/24 = 128
print(kv_cache_bytes(32, 8, 128, 131_072) / 2**30) # ~16 GiB at the full 128K, bf16So a single phi-4 reasoning request that uses its full 32K window needs about 6 GB of KV cache on top of the weights, and Phi-4-mini at its full 128K context needs about 16 GB, more than twice its own bf16 weights. Set the serving engine's maximum model length to what your application actually uses rather than the model's maximum; for example, vLLM's --max-model-len flag caps it, which lets the engine fit more concurrent requests in the same memory.
Worked example: choosing for a field-service laptop app
A company wants an offline assistant on technicians' laptops with 16 GB of RAM and no discrete GPU. It must answer questions from equipment manuals, classify fault reports into 40 codes, and help with wiring calculations.
Phi-4-mini-instruct at 4-bit, around 2 to 2.3 GB, is the natural base. Manual questions are a knowledge problem, which is exactly where a small curated-data model is weak, so they go through local retrieval over the manuals and the model answers from retrieved passages. Fault classification is a narrow task: fine-tune the mini model, following the fine-tuning guide, and constrain output to the 40 labels. Wiring calculations are arithmetic; rather than adding a second, math-only model, give the assistant a calculator tool, which is cheaper and exact.
Context is capped at 8K tokens, which keeps the KV cache to about 1 GB and leaves memory for the operating system and the manuals index. phi-4 at 14B would not fit comfortably next to everything else in 16 GB, and the evaluation below showed it was not needed. The GGUF runtime guide covers running quantized models on CPU.
Evaluating a Phi model honestly
Published benchmark scores are a starting point, not a decision. Any model trained heavily on synthetic, exam-like data is particularly likely to look strong on exam-like benchmarks, so test on your own tasks, and keep separate scores for reasoning, factual recall and output-format validity, because Phi models often score very differently across them.
SETS = {
"reasoning": load_jsonl("evals/multi_step_tasks.jsonl"), # your task, with answers
"recall": load_jsonl("evals/closed_book_facts.jsonl"), # facts the app relies on
"format": load_jsonl("evals/output_contract.jsonl"), # JSON / label validity
}
def run(model_fn, sets=SETS):
report = {}
for name, items in sets.items():
ok = sum(grade(item, model_fn(item["prompt"])) for item in items)
report[name] = round(ok / len(items), 3)
return report
# Compare candidates on the same sets; a Phi model that wins "reasoning" but
# loses "recall" is telling you to add retrieval, not to pick a bigger model.
for name, fn in {"phi-4-mini": phi_mini, "phi-4": phi_14b, "incumbent": current}.items():
print(name, run(fn))A result where the Phi model wins on reasoning and loses on recall is useful information: it says add retrieval, not use a bigger model. A result where it fails the format set usually points to a chat-template or sampling mismatch, the most common deployment error with this family.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Confident wrong facts | Small curated-data model asked for recall | Retrieval; answer only from supplied context |
| Reasoning model returns nothing useful | Token budget ended inside thinking | Larger budget; detect truncation; fall back |
| Poor quality from a reasoning model | Greedy decoding or missing system prompt | Use the card's prompt and sampling settings |
| Math model fails at general chat | mini-reasoning used outside its scope | Route general tasks to Phi-4-mini-instruct |
| Out of memory under load | KV cache sized for maximum context | Cap max model length; budget per request |
| Garbled turns after an upgrade | Template differs between generations | Always use the tokenizer's chat template |
What to do next
- Write down the job: reasoning over supplied context, knowledge recall, math, or perception. That choice selects the row in the table above.
- Download the candidate from the official microsoft organisation on Hugging Face and read its card and LICENSE file.
- Build three small evaluation sets, reasoning, recall and format, from real tasks, and score the candidate against your current model.
- For reasoning models, use the card's system prompt and sampling settings, set a token budget and handle truncation explicitly.
- Compute weight and KV-cache memory for your context length, and cap the serving engine's maximum model length to match.
- If the task is narrow, read the distillation data recipe and fine-tune the mini model rather than moving to a larger one.