Phi is Microsoft's family of small language models, and it started as an argument rather than a product. The argument was that the quality of training data matters more than its quantity: a model trained on carefully filtered and synthetic textbook-style text can reason far better than its size suggests. Three years and many releases later, Phi models are used on laptops, phones and edge servers, and the family includes general instruction models, multimodal models and dedicated reasoning models.

This article explains the idea behind the family, walks through the generations with figures taken from the model cards, maps each model to the jobs it suits and covers the practical details of running one: the reasoning output format, memory sizing and how to evaluate a Phi model without being misled. Fine-tuning and chat templates are covered in the Phi-4 fine-tuning guide, and the comparison with other small families in the Phi, Qwen and Gemma comparison.

Advertisement

The thesis: data quality over scale

The first model, phi-1, was a 1.3-billion-parameter model for Python code, introduced in June 2023 in a paper titled "Textbooks Are All You Need". Its training data combined web code filtered for educational value with synthetic textbooks and exercises written by a larger model. phi-1.5, in September 2023, kept the size and extended the recipe to general natural language: its card describes a 30-billion-token dataset, seen for 150 billion training tokens, drawn from the phi-1 sources plus synthetic NLP texts, and deliberately excluding generic web crawls such as Common Crawl. phi-2, in December 2023, grew to 2.7 billion parameters and 1.4 trillion training tokens, mixing synthetic data with web data filtered with the help of a larger model. phi-1.5 and phi-2 were released as base models without instruction tuning or RLHF.

The idea has a predictable consequence that is worth understanding before choosing a Phi model. Synthetic, textbook-style data teaches patterns of explanation and reasoning very efficiently, but a small model trained on curated text has limited room for broad world knowledge; this is a known trade-off of the approach. In practice, Phi models tend to do well on structured reasoning, mathematics and code relative to their size, and less well at recalling obscure facts. Design for that: pair them with retrieval when the application needs knowledge, and use them alone when it needs reasoning over supplied context.

The generations, with figures from the cards

The Phi family, 2023 to 2026: from code-only research models to reasoning and vision2023202420252026phi-11.3B, codephi-1.51.3B, basephi-22.7B, basePhi-3mini 3.8B to 14BPhi-3 visionimage inputPhi-3.5mini, MoE, visionphi-414B, Dec 2024Phi-4-mini3.8B, 128KPhi-4-multimodal5.6B, image, audioPhi-4-reasoning14B, plus RL variantmini-reasoning3.8B, math onlymini-flash-reasoning3.8B hybrid, mathreasoning-vision-15BMarch 2026Grey: research base models. Blue: first instruction-tuned generation. Green: general Phi-4 models. Yellow: reasoning models.Every model card on Hugging Face names its own licence; the current models are MIT, but check each repository you download.
The Phi lineage. Each later generation kept the data-first approach and added instruction tuning, long context, multimodality and dedicated reasoning training.
ModelReleasedParametersContextNotes
phi-1June 20231.3BPython code; research
phi-1.5Sept 20231.3BBase model; no web crawl
phi-2Dec 20232.7B2,048Base model; 1.4T tokens
Phi-3-mini20243.8B4K and 128K variantsFirst instruction-tuned generation; small 7B and medium 14B also released
Phi-3.5-MoEAug 202416 x 3.8B, about 42B total128K6.6B active with 2 experts
phi-4Dec 202414B16K9.8T training tokens
Phi-4-mini-instructFeb 20253.8B128KGeneral model; function calling
Phi-4-multimodal-instructFeb 20255.6B128KText, image and speech input
Phi-4-reasoning and -plusApr 30, 202514B32KFine-tuned from phi-4 on reasoning traces; plus adds RL
Phi-4-mini-reasoningApr 20253.8B128KMath only, synthetic math data
Phi-4-mini-flash-reasoningJune 20253.8B64KHybrid SambaY architecture; math only
Phi-4-reasoning-vision-15BMar 4, 202615B16,384Vision reasoning; thinking and non-thinking modes

Two notes on reading this table. The mini-reasoning and mini-flash-reasoning cards both say the models were designed and tested for mathematical reasoning only; they are not general assistants. And for Phi-4-mini-flash-reasoning, Microsoft reports up to ten times higher throughput than its predecessor in one specific setting, a 2K-token prompt with 32K tokens generated on vLLM; treat that as a long-generation result and measure your own workload.

Advertisement

Which Phi for which job

JobStart withWhy
Classification, extraction, routing on a laptop or edge boxPhi-4-mini-instruct3.8B fits in a few GB at 4-bit; long context
Higher-quality general assistant on one GPUphi-414B, stronger reasoning; 16K context
Hard multi-step reasoning, math, sciencePhi-4-reasoning or -plusTrained to think first; needs long output budgets
Math tutoring on constrained hardwarePhi-4-mini-reasoning or mini-flash-reasoningMath-specialised at 3.8B
Speech or image input, small footprintPhi-4-multimodal-instructOne 5.6B model for text, image and audio
Screenshots, GUIs, charts with reasoningPhi-4-reasoning-vision-15BVision plus switchable thinking

If two rows seem to fit, start with the smaller model and move up only when your evaluation shows the gap. The step from 3.8B to 14B roughly quadruples weight memory and lowers throughput accordingly, and for tasks like classification a fine-tuned mini model often matches the larger one.

The reasoning output contract

Phi-4-reasoning does not answer immediately. Its card describes a response with two parts: a chain of thought inside <think> tags, then a solution section. The card also specifies a system prompt, to be used with its chat template, and sampling settings: temperature 0.8, top-k 50, top-p 0.95, with up to 32,768 new tokens for complex queries. Those settings are part of the model's contract, not tuning suggestions: greedy decoding and short budgets are common reasons for reports of poor quality.

from transformers import AutoModelForCausalLM, AutoTokenizer

MODEL = "microsoft/Phi-4-reasoning"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype="auto", device_map="auto")

SYSTEM = open("phi4_reasoning_system_prompt.txt").read()   # copy verbatim from the model card

def solve(question, max_new_tokens=32768):
    messages = [{"role": "system", "content": SYSTEM},
                {"role": "user", "content": question}]
    ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
    out = model.generate(ids, max_new_tokens=max_new_tokens, do_sample=True,
                         temperature=0.8, top_k=50, top_p=0.95)      # card's recommended settings
    new_tokens = out[0, ids.shape[1]:]
    text = tok.decode(new_tokens, skip_special_tokens=True)
    thinking, sep, answer = text.partition("</think>")     # split on the closing tag only
    if not sep:
        # Budget ran out inside the thinking section: there is no answer to show.
        return {"status": "truncated", "tokens": len(new_tokens), "answer": None}
    return {"status": "ok", "tokens": len(new_tokens),
            "thinking": thinking.replace("<think>", "").strip(), "answer": answer.strip()}

The truncation branch matters in production. If the token budget runs out while the model is still thinking, there is no answer, and showing the partial reasoning to a user is worse than an honest failure. Log the token count of every response, alert when the truncation rate rises, and decide per route whether to raise the budget, retry once, or fall back to phi-4 for a direct answer. Strip the thinking section before storing conversation history, or each turn will carry thousands of tokens of old reasoning into the next prompt.

Phi-4-reasoning-vision-15B adds a choice: it can think inside <think> blocks for mathematical and scientific problems, or answer directly in a <nothink> mode for perception tasks such as captioning and locating interface elements. Use the direct mode for perception; reasoning about a caption only adds latency.

Sizing memory: weights and KV cache

Weight memory is parameters times bytes per parameter: phi-4 at 14 billion parameters is about 28 GB in bf16 and roughly 8 to 9 GB at 4-bit including quantization scales; Phi-4-mini at 3.8 billion is about 7.6 GB in bf16 and around 2 to 2.3 GB at 4-bit. The quantization guide explains the format choices.

The part people forget is the KV cache, which grows with every token of context, and reasoning models generate a lot of tokens. It can be computed from the model configuration:

def kv_cache_bytes(layers, kv_heads, head_dim, tokens, bytes_per_value=2, batch=1):
    # 2 = one key and one value vector per layer per KV head per token
    return 2 * layers * kv_heads * head_dim * bytes_per_value * tokens * batch

# phi-4 / Phi-4-reasoning: 40 layers, 10 KV heads, head_dim 5120/40 = 128
print(kv_cache_bytes(40, 10, 128, 32_768) / 2**30)    # ~6.25 GiB at 32K tokens, bf16

# Phi-4-mini: 32 layers, 8 KV heads, head_dim 3072/24 = 128
print(kv_cache_bytes(32, 8, 128, 131_072) / 2**30)    # ~16 GiB at the full 128K, bf16

So a single phi-4 reasoning request that uses its full 32K window needs about 6 GB of KV cache on top of the weights, and Phi-4-mini at its full 128K context needs about 16 GB, more than twice its own bf16 weights. Set the serving engine's maximum model length to what your application actually uses rather than the model's maximum; for example, vLLM's --max-model-len flag caps it, which lets the engine fit more concurrent requests in the same memory.

Worked example: choosing for a field-service laptop app

A company wants an offline assistant on technicians' laptops with 16 GB of RAM and no discrete GPU. It must answer questions from equipment manuals, classify fault reports into 40 codes, and help with wiring calculations.

Phi-4-mini-instruct at 4-bit, around 2 to 2.3 GB, is the natural base. Manual questions are a knowledge problem, which is exactly where a small curated-data model is weak, so they go through local retrieval over the manuals and the model answers from retrieved passages. Fault classification is a narrow task: fine-tune the mini model, following the fine-tuning guide, and constrain output to the 40 labels. Wiring calculations are arithmetic; rather than adding a second, math-only model, give the assistant a calculator tool, which is cheaper and exact.

Context is capped at 8K tokens, which keeps the KV cache to about 1 GB and leaves memory for the operating system and the manuals index. phi-4 at 14B would not fit comfortably next to everything else in 16 GB, and the evaluation below showed it was not needed. The GGUF runtime guide covers running quantized models on CPU.

Evaluating a Phi model honestly

Published benchmark scores are a starting point, not a decision. Any model trained heavily on synthetic, exam-like data is particularly likely to look strong on exam-like benchmarks, so test on your own tasks, and keep separate scores for reasoning, factual recall and output-format validity, because Phi models often score very differently across them.

SETS = {
    "reasoning": load_jsonl("evals/multi_step_tasks.jsonl"),    # your task, with answers
    "recall":    load_jsonl("evals/closed_book_facts.jsonl"),    # facts the app relies on
    "format":    load_jsonl("evals/output_contract.jsonl"),      # JSON / label validity
}

def run(model_fn, sets=SETS):
    report = {}
    for name, items in sets.items():
        ok = sum(grade(item, model_fn(item["prompt"])) for item in items)
        report[name] = round(ok / len(items), 3)
    return report

# Compare candidates on the same sets; a Phi model that wins "reasoning" but
# loses "recall" is telling you to add retrieval, not to pick a bigger model.
for name, fn in {"phi-4-mini": phi_mini, "phi-4": phi_14b, "incumbent": current}.items():
    print(name, run(fn))

A result where the Phi model wins on reasoning and loses on recall is useful information: it says add retrieval, not use a bigger model. A result where it fails the format set usually points to a chat-template or sampling mismatch, the most common deployment error with this family.

Failure modes

SymptomCauseFix
Confident wrong factsSmall curated-data model asked for recallRetrieval; answer only from supplied context
Reasoning model returns nothing usefulToken budget ended inside thinkingLarger budget; detect truncation; fall back
Poor quality from a reasoning modelGreedy decoding or missing system promptUse the card's prompt and sampling settings
Math model fails at general chatmini-reasoning used outside its scopeRoute general tasks to Phi-4-mini-instruct
Out of memory under loadKV cache sized for maximum contextCap max model length; budget per request
Garbled turns after an upgradeTemplate differs between generationsAlways use the tokenizer's chat template

What to do next

  1. Write down the job: reasoning over supplied context, knowledge recall, math, or perception. That choice selects the row in the table above.
  2. Download the candidate from the official microsoft organisation on Hugging Face and read its card and LICENSE file.
  3. Build three small evaluation sets, reasoning, recall and format, from real tasks, and score the candidate against your current model.
  4. For reasoning models, use the card's system prompt and sampling settings, set a token budget and handle truncation explicitly.
  5. Compute weight and KV-cache memory for your context length, and cap the serving engine's maximum model length to match.
  6. If the task is narrow, read the distillation data recipe and fine-tune the mini model rather than moving to a larger one.
Key takeaway: Phi models are built on a data-first thesis: curated and synthetic textbook-style training makes small models reason well, at the cost of limited stored knowledge. Choose by job: Phi-4-mini for edge and narrow tasks, phi-4 for a stronger general model, the reasoning variants for hard problems with long output budgets, the math-only models only for math, and the multimodal and vision models for images and audio. Pair them with retrieval for facts, honour the reasoning contract, size the KV cache, and evaluate on your own tasks.