Microsoft's Phi-4 models are small enough to fine-tune on one GPU, and they are easy to fine-tune badly. The two usual starting points differ in size, vocabulary, context handling and chat format, and both use fused projection layers that many generic LoRA recipes do not target. A run that ignores these details still reports a falling loss, then produces a model that never stops generating or has quietly learned nothing.
This article covers what is specific to Phi-4. It starts from the published configuration files, turns them into adapter and memory arithmetic, builds labels correctly for each chat format, and gives a complete LoRA loop. The general choice between full fine-tuning, adapters and preference tuning is covered in SLM fine-tuning, so here we assume you have already decided that supervised fine-tuning with an adapter is right.
Which Phi-4 you are fine-tuning
The name Phi-4 covers several models. Two are the usual fine-tuning bases. phi-4 is a 14-billion-parameter dense decoder released in December 2024, trained on 9.8 trillion tokens with a 16K-token context. Phi-4-mini-instruct is a 3.8-billion-parameter model released in February 2025, trained on 5 trillion tokens with a 128K-token context. Both are MIT-licensed. The family also includes multimodal and reasoning variants, whose training recipes differ, so check the model card of the exact checkpoint before reusing anything here.
Both models load as Phi3ForCausalLM with model_type: "phi3", so the same code path handles them. The configuration values that affect fine-tuning are these:
| Config value | phi-4 | Phi-4-mini-instruct | Why it matters for fine-tuning |
|---|---|---|---|
| hidden / layers | 5120 / 40 | 3072 / 32 | Adapter size and activation memory |
| attention heads / KV heads | 40 / 10 | 24 / 8 | Grouped-query attention: the fused qkv output is not 3 x hidden |
| vocab_size | 100,352 | 200,064 | Logit memory per token; mini's is twice as large |
| tie_word_embeddings | false | true | On mini, input embeddings and output head are one tensor |
| max_position_embeddings | 16,384 | 131,072 (LongRoPE, original 4,096) | Long-context behaviour depends on rope scaling |
| pad / eos token ids | 100349 / 100265 | 199999 / 199999 | Mini's pad id equals its EOS id |
Fused layers: what LoRA must target
Llama-style models have separate q_proj, k_proj and v_proj layers and separate gate_proj and up_proj layers. The Phi-3 implementation in transformers, which Phi-4 uses, fuses them. Each attention block has one qkv_proj and one o_proj; each MLP has one gate_up_proj, producing twice the intermediate size, and one down_proj.
This matters because many copied recipes list target_modules=["q_proj", "v_proj"]. Recent PEFT versions raise an error when none of the requested names exist in the model, but a list that mixes one valid name with several invalid ones trains only the valid one, and older tooling can fail silently. Always call print_trainable_parameters() and compare the count with the arithmetic below. On phi-4, qkv_proj maps 5,120 inputs to 7,680 outputs: 40 query heads of 128 dimensions plus 10 key and 10 value heads. A LoRA pair on a layer with d_in inputs and d_out outputs adds r(d_in + d_out) parameters, so it is easy to compute what you should see:
def lora_params(hidden, inter, heads, kv_heads, layers, r):
"""Trainable LoRA parameters when targeting all four Phi-3-style linear layers.
A LoRA pair on a d_in x d_out layer adds r * (d_in + d_out) parameters."""
head_dim = hidden // heads
qkv_out = heads * head_dim + 2 * kv_heads * head_dim # fused q, k and v
per_layer = (
r * (hidden + qkv_out) # qkv_proj
+ r * (hidden + hidden) # o_proj
+ r * (hidden + 2 * inter) # gate_up_proj (gate and up fused)
+ r * (inter + hidden) # down_proj
)
return per_layer * layers
# Shapes from each model's config.json
print(lora_params(5120, 17920, 40, 10, 40, 16)) # phi-4: 55,705,600
print(lora_params(3072, 8192, 24, 8, 32, 16)) # Phi-4-mini: 23,068,672
# Logits are often the largest single activation: seq_len x vocab x 4 bytes in fp32
print(4096 * 200064 * 4 / 1e9) # Phi-4-mini, one 4k sequence: ~3.3 GB
print(4096 * 100352 * 4 / 1e9) # phi-4, one 4k sequence: ~1.6 GBAt rank 16 on all four layer types, phi-4 trains about 55.7 million parameters, roughly 0.4 percent of the model, and Phi-4-mini trains about 23.1 million. If the printed count is far lower, some target names did not match. Note that an adapter on the fused qkv_proj adapts queries, keys and values together; adapting values alone would need custom code.
Memory arithmetic for both models
In bf16, phi-4's weights take about 28 GB and Phi-4-mini's about 7.6 GB. Full fine-tuning with AdamW in mixed precision needs roughly 16 bytes per parameter: about 224 GB for phi-4, meaning several sharded 80 GB GPUs, and about 61 GB for mini.
With LoRA the frozen base stays in bf16 and only the adapter carries gradients and optimizer state. The remaining costs are activations, which gradient checkpointing reduces, and logits: sequence length times vocabulary, often upcast to fp32 for the loss. For a 4,096-token sequence that is about 3.3 GB on Phi-4-mini, because of its 200,064-token vocabulary, versus 1.6 GB on phi-4. A small model with a large vocabulary can hit out-of-memory on logits first; use a shorter maximum length or a smaller micro-batch with gradient accumulation.
Rough planning figures: LoRA on Phi-4-mini fits a 24 GB GPU at moderate lengths; LoRA on phi-4 in bf16 needs 40 to 80 GB; QLoRA keeps the phi-4 base in 4-bit, around 8 GB, and makes a 24 GB GPU workable, with costs covered in QLoRA in depth. These are estimates, so measure peak memory over a few steps first.
Two chat formats, and why you never write them by hand
The two models use different turn markers. phi-4 uses <|im_start|>role<|im_sep|>content<|im_end|>, so a turn looks like <|im_start|>user<|im_sep|>Hello<|im_end|>. Phi-4-mini uses <|system|>…<|end|><|user|>…<|end|><|assistant|>, and defines tool calling by placing JSON tool definitions between <|tool|> and <|/tool|> inside the system turn.
Training on one format and serving with the other is a common, silent failure: the model has never seen the markers at inference time, so it drifts. Hand-written f-string templates cause the same drift through a missing newline or special token. Use the tokenizer's own apply_chat_template for both training and inference, and read the rendered string for one example before training. The general mechanics of chat templates are covered in chat templates in depth.
Labels, and the pad/EOS trap
Supervised fine-tuning should compute loss only on the assistant's reply. Loss on the prompt teaches the model to generate user messages and wastes capacity. The dependable way to build labels without relying on template-specific tags is to render the conversation twice: once up to the generation prompt, and once in full. The first rendering is a prefix of the second, and everything after it is the reply, including the token that closes the turn. That closing token must be in the labels, because it is how the model learns to stop.
Phi-4-mini adds a trap. Its config sets pad_token_id and eos_token_id to the same id, 199999. A common collator masks labels with labels[input_ids == pad_token_id] = -100. On mini, that also masks every real end-of-text token in the data. If your targets rely on that token to terminate, the model never sees a loss on it and learns to run on. phi-4 has a separate pad id, so the same collator appears to work there, which is why this bug often surfaces only after switching models. Mask padding by position, as below, and the problem cannot occur on either model:
import torch
IGNORE = -100
def build_example(tok, messages, max_len=4096):
"""messages: [{"role": "system"|"user"|"assistant", "content": str}, ...],
ending with the assistant turn we want to learn. Loss is on that turn only."""
prompt_ids = tok.apply_chat_template(
messages[:-1], add_generation_prompt=True, tokenize=True, return_dict=False)
full_ids = tok.apply_chat_template(messages, tokenize=True, return_dict=False)
if full_ids[:len(prompt_ids)] != prompt_ids:
raise ValueError("template prefix mismatch; do not train on this example")
labels = [IGNORE] * len(prompt_ids) + full_ids[len(prompt_ids):]
if len(full_ids) > max_len:
return None # drop, never truncate away the turn end
return {"input_ids": full_ids, "labels": labels}
def collate(batch, pad_id):
n = max(len(b["input_ids"]) for b in batch)
ids, labels, attn = [], [], []
for b in batch:
k = n - len(b["input_ids"])
ids.append(b["input_ids"] + [pad_id] * k)
# mask padding by POSITION, never by comparing ids to pad_id:
# on Phi-4-mini the pad id is also the EOS id
labels.append(b["labels"] + [IGNORE] * k)
attn.append([1] * len(b["input_ids"]) + [0] * k)
return {"input_ids": torch.tensor(ids), "labels": torch.tensor(labels),
"attention_mask": torch.tensor(attn)}Drop over-length examples instead of truncating them, because truncation removes the turn end. The prefix check refuses any example the template renders inconsistently instead of mislabelling it.
A complete LoRA training loop
The loop below uses only stable calls from transformers and PEFT, so it does not depend on version-specific trainer flags. It targets the fused modules, enables gradient checkpointing, and saves both the adapter and a merged model. train_loader, EPOCHS, ACCUM and evaluate are yours to define; the loader uses the collator above.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, get_cosine_schedule_with_warmup
from peft import LoraConfig, get_peft_model
MODEL = "microsoft/Phi-4-mini-instruct" # or "microsoft/phi-4"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, torch_dtype=torch.bfloat16, device_map="auto")
model.gradient_checkpointing_enable()
model.enable_input_require_grads() # needed with checkpointing + frozen embeddings
cfg = LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
target_modules=["qkv_proj", "o_proj", "gate_up_proj", "down_proj"], # fused names in Phi-3/Phi-4
)
model = get_peft_model(model, cfg)
model.print_trainable_parameters() # sanity check: ~23M on mini, ~56M on phi-4
opt = torch.optim.AdamW([p for p in model.parameters() if p.requires_grad], lr=2e-4, weight_decay=0.0)
steps = len(train_loader) * EPOCHS // ACCUM
sched = get_cosine_schedule_with_warmup(opt, int(0.03 * steps), steps)
model.train()
for epoch in range(EPOCHS):
for i, batch in enumerate(train_loader):
batch = {k: v.to(model.device) for k, v in batch.items()}
loss = model(**batch).loss / ACCUM
loss.backward()
if (i + 1) % ACCUM == 0:
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
opt.step(); sched.step(); opt.zero_grad(set_to_none=True)
evaluate(model, tok, val_set) # task metric AND a general regression set
model.save_pretrained("out/adapter") # adapter weights only, not the base
merged = model.merge_and_unload() # only when serving without adapters
merged.save_pretrained("out/merged"); tok.save_pretrained("out/merged")Reasonable starting values for a task dataset of a few thousand examples are rank 16, alpha 32, learning rate 1e-4 to 2e-4, two to three epochs, and an effective batch of 16 to 32 sequences. The adapter recipe itself is explained in LoRA in depth. If you quantize the base for QLoRA, merge the trained adapter into a bf16 copy of the base, not into the 4-bit weights, and then quantize the merged model for deployment.
Worked example: support-ticket triage on Phi-4-mini
Suppose you need to turn free-text support tickets into a JSON object with a category, a priority and a one-line summary, on a CPU-only edge box. Prompting the base model gets the idea right but produces invalid JSON on a noticeable fraction of tickets and misuses your category names. This is a good fine-tuning target: the output format is fixed, the labels are cheap to produce, and a 3.8B model is small enough to serve after quantization.
Collect 3,000 historical tickets with their final human categorization and write the target JSON for each, under a fixed system message listing the allowed categories. Hold out 300 tickets split by time, not at random. Train LoRA at rank 16 for three epochs; at an effective batch of 16 2,700 training tickets make about 500 optimizer steps. Measure JSON parse rate, category accuracy and priority accuracy, plus a few dozen general prompts to catch forgetting.
Read outputs, not only metrics: replies that run past the closing brace point to the turn-end label problem, and stray turn markers point to a template mismatch. Once the numbers pass, merge, quantize and re-run the same evaluation, because quantization can undo a narrow fine-tune; see SLM evaluation pitfalls.
Failure modes
- The model never stops. The turn-end token was masked or truncated away. Check the labels of one batch by decoding the positions where labels are not -100.
- Loss falls, behaviour does not change. The adapter targeted too few layers, or none. Compare the printed trainable count with the arithmetic.
- Output contains stray markers. Training and inference used different chat formats, often because the code was ported between phi-4 and mini.
- Out-of-memory on a small model. Mini's logit tensor is large; reduce the maximum length or the micro-batch.
- Long-context quality drops on mini. Training only on short sequences can change behaviour at long range. If your use needs long inputs, include long examples, and keep the model's rope scaling configuration unchanged.
- General skills regress. The learning rate was too high or training ran too long. Lower the rate, add a small share of general instruction data, or stop earlier.
Trade-offs: phi-4 or Phi-4-mini
Choose phi-4 when the task needs reasoning depth and you can serve a 14B model. Choose Phi-4-mini when latency, memory or edge deployment dominate, or when you need more than 16K of context. Fine-tune mini first, because runs are cheap, and move to phi-4 only if evaluation plateaus below your bar. For preference tuning afterwards, the Phi-4 model card reports that its own post-training combined supervised fine-tuning with iterative DPO; see DPO alignment.
What to do next
- Pick the base from the table above and record its exact checkpoint name and revision.
- Render three training examples with
apply_chat_templateand read them character by character. - Build labels with the prefix method and decode the labelled span of one batch to confirm it ends with the turn-end token.
- Attach LoRA to
qkv_proj,o_proj,gate_up_projanddown_proj, and check the trainable count against the arithmetic. - Run 20 steps and record peak memory before launching the full run.
- Evaluate on a time-split held-out set and a general regression set, then re-evaluate after merging and quantizing.