Full fine-tuning updates every weight of a pretrained model on new data. It is the most expressive way to adapt a model and the most expensive: the optimizer needs state for every parameter, the checkpoints are full model copies, and a mistake can erase capabilities that took trillions of tokens to learn. Parameter-efficient methods such as LoRA exist because of that cost, yet full fine-tuning remains the quality reference they are measured against, and for some jobs it is still the right tool.
This article covers full fine-tuning as an engineering project on GPUs: how to budget memory across a cluster and choose a sharding stage, how to prepare data so that the model learns the right thing, a training recipe with pseudocode, how to evaluate, what goes wrong, and when full fine-tuning actually beats LoRA. The per-bucket memory derivation and a 7B worked example are in Fine-Tuning Math; supervised fine-tuning as a method is in SFT.
The memory budget: 16 bytes per parameter, then activations
With mixed-precision AdamW, the standard setup, each trainable parameter carries BF16 weights (2 bytes), BF16 gradients (2), an FP32 master copy of the weights (4) and two FP32 Adam moments (8): 16 bytes of model state. Some setups keep gradients in FP32 or add buffers, and activations, temporary workspace and allocator fragmentation come on top, which is why rules of thumb of 18 to 20 bytes per parameter circulate. Plan with 16 for states and measure the rest. A 13B model therefore needs about 208 GB of model state before a single activation is stored, and a 70B model about 1.1 TB; neither fits on one 80 GB GPU.
Sharding divides those buckets across data-parallel ranks. ZeRO stage 1 shards the optimizer state, stage 2 also shards gradients, and stage 3, equivalent to PyTorch FSDP's full sharding, also shards the weights and gathers each layer's weights just before use. The function below computes per-GPU state memory. For 13B on eight 80 GB GPUs, stage 1 needs about 71.5 GB per GPU and leaves no room for activations; stage 2 needs about 49 GB and leaves roughly 30 GB for activations with checkpointing; stage 3 needs 26 GB and leaves the most headroom at the cost of more communication.
# Per-GPU memory for model states under mixed-precision AdamW (16 bytes/param):
# BF16 weights (2) + BF16 grads (2) + FP32 master weights (4) + FP32 m (4) + FP32 v (4).
def per_gpu_gb(params, n_gpus, stage):
P = params
if stage == 0: # plain data parallel: everything replicated
b = 16 * P
elif stage == 1: # shard optimizer state (master + m + v)
b = 4 * P + 12 * P / n_gpus
elif stage == 2: # also shard gradients
b = 2 * P + 14 * P / n_gpus
elif stage == 3: # also shard weights (FSDP FULL_SHARD)
b = 16 * P / n_gpus
return b / 1e9 # activations, buffers and allocator slack come on top
for stage in (1, 2, 3):
print(stage, round(per_gpu_gb(13e9, 8, stage), 1))
# 13B on 8 GPUs -> stage 1: 71.5 GB, stage 2: 48.8 GB, stage 3: 26.0 GBFor 70B, 1.1 TB divided across 16 GPUs is 70 GB per GPU, which leaves no room for activations; 32 GPUs at full sharding bring it to 35 GB. The alternatives on fewer GPUs are CPU offload of optimizer state (ZeRO-Offload, limited by host memory and PCIe bandwidth) or 8-bit optimizer states, which cut the moments from 8 bytes to about 2 per parameter. See ZeRO, FSDP and activation checkpointing. Models of up to about 1.5B parameters fit on a single 80 GB GPU with room for activations, which makes them a cheap place to debug the pipeline.
The pipeline end to end
A full fine-tune is a data pipeline feeding a sharded training loop feeding an evaluation loop. The stages below each produce an artifact you can inspect: a cleaned dataset, tokenized and packed shards, sharded checkpoints and evaluation logs. Keeping them separate lets you rerun training without re-tokenizing, and rerun evaluation on any checkpoint.
Data preparation: most of the quality is decided here
Because every weight moves, full fine-tuning learns whatever the data contains, including its mistakes, formatting quirks and repeated boilerplate. Quality beats quantity for instruction tuning: the LIMA study (Zhou and colleagues, 2023) tuned a 65B model on 1,000 carefully chosen examples and obtained a strong assistant. Volume matters when the goal is new knowledge or a new domain, where billions of tokens may be needed.
Clean before tokenizing. Remove exact and near-duplicate examples, which otherwise get memorized. Decontaminate: drop training examples that overlap your evaluation sets, or your scores will be inflated. Filter by length and language, and check a random sample by hand. For chat data, render every example through the model's own chat template, the one its tokenizer defines, and verify that each assistant turn ends with the end-of-turn token; a missing end token teaches the model never to stop.
Mask the loss. In supervised fine-tuning, the model should learn to produce assistant turns, not to predict the user's prompt. Set label positions for system and user tokens to an ignore value so they contribute context but no gradient. Pack short examples into full-length sequences to avoid wasting compute on padding, but restart position ids at each document boundary and use an attention kernel that respects those boundaries, so examples do not attend to one another.
IGNORE = -100 # label value that PyTorch cross-entropy skips
def encode(example, tok):
msgs = example["messages"]
# Render the whole conversation once, exactly as it will be served.
ids = render(tok, msgs, add_generation_prompt=False)
labels = [IGNORE] * len(ids)
for i, turn in enumerate(msgs):
if turn["role"] != "assistant":
continue # system and user tokens are context only
start = len(render(tok, msgs[:i], add_generation_prompt=True))
end = len(render(tok, msgs[: i + 1], add_generation_prompt=False))
labels[start:end] = ids[start:end] # includes the end-of-turn token
return ids, labels
# render() wraps the tokenizer's chat template; assert that each prefix render
# is a true prefix of the full one, since some templates rewrite earlier turns.
def pack(examples, max_len):
buf_ids, buf_lab, buf_pos = [], [], []
for ids, labels in examples:
if len(ids) > max_len:
continue # or truncate on a turn boundary
if len(buf_ids) + len(ids) > max_len:
yield buf_ids, buf_lab, buf_pos
buf_ids, buf_lab, buf_pos = [], [], []
buf_ids += ids
buf_lab += labels
buf_pos += list(range(len(ids))) # positions restart per document
if buf_ids:
yield buf_ids, buf_lab, buf_pos
# Restarting positions lets varlen attention kernels keep documents from
# attending to each other; without that, packed examples leak context.To limit forgetting, mix in a small fraction of general data, such as instruction data similar to what the base model was tuned on or a slice of pretraining-style text. This replay costs a few percent more tokens and is often the cheapest protection against regressions.
The training recipe
Full fine-tuning uses much lower learning rates than pretraining or LoRA because every weight moves and the model starts from a good optimum. Common starting points for dense models in the 7B to 70B range are 1e-5 to 2e-5 peak learning rate with AdamW, a short linear warmup of about 3 percent of steps, cosine decay to about 10 percent of the peak, gradient clipping at 1.0, and one to three epochs. Larger models generally want the lower end. Treat these as a starting grid for a short sweep, not as constants; hyperparameter search covers how to run it cheaply.
Batch size is best counted in tokens: global batches of roughly half a million to a few million tokens are typical, reached with gradient accumulation when memory limits the micro-batch. Train in BF16 with FP32 master weights, which avoids the loss scaling FP16 needs (see mixed precision). Checkpoint activations for every transformer block when memory is tight; it adds roughly one extra forward pass of compute. Weight decay is often 0 to 0.1 for fine-tuning and matters less than the learning rate.
model = load_pretrained("base-13b", dtype=bf16)
model = shard(model, strategy="full_shard") # FSDP / ZeRO-3
enable_activation_checkpointing(model, every_block=True)
opt = AdamW(model.parameters(), lr=1e-5, betas=(0.9, 0.95), weight_decay=0.0)
sched = warmup_cosine(opt, warmup=0.03 * total_steps, total=total_steps, min_ratio=0.1)
for step, batch in enumerate(packed_loader): # ~0.5-1M tokens per global batch
for micro in split(batch, accum_steps):
loss = model(micro.ids, positions=micro.pos, labels=micro.labels).loss
(loss / accum_steps).backward()
clip_grad_norm_(model.parameters(), 1.0)
opt.step(); sched.step(); opt.zero_grad(set_to_none=True)
if step % eval_every == 0:
log(val_loss(model), task_eval(model), regression_suite(model))
save_sharded_checkpoint(model, opt, sched, step)
Evaluation and checkpoint selection
Track three things at every evaluation step. Held-out loss on data from the same distribution catches divergence and overfitting: training loss that keeps falling while validation loss rises by the second or third epoch is the classic signature. Task metrics, measured by generating answers for a fixed set of prompts, tell you whether the model does the job. A regression suite of general benchmarks and a few safety prompts tells you what the tuning cost elsewhere. The best checkpoint by task metric is often not the last one.
Save sharded checkpoints that include optimizer and scheduler state so training can resume, and consolidate only the chosen step into a single BF16 file for export. A 13B model's consolidated BF16 weights are about 26 GB; with optimizer state the sharded checkpoint is about 208 GB, so retain only a few.
When full fine-tuning beats LoRA
LoRA trains a low-rank update to frozen weights and matches full fine-tuning on many instruction-tuning tasks at a fraction of the memory; see LoRA fine-tuning. The gap opens when the target is far from what the base model knows and the data is large. Biderman and colleagues (2024), in LoRA Learns Less and Forgets Less, compared the two on code and math with both continued pretraining and instruction tuning at scales up to billions of tokens, and found full fine-tuning clearly more accurate and more sample-efficient, especially for continued pretraining, while LoRA preserved more of the base model's other abilities. They also measured the weight change learned by full fine-tuning and found it of much higher rank than typical LoRA configurations.
- Choose full fine-tuning for continued pretraining on a new domain or language, for large datasets (hundreds of millions of tokens or more), when you are producing a new base model for others to build on, and when the last points of accuracy on a single task justify the cost.
- Choose LoRA or QLoRA for small datasets, many tasks served from one base model, limited GPUs, and when preserving general capability matters more than peak task accuracy; see QLoRA.
- Check the middle ground. Higher LoRA ranks, adapters on all linear layers, or full fine-tuning of only the top blocks can close much of the gap for less memory.
Failure modes
- Catastrophic forgetting. General benchmarks drop sharply; lower the learning rate, shorten training, and mix in replay data.
- Divergence or loss spikes. Too high a learning rate or too short a warmup for this model size; restart from the last good checkpoint with a lower peak.
- Template mismatch. Training and serving format prompts differently and quality collapses at inference; render both through the same tokenizer template and test the served prompt string.
- Runaway generations. Examples lack the end-of-turn token, or it was masked out of the loss; check a tokenized example by eye.
- Cross-contaminated packing. Packed examples attend to each other and the model learns spurious continuations; restart positions and use document-aware attention.
- Overfitting in extra epochs. Small datasets memorize after a few passes; select the checkpoint by validation and task metrics.
- New tokens without embedding resize. Adding special tokens to the tokenizer without resizing and training the embedding and output matrices crashes or produces garbage.
Operational guidance and trade-offs
Debug on a small sibling model before renting the cluster: tokenize, pack and overfit a hundred examples on a 1B model, confirm the loss goes near zero and the generations stop cleanly, then scale. Measure throughput and memory in the first hundred steps of the real run and extrapolate the cost; fine-tuning costs and AdamW math help with the arithmetic. Keep the data, template, hyperparameters and seed in version control with each checkpoint so the result can be reproduced.
The trade-offs are stark. Full fine-tuning gives the highest ceiling and one self-contained model, but costs several times the memory of LoRA, produces full-size checkpoints per task and forgets more. Stage 3 sharding fits the largest models at the price of all-gather traffic on every forward and backward pass; stage 2 is faster when it fits. Offload fits on fewer GPUs but moves the bottleneck to PCIe. Choose the cheapest setup that meets the memory budget, and choose full fine-tuning itself only when the data and the target justify it.