The training recipe for supervised fine-tuning fits on a slide. Running it repeatedly, on shared GPUs, with results you can trust and reproduce, is a different job. Most failed fine-tunes do not fail in the optimizer; they fail because the run used a different base revision than the one evaluated, ran out of memory on step 1,900, masked the wrong tokens, lost six hours to a preemption, or shipped an adapter that quietly made the model worse at everything it was not trained on.

This article is the runbook for that job. It walks the lifecycle of a run, works a memory plan for an 8B model with LoRA and QLoRA, shows a training loop instrumented for operations, and catalogues the failures you will meet. Cost estimation and the memory budget of full fine-tuning have their own articles; this one is about running jobs well.

Advertisement

The lifecycle of a fine-tuning run

Datasetversioned, hashedRun specconfig + base revMemory planfits? batch, seqLaunchtorchrun / schedulerTrain loopstep, log, guardCheckpointatomic, resumableEval gatetask + regressionRegistryadapter + lineageServemerge or multi-LoRAMonitoringloss, grad norm, tokens/s, MFU, memorymetricsEvery arrow is a place where a run can silently go wrong: wrong data, wrong base revision, OOM, NaN, lost progress, regression.Fine-tuning ops is the discipline of making each one checked, logged and recoverable.
A fine-tuning run as a pipeline. Monitoring spans the training loop and the checkpoint stage; the eval gate decides whether anything reaches the registry.

Each stage has one job. The run spec pins every input. The memory plan proves the configuration fits before you queue for GPUs. The training loop turns steps into checkpoints while emitting the signals that tell you whether to stop early. The eval gate compares the result against the base model on the target task and on capabilities you must not lose. The registry stores the adapter with enough lineage to rebuild it. Skipping any stage does not save time; it moves the failure later, where it costs more.

Pin everything in a run spec

A run is reproducible when one file determines it. Store the spec next to the checkpoints and the metrics, and refuse to start a run whose inputs are not pinned by content hash or immutable revision.

run_id: support-tone-r16-2026-10-01a
base_model: {repo: meta-llama/Llama-3.1-8B-Instruct, revision: <commit sha>}
tokenizer: {same_as_base: true, chat_template_sha256: <hash>}
data:
  train: {uri: s3://ft-data/support/v7/train.jsonl, sha256: <hash>}
  eval:  {uri: s3://ft-data/support/v7/eval.jsonl,  sha256: <hash>}
  max_seq_len: 4096
  packing: true
  loss_on: assistant_only
method: {type: lora, r: 16, alpha: 32, dropout: 0.05,
         targets: [q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj]}
quantization: none            # or nf4 for QLoRA
optim: {name: adamw, lr: 2.0e-4, warmup_steps: 50, schedule: cosine,
        weight_decay: 0.0, grad_clip: 1.0}
batch: {micro: 1, grad_accum: 16, world_size: 1}
precision: bf16
seed: 1234
checkpoint: {every_steps: 200, keep_last: 3}
eval_gate: {task_metric_min_gain: 0.03, regression_max_drop: 0.01}

The chat template hash is there because a changed template changes every training token while leaving the data file identical. The effective batch is micro-batch times accumulation times world size, here 16 sequences of up to 4,096 tokens; record it explicitly because learning rates only transfer between runs with the same effective batch.

Advertisement

Worked memory plan: an 8B model on one GPU

Take the published Llama-3.1-8B shape: 32 layers, hidden size 4,096, 8 key-value heads of dimension 128 (so the k and v projections output 1,024), MLP width 14,336, vocabulary 128,256, about 8.0B parameters. A LoRA adapter of rank r on a weight of shape in by out adds r times (in + out) parameters.

Per layer at r = 16: q and o projections 16 x 8,192 = 131,072 each; k and v 16 x 5,120 = 81,920 each; gate, up and down 16 x 18,432 = 294,912 each. That is 1,310,720 per layer, and about 41.9M trainable parameters across 32 layers, roughly 0.5 percent of the model. Kept in fp32 with AdamW, each trainable parameter costs about 16 bytes (weight, gradient, two moments): about 0.67 GB. The adapter is not the memory problem.

The frozen base is. In bf16 it is about 16 GB. Quantised to 4-bit NF4 as in QLoRA, the quantised layers shrink to roughly a quarter; with quantisation constants and the embeddings and output head usually left in higher precision, plan on roughly 5 to 6 GB. Full fine-tuning, by contrast, needs about 16 bytes for every one of the 8B parameters before activations, which is why it needs sharding across several GPUs; full fine-tuning in depth works that budget.

Activations depend on sequence length. With gradient checkpointing, the saved layer inputs cost 32 layers x 4,096 tokens x 4,096 x 2 bytes, about 1 GiB per sequence, plus the activations of the one layer being recomputed. Then comes the trap: the output logits are 4,096 x 128,256 values, about 1.05 GB in bf16, and many loss implementations upcast them to fp32 (2.1 GB) and materialise a gradient of the same size. The loss step alone can transiently need 4 GB or more, larger than all the checkpointed activations together.

Item (micro-batch 1, 4,096 tokens)LoRA, bf16 baseQLoRA, NF4 base
Base weightsabout 16 GBabout 5-6 GB
Adapter + AdamW stateabout 0.7 GBabout 0.7 GB
Checkpointed activationsabout 1 GBabout 1 GB
Recomputed layer + logits/loss peakabout 4-6 GBabout 4-6 GB
CUDA context, allocator slack1-2 GB1-2 GB
Planning totalabout 23-26 GBabout 12-15 GB

These are planning estimates, not measurements; kernels, allocator behaviour and library versions move them. The procedure is what matters: estimate, then run ten steps at the longest sequence length in the dataset and read torch.cuda.max_memory_allocated(). Testing with short sequences is how runs die at step 1,900 when the first long example arrives. If the logits dominate, a chunked or fused cross-entropy that never materialises the full fp32 tensor recovers several GB.

Data path: templates, masking and packing

Three data decisions change what the model learns more than any hyperparameter. Use the base model's chat template, applied by its tokenizer, never a hand-written format. Mask the loss so only assistant tokens contribute, otherwise the model also learns to imitate user turns and the system prompt. End every example with the end-of-turn token, or the fine-tuned model will not know when to stop.

def build_example(tok, messages, max_len):
    ids, labels = [], []
    for i, m in enumerate(messages):
        # render the prefix up to and including this turn, take only the new tokens
        upto = tok.apply_chat_template(messages[: i + 1], tokenize=True)
        assert upto[:len(ids)] == ids, "chat template is not prefix-stable"
        new = upto[len(ids):]
        ids += new
        labels += new if m["role"] == "assistant" else [-100] * len(new)
    if len(ids) > max_len:
        return None                      # drop, do not truncate mid-answer
    assert any(l != -100 for l in labels), "example has no trainable tokens"
    return {"input_ids": ids, "labels": labels}

Packing concatenates short examples into full-length sequences so no compute is wasted on padding. It is a large throughput win on chat data, where most examples are short, but without document-boundary-aware attention, tokens can attend across example boundaries. Use your framework's packing with per-example position resets and boundary masks where available, and log the padding fraction either way: it is the cheapest efficiency signal you have.

A training loop instrumented for operations

The loop below uses PEFT and plain PyTorch. The operational additions are the point: a loss and gradient-norm guard, throughput and memory logging, periodic atomic checkpoints, and a preemption handler that checkpoints and exits cleanly on SIGTERM.

import math, signal, threading, time, torch
from peft import LoraConfig, get_peft_model

model = load_base(spec)                         # bf16, or 4-bit via BitsAndBytesConfig
model.gradient_checkpointing_enable()
model = get_peft_model(model, LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
                                         target_modules=spec.targets, task_type="CAUSAL_LM"))
opt = torch.optim.AdamW([p for p in model.parameters() if p.requires_grad], lr=spec.lr)
sched = make_scheduler(opt, spec)               # warmup + cosine
step = resume_if_present(model, opt, sched, loader)   # restores RNG and data position

stop = threading.Event()
signal.signal(signal.SIGTERM, lambda *_: stop.set())
t0, tokens, skipped = time.time(), 0, 0
for micro, batch in enumerate(loader, start=1):
    loss = model(**batch).loss / spec.grad_accum
    loss.backward()
    tokens += int(batch["attention_mask"].sum())
    if micro % spec.grad_accum:
        continue
    gnorm = torch.nn.utils.clip_grad_norm_(model.parameters(), spec.grad_clip)
    if not math.isfinite(gnorm) or not math.isfinite(loss.item()):
        opt.zero_grad(set_to_none=True)          # skip the step, count it, alert
        skipped += 1
        continue
    opt.step(); sched.step(); opt.zero_grad(set_to_none=True); step += 1
    if step % 10 == 0:
        dt = time.time() - t0
        log(step=step, loss=loss.item() * spec.grad_accum, gnorm=float(gnorm),
            lr=sched.get_last_lr()[0], tok_s=tokens / dt, skipped=skipped,
            peak_gb=torch.cuda.max_memory_allocated() / 2**30)
        t0, tokens = time.time(), 0
    if step % spec.ckpt_every == 0 or stop.is_set():
        save_checkpoint_atomic(model, opt, sched, loader, step)
        if stop.is_set():
            break

For QLoRA, load the base with BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_use_double_quant=True, bnb_4bit_compute_dtype=torch.bfloat16) and call PEFT's prepare_model_for_kbit_training before adding the adapter. Skipped steps should be rare; if more than a handful occur, stop and investigate rather than letting the guard hide a real instability.

Checkpoints you can actually resume from

A checkpoint that only holds adapter weights lets you evaluate, not resume. To resume a run exactly, save the adapter weights, optimizer state, scheduler state, the step counter, every RNG state (Python, NumPy, torch CPU and CUDA), and the data loader's position, plus the run spec. Write to a temporary directory and rename it into place only after every file is flushed, so a crash during a save never leaves a half-written checkpoint as the newest one. Keep the last few and the best by eval loss.

Test resume deliberately: run 100 steps, kill the job at step 60, resume, and compare the loss curve with an uninterrupted run. If the curves diverge beyond numerical noise, something is not being restored, most often the data position, which makes the resumed run see some examples twice and others never.

Monitoring: the signals that matter

SignalHealthyInvestigate when
Training lossFalls quickly, then flattensSpikes, goes flat at the start, or approaches zero
Gradient normStable after warmupGrows steadily or jumps by an order of magnitude
Eval lossTracks training lossRises while training loss falls (overfitting)
Tokens per secondSteadyDrops step to step: data stalls, throttling, a slow node
Peak memoryFlat after the first stepsCreeps upward: a leak, or growing sequence lengths
Skipped stepsZero or near zeroAny sustained count

Convert tokens per second into model FLOPs utilisation to know whether the GPU is being used well. LoRA training with gradient checkpointing costs roughly 6N FLOPs per token for an N-parameter model: 2N forward, about 2N to propagate gradients back through the frozen weights, and 2N for recomputation. For 8B parameters that is about 4.8e10 FLOPs per token, so a hypothetical 3,000 tokens per second is about 144 TFLOP/s, around 15 percent of an H100 SXM's roughly 989 TFLOP/s dense BF16 peak. Low utilisation at micro-batch 1 is normal; when throughput matters, raise the micro-batch before adding GPUs. Anatomy of a training step explains where that time goes.

Eval gates and shipping

An adapter ships only if it passes two tests against the pinned base model: a task gain on held-out examples from the target distribution, and a regression check on a fixed suite covering what the base model already does well, such as instruction following, refusals and general knowledge. Training loss is not an eval; a model that memorises the training set has excellent loss.

Register the passing adapter with its run spec, data hashes, base revision, eval results and checkpoint step. At serving time you either merge the adapter into the base weights, which costs nothing per token but creates a full model copy per adapter, or serve it unmerged so one base can carry many adapters, described in multi-LoRA serving. Never merge into a quantised base and ship the result without re-evaluating it.

Failure catalogue

  • OOM late in the run. The longest example arrived. Plan with the maximum length, cap or drop over-length examples, and fix the logits peak.
  • Loss near zero. Usually a masking bug training on prompt tokens, or train and eval overlap. Inspect labels on real batches.
  • Endless generations. Missing end-of-turn tokens in training data, or a template mismatch between training and serving.
  • NaN or loss spikes. Learning rate too high for the adapter rank, fp16 overflow (prefer bf16, see mixed precision training), or a corrupt example. Log the batch indices of skipped steps.
  • Lost progress. Preemption without a SIGTERM handler, or non-atomic saves. Test the kill-and-resume path before you need it.
  • Quiet regressions. The task metric rises while general ability falls. That is what the regression suite is for.

On trade-offs: LoRA in bf16 is the default; QLoRA trades some throughput (dequantisation on every forward) for fitting larger models on smaller GPUs; full fine-tuning costs several times more memory and is worth it when the distribution shift is large. Estimating fine-tuning costs turns these choices into spend.

What to do next

  1. Write a run spec template that pins base revision, tokenizer template hash and data hashes, and refuse unpinned runs.
  2. Measure peak memory for ten steps at the dataset's maximum sequence length before queueing a long run.
  3. Inspect the labels of five real batches to confirm only assistant tokens are trained and every example ends with the end-of-turn token.
  4. Add the gradient-norm guard, throughput and memory logging, atomic checkpoints and a SIGTERM handler to your loop.
  5. Run the kill-at-step-60 resume test once per framework upgrade.
  6. Build a regression suite and make the eval gate a required step before any adapter is registered.
Key takeaway: Fine-tuning ops turns a short training recipe into runs you can trust: pin every input in a run spec, plan memory at the longest sequence and watch the logits peak, mask and template data correctly, guard and instrument the loop, write checkpoints that genuinely resume, and let an eval gate with a regression suite decide what ships.