Low-Rank Adaptation, LoRA, fine-tunes a language model by freezing every original weight and learning a small correction next to some of them. For a small language model, one with a few hundred million to a few billion parameters, that turns fine-tuning into a job for a single consumer GPU and turns the result into a file of tens of megabytes that can be swapped, versioned and shipped separately from the base.

The mathematics of LoRA is covered in LoRA: The Math, and the case for and against parameter-efficient methods in Full FT vs PEFT trade-offs. This article is the working procedure for a small model: how to build the data, what to adapt, how much memory it takes, how to train and evaluate, and how to get the result onto a device. The worked numbers use Llama-3.2-3B, but every step applies to any decoder-only model of similar size.

Advertisement

What LoRA changes, in one paragraph

A linear layer computes y = Wx with W of shape out by in. LoRA keeps W frozen and adds two trainable matrices: A with shape r by in and B with shape out by r, where the rank r is small, typically 8 to 64. The layer now computes y = Wx + (alpha / r) B A x. B is initialised to zero, so at step zero the model is exactly the base model and training starts from its behaviour rather than from noise. Only A and B receive gradients and optimizer state, which is where almost all the memory saving comes from. After training, B A can be added into W, leaving an ordinary model with no extra inference cost.

Hugging Face PEFT, the library used below, defaults to r=8 and lora_alpha=8, and offers use_rslora=True, which scales by alpha divided by the square root of r instead of by r, so that raising the rank does not quietly shrink the update.

Why small models are a special case

Three things differ from fine-tuning a 70B model. First, the hidden size is small (3,072 for Llama-3.2-3B), so a given rank is a larger fraction of the layer's dimension, and adapters are proportionally more expressive. Second, a small model has less spare capacity: it is easier to overwrite what it already knows, so catastrophic forgetting is a practical concern rather than a footnote. Third, full fine-tuning is sometimes affordable at this size, which makes LoRA a choice rather than a necessity.

LoRA still wins most small-model projects for operational reasons. The adapter is small enough to version like code. Several task adapters can share one base in memory, which multi-LoRA serving exploits. And turning the adapter off gives you the exact base model back for comparison, which is the cheapest regression test you will ever have.

Advertisement

The pipeline end to end

LoRA fine-tuning pipeline for a small language modelExamplesprompt, response pairsChat template+ mask prompt (-100)Frozen basebf16, no gradientsAdapters A, Bfp32, trainableTraining looploss on response tokens, AdamW on adapters onlyTask evalheld-out, same templateRegression evalbase skills, adapter off vs onShipchoose one path belowMerge in bf16then quantize for edgeKeep separatemulti-adapter servingW' = W + (alpha / r) B AB starts at zero, so step 0 equals the base model
Data is rendered with the model's chat template and the prompt tokens are masked out of the loss. The frozen base and trainable adapters feed a training loop that is evaluated on the task and on regressions, then either merged and quantized for devices or kept separate for multi-adapter serving.

Each box is a place where projects go wrong. The template and masking determine what the model learns. The frozen base must be the exact checkpoint you will ship on. The evaluation must compare against the base with the adapter off. And the order of merge and quantization decides whether the device runs the model you evaluated.

Worked budget: parameters and memory for a 3B model

Start from the config: hidden size 3,072, MLP width 8,192, 28 layers, 24 query heads and 8 key-value heads of dimension 128, so the key and value projections are 3,072 by 1,024. Each LoRA pair on a layer with in and out features adds r times (in + out) parameters. The whole count is a few lines:

# Llama-3.2-3B shapes from its published config
d, ffn, layers, kv = 3072, 8192, 28, 8 * 128        # kv width = 8 KV heads x 128
shapes = {                                          # (in_features, out_features)
    "q_proj": (d, d), "k_proj": (d, kv), "v_proj": (d, kv), "o_proj": (d, d),
    "gate_proj": (d, ffn), "up_proj": (d, ffn), "down_proj": (ffn, d),
}

def lora_params(r, targets):
    # A is r x in, B is out x r
    return layers * sum(r * (i + o) for name, (i, o) in shapes.items() if name in targets)

print(lora_params(16, shapes))                      # 24,313,856 (all linear layers)
print(lora_params(16, ["q_proj", "v_proj"]))        # 4,587,520  (attention q, v only)

The base has about 3.21 billion parameters: 28 layers of 100.7 million each plus a 394-million-parameter embedding matrix shared with the output head. At r=16 over every linear layer the adapter is 24.3 million parameters, about 0.76 percent of the model. Adapting only the query and value projections, as the original LoRA paper did, gives 4.6 million.

Memory for training with the all-linear r=16 adapter in bf16:

ItemSizeNote
Frozen base weightsabout 6.4 GB3.21B parameters x 2 bytes
Adapter weightsabout 0.1 GB24.3M x 4 bytes (fp32)
Adapter gradientsabout 0.1 GBsame shape as the weights
AdamW stateabout 0.2 GBtwo fp32 moments
Activations, checkpointedestimate: a few GBgrows with micro-batch x sequence length
Logitsestimate: 4 GB or more16,384 tokens x 128,256 vocab x 2 bytes, more if upcast

The surprise for most people is the last row. Small models often keep a large vocabulary, and the output logits for a micro-batch of 8 sequences of 2,048 tokens are about 4.2 GB in bf16, twice that if the loss upcasts to fp32. On a 24 GB card, keep the micro-batch small and reach the effective batch size you want with gradient accumulation. Gradient checkpointing trades roughly one extra forward pass for most of the activation memory and is almost always worth it here.

Data: the template and the mask

The model learns whatever the loss rewards. If you compute the loss on every token, you also train it to predict the user's prompt, which wastes capacity and can teach it to ramble in the user's voice. The standard fix is to set the label of every prompt token to -100, the value PyTorch's cross-entropy ignores, so only response tokens contribute.

IGNORE = -100   # PyTorch cross-entropy skips this label

def build_example(tok, messages, max_len=2048):
    # messages: [{"role": "user", ...}, {"role": "assistant", ...}]
    prompt_ids = tok.apply_chat_template(messages[:-1], add_generation_prompt=True)
    full_ids = tok.apply_chat_template(messages)
    assert full_ids[:len(prompt_ids)] == prompt_ids, "template is not prefix-stable"
    labels = [IGNORE] * len(prompt_ids) + full_ids[len(prompt_ids):]
    return {"input_ids": full_ids[:max_len], "labels": labels[:max_len]}

Render every example with the tokenizer's own chat template, the one the model will see at inference. Training on a hand-rolled format and serving with the official template is one of the most common reasons a fine-tune looks good in the notebook and poor in the app. The assertion guards a subtle case: some templates render the prompt differently when an assistant turn follows, and the mask would then be off by several tokens.

Quality beats quantity at this size. A few thousand clean, consistent examples usually beat a hundred thousand scraped ones. Deduplicate, hold out a test set before you start, and make sure the end-of-turn token is in the labels, or the model will not learn to stop.

The training loop

Trainer libraries wrap this loop and change their flags often, so here is the loop itself in plain PyTorch on top of PEFT. The values are starting points to adjust, not recommendations for every dataset.

import torch
from transformers import AutoModelForCausalLM
from peft import LoraConfig, get_peft_model

base = AutoModelForCausalLM.from_pretrained(BASE_ID, torch_dtype=torch.bfloat16).cuda()
base.gradient_checkpointing_enable()
base.enable_input_require_grads()    # checkpointing with a frozen embedding layer

cfg = LoraConfig(
    r=16, lora_alpha=32, lora_dropout=0.05, bias="none",
    target_modules="all-linear", task_type="CAUSAL_LM",
)
model = get_peft_model(base, cfg)
model.print_trainable_parameters()   # expect roughly 0.76% for a 3B Llama

opt = torch.optim.AdamW([p for p in model.parameters() if p.requires_grad],
                        lr=2e-4, weight_decay=0.0)
sched = cosine_with_warmup(opt, warmup=50, total=total_steps)

model.train()
for step, batch in enumerate(loader):              # micro-batches, padded
    out = model(input_ids=batch["input_ids"].cuda(),
                attention_mask=batch["attention_mask"].cuda(),
                labels=batch["labels"].cuda())
    (out.loss / accum).backward()
    if (step + 1) % accum == 0:
        torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
        opt.step(); sched.step(); opt.zero_grad(set_to_none=True)
    if step % eval_every == 0:
        evaluate(model)                            # task set + regression set

model.save_pretrained("adapter/")                  # only the adapter weights

Note what is not there: no optimizer state for the base, no gradient for the base, and a save that writes only the adapter. target_modules="all-linear" adapts every linear layer except the output head, which is usually what you want for small models; the output head and embeddings stay frozen unless you add tokens, in which case list them in modules_to_save so they are trained and saved in full.

Choosing rank, alpha, targets and learning rate

KnobStarting pointMove it when
Targetsall linear layersMemory is tight: fall back to attention projections only
Rank r16Task is narrow (format, tone): try 8. Task adds knowledge or skill: try 32 or 64
Alpha2 x r, or use rsLoRAKeep alpha / r fixed when you change r, or the effective step size changes
Learning rate1e-4 to 2e-4Loss spikes or diverges: lower it. Plateau early: raise it modestly
Epochs1 to 3Held-out loss rises while training loss falls: stop earlier
Dropout0 to 0.1Small dataset and visible overfitting: raise it

Run the sweep as a controlled experiment: change one knob, keep the data order and seed fixed, and compare held-out task metrics, not training loss. For small models, rank matters less than data quality and target coverage; doubling the rank rarely rescues a weak dataset.

Evaluate twice: the task and what you might have broken

Measure the task on a held-out set rendered with the same template, using the metric the product cares about: exact match for extraction, pass rate for code, a rubric for tone. Then measure regressions: a fixed suite of general prompts, safety prompts and any capability the product depends on. The PEFT model can switch its adapter off, so both runs use identical weights and identical code, and any difference is the adapter's doing.

from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained(BASE_ID, torch_dtype=torch.bfloat16)
model = PeftModel.from_pretrained(base, "adapter/")

with model.disable_adapter():                      # same weights, adapter off
    base_score = run_regression_suite(model)
tuned_score = run_regression_suite(model)

merged = model.merge_and_unload()                  # returns a plain model; assign it
merged.save_pretrained("merged-bf16/")
# then convert merged-bf16/ with your runtime's converter and quantize it

If the regression suite drops, lower the learning rate, reduce epochs, lower the rank, or mix a small share of general instruction data into the training set. A fine-tune that wins the task and loses the ability to refuse, or to follow a simple format, is not a win.

Shipping: merge first, then quantize

On a device you usually want one quantized file with no adapter logic. The order matters. Merge the adapter into the base in bf16, save the merged model, then quantize it with the runtime's own converter. Quantizing the base and then adding a full-precision adapter at run time is possible in some runtimes, but it is a different model from the one you evaluated, so evaluate it separately if you do it. See SLM edge quantization for choosing a format and calibrating on-device.

If you trained with QLoRA, the adapter learned to correct a 4-bit base. Merging it into a bf16 base gives a slightly different model from the one that was trained; usually close, but run the evaluation again after merging rather than assuming.

Keep adapters separate instead when one server hosts many tasks or tenants on one base. Then the adapter file is the deployable unit, and the base is loaded once.

Failure modes

  • Output never stops. The end-of-turn token was masked or missing from labels. Check the last labels of a rendered example.
  • Great notebook results, poor app results. Train and serve templates differ, or the app uses a different system prompt. Render both and diff the token ids.
  • Adapter has no effect. Target module names did not match, so nothing was adapted. Check the trainable-parameter printout before the first step.
  • Quality drops after quantization. The adapter was merged into an already-quantized base, or the quantization format is too aggressive for the merged weights. Merge in bf16, then compare formats.
  • Changing the rank changed everything. Alpha was left fixed, so the effective scale alpha/r moved. Keep the ratio or use rsLoRA.

What to do next

  1. Write the task metric and the regression suite before collecting training data, and record base-model scores for both.
  2. Build 1,000 to 5,000 clean examples, render them with the official chat template, and assert that the prompt mask is correct on a sample.
  3. Train an r=16 all-linear adapter with the loop above, with a small micro-batch and gradient accumulation, and watch held-out loss.
  4. Compare adapter on versus off on both suites; adjust learning rate, epochs or data mix if regressions appear.
  5. Merge in bf16, quantize with your runtime's converter, and rerun both suites on the quantized file on target hardware.
  6. Version the adapter, the exact base checkpoint id and the template together, because the adapter is meaningless without the other two.
Key takeaway: For small language models, LoRA turns fine-tuning into a single-GPU job and the result into a small versioned file. The hard parts are not the matrices but the procedure: render data with the real chat template, mask the prompt, adapt all linear layers at a modest rank, keep micro-batches small because the logits are large, evaluate against the adapter-off base for regressions, and merge in bf16 before quantizing so the device runs the model you actually measured.