Fine-tuning a small language model used to mean updating every weight. For a 1-billion-parameter model that is a 20 GB training job and a 2.5 GB file per task. Parameter-efficient fine-tuning (PEFT) freezes the pretrained weights and trains a small set of added or selected parameters instead, typically well under 1% of the model. The base stays shared, each task becomes a file of a few megabytes, and the job fits on a laptop GPU.
LoRA is the method most people reach for, and this site covers it in depth in LoRA for small language models and implementing LoRA in code. This article is about the choice before that one: what the whole family does, where each method plugs into the network, how many parameters each actually adds to a concrete model, what you save and what you do not, and how to decide.
The idea from first principles
A pretrained model already contains most of what a downstream task needs. Fine-tuning mostly steers it: towards a format, a domain vocabulary or a style of answer. Empirically, that steering lives in a small subspace. The intrinsic-dimension experiments summarised in PEFT theory found that many tasks can be learned by optimising a few hundred to a few thousand directions in weight space. PEFT methods are different guesses about which small set of parameters can express that steering.
There are three families. Additive methods insert new parameters: soft prompts, prefixes, adapter layers. Selective methods unfreeze a chosen subset of existing ones, such as BitFit's biases. Reparameterised methods train a compact description of a weight update, such as LoRA's two low-rank matrices, which can be folded back into the original weight after training. The family decides two practical things: whether inference gets slower, and whether one model can serve many tasks at once.
The methods, one paragraph each
LoRA learns an update dW = B A for a frozen weight W of shape d_out by d_in, with A of shape r by d_in and B of shape d_out by r. B starts at zero, so training begins from the unmodified model. The update is scaled by alpha / r (or alpha / sqrt(r) with rank-stabilised LoRA) and can be merged into W for zero inference overhead.
DoRA splits each weight into a magnitude vector and a direction, applies LoRA to the direction and trains the magnitude separately. It adds one value per output feature, costs more compute per step, is reported to close part of the gap to full fine-tuning at low ranks, and merges.
IA3 learns vectors that rescale activations elementwise: the outputs of the key and value projections and the input to the feed-forward down projection. It is far smaller than LoRA and merges, because scaling an input or output equals scaling a weight's columns or rows.
Prompt tuning prepends trainable embedding vectors (virtual tokens) to the input. Only the first layer sees them directly. Prefix tuning instead adds trainable key and value vectors to every attention layer, so each layer can be steered. Neither merges, the virtual tokens occupy context at every request, and the original prompt-tuning paper found the gap to full tuning closes only as models reach billions of parameters, so small models suffer most.
Bottleneck adapters insert a small down-project, nonlinearity, up-project block after attention and feed-forward sublayers. The nonlinearity makes them expressive but prevents merging, so they add latency. They live in the separate Adapters library rather than Hugging Face PEFT. BitFit trains only bias terms, and modern decoders such as Llama-style models have no biases in their linear layers, so on those it has nothing to train.
Exact parameter counts for a 1B-class model
Take a decoder with the shape of current 1B-class models: hidden size 2048, 16 layers, feed-forward width 8192, 32 query heads and 8 key-value heads of dimension 64, so the key and value projections map 2048 to 512. The full model is about 1.24 billion parameters including embeddings. LoRA adds r times (d_in + d_out) per adapted matrix. At r = 16 on all seven linear layers in a block:
| Matrix | Shape (out x in) | LoRA r=16 params | DoRA extra | IA3 params |
|---|---|---|---|---|
| q_proj, o_proj | 2048 x 2048 | 65,536 each | 2,048 each | - |
| k_proj, v_proj | 512 x 2048 | 40,960 each | 512 each | 512 each |
| gate_proj, up_proj | 8192 x 2048 | 163,840 each | 8,192 each | - |
| down_proj | 2048 x 8192 | 163,840 | 2,048 | 8,192 (input side) |
| Per block | 704,512 | 23,552 | 9,216 | |
| 16 blocks | 11,272,192 (0.91%) | 376,832 | 147,456 (0.012%) |
For the input-side methods: 20 virtual tokens of prompt tuning is 20 x 2048 = 40,960 parameters. Prefix tuning with 20 virtual tokens stores a key and a value per token per layer; with grouped-query attention each is 512 wide, so 20 x 16 x 2 x 512 = 327,680, before any reparameterisation network used only during training. BitFit on this architecture would train the RMSNorm scales at most, 67,584 values, since there are no biases.
Two things stand out. The spread is huge: IA3 is roughly 76 times smaller than r = 16 LoRA. And none of these numbers is the memory bill, which is the next point.
What PEFT saves, and what it does not
Full fine-tuning with AdamW in mixed precision costs about 16 bytes per parameter: 2 for bf16 weights, 2 for gradients, 4 for an fp32 master copy and 8 for the two Adam moments. For 1.24B parameters that is about 19.8 GB before a single activation is stored. With LoRA the frozen base needs only its 2 bytes per parameter, about 2.5 GB, and the 16-byte overhead applies to 11.3M parameters, about 180 MB. That is where the famous savings come from: gradients and optimizer state for frozen weights disappear.
What does not disappear: activations and most of the compute. To get a gradient for an adapter in layer 3, backpropagation must still pass through layers 16 down to 3, which means storing activations for every layer during the forward pass. A step costs the forward pass plus the backward pass through activations; only the weight-gradient half of the backward work is skipped for frozen matrices. As a rough rule, a PEFT step is about two thirds of a full fine-tuning step, not one percent of it. Sequence length and batch size, which drive activations, are still your main memory lever, and gradient checkpointing is still worth turning on.
Prompt and prefix tuning have one extra cost people miss: they lengthen every sequence. Twenty virtual tokens on a 256-token classification input is an 8% longer sequence in training and in serving, forever. The full trade-off space, including when full fine-tuning still wins, is laid out in full fine-tuning versus PEFT.
Choosing a method
| Situation | Start with | Why |
|---|---|---|
| General instruction or domain tuning of a 0.5-4B model | LoRA, all linear layers, r 8-32 | Best-tested quality per parameter; merges for free serving |
| LoRA plateaus below full fine-tuning at the rank you can afford | DoRA at the same rank | Small parameter increase, often recovers part of the gap; slower steps |
| Hundreds of per-customer variants, tiny storage budget | IA3 | Adapters of a few hundred KB; merges; fast to train |
| You cannot change the weights or the serving graph at all | Prompt tuning | Only an embedding tensor ships; works through APIs that accept input embeddings |
| Base must stay 4-bit to fit in memory | QLoRA | LoRA over a quantized frozen base; see the QLoRA page |
| Large distribution shift, new language, enough GPUs | Full fine-tuning | PEFT's low-rank prior can be the bottleneck |
Run LoRA first as the reference, because it has the most tooling and published hyperparameters. Then try the cheaper method and keep it only if it matches on your evaluation set. Going the other way, starting with prompt tuning on a 1B model and concluding "fine-tuning does not help", is a common and wrong conclusion. If memory forces a quantized base, QLoRA for small models covers when that is worth it.
Configs in Hugging Face PEFT
The peft library wraps a Transformers model so the chosen parameters are trainable and everything else is frozen. The three configs below were checked against the library's reference documentation. Module names are model-specific: these are Llama-style names, and you should print model.named_modules() for your model before trusting them.
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import LoraConfig, IA3Config, PromptTuningConfig, PromptTuningInit, get_peft_model
base_id = "path/or/hub-id-of-your-1b-model"
tok = AutoTokenizer.from_pretrained(base_id)
model = AutoModelForCausalLM.from_pretrained(base_id, torch_dtype="bfloat16")
lora = LoraConfig(task_type="CAUSAL_LM", r=16, lora_alpha=32, lora_dropout=0.05,
target_modules="all-linear", # every linear layer except the LM head
use_dora=False) # set True to try DoRA at the same rank
ia3 = IA3Config(task_type="CAUSAL_LM",
target_modules=["k_proj", "v_proj", "down_proj"],
feedforward_modules=["down_proj"]) # scaled on the input side
init_text = "Classify the support ticket as billing, outage, account or other:"
prompt = PromptTuningConfig(task_type="CAUSAL_LM",
prompt_tuning_init=PromptTuningInit.TEXT,
prompt_tuning_init_text=init_text,
num_virtual_tokens=len(tok(init_text)["input_ids"]),
tokenizer_name_or_path=base_id)
model = get_peft_model(model, lora) # or ia3, or prompt
model.print_trainable_parameters() # compare with the table aboveInitialising the soft prompt from real task text matters on small models, because random soft tokens can sit far from any real embedding. LoRA has its own initialisation options (init_lora_weights="pissa", "olora", "loftq" and others) that start the adapter from a decomposition of the weights instead of zero. Treat them as experiments once the default works.
A training loop that respects the frozen base
Any trainer works with a PEFT model, but a plain loop makes the moving parts visible. Two details are PEFT-specific: the optimizer should see only trainable parameters, and with gradient checkpointing the input embeddings must be marked as requiring gradients, or a fully frozen first layer cuts the graph and the adapters receive no gradient.
import torch
from torch.utils.data import DataLoader
model.gradient_checkpointing_enable()
model.enable_input_require_grads() # needed when the embedding layer is frozen
params = [p for p in model.parameters() if p.requires_grad]
# LoRA / IA3: lr around 1e-4 to 3e-4. Prompt tuning usually needs far higher, 1e-2 or more.
opt = torch.optim.AdamW(params, lr=2e-4, weight_decay=0.0)
loader = DataLoader(train_ds, batch_size=8, shuffle=True, collate_fn=collate)
model.train()
for step, batch in enumerate(loader):
# labels = input_ids with prompt tokens and padding set to -100, so loss covers answers only
out = model(input_ids=batch["input_ids"].cuda(),
attention_mask=batch["attention_mask"].cuda(),
labels=batch["labels"].cuda())
out.loss.backward()
torch.nn.utils.clip_grad_norm_(params, 1.0)
opt.step(); opt.zero_grad(set_to_none=True)
if step % 50 == 0:
print(step, round(out.loss.item(), 4))
model.save_pretrained("adapters/ticket-router-v1") # adapter weights plus adapter_config.jsonThe saved directory holds only the trainable tensors and a config that records the base model. To ship a merged model, load the base, attach the adapter with PeftModel.from_pretrained, call merge_and_unload() and save the result. Prompt and prefix tuning cannot be merged; they ship as adapters and are loaded at inference.
Worked example: routing support tickets on a 1B model
Suppose a team needs a four-way ticket classifier that runs on CPU at the edge, with 6,000 labelled tickets, and the base model scores 71% zero-shot on a held-out set of 600. They run three PEFT jobs on the same data on one 24 GB GPU. The accuracy figures below are illustrative, chosen to show the ordering these methods typically produce, not measured results.
LoRA, r = 16, all linear layers: 11.3M trainable parameters, 45 MB adapter in fp32, 93% accuracy. IA3: 147K parameters, under 1 MB, 89%. Prompt tuning with 14 text-initialised virtual tokens: 29K parameters, 82% after raising the learning rate to 3e-2; at 2e-4 it barely moved from zero-shot.
They ship IA3 merged into the base, because the four-point gap does not matter for routing and per-region variants are now tiny. The lesson generalises: let the strongest method set the ceiling, then buy back cost only as far as the evaluation allows.
Failure modes
- Wrong target modules. Names differ between architectures: some models use a fused
qkv_proj, GPT-2 usesc_attnConv1D layers that store weights transposed. A typo can match nothing or the wrong layers. Always readprint_trainable_parameters()and compare with your own count. - New tokens without trainable embeddings. Adding special tokens resizes the embedding matrix, but frozen embeddings leave the new rows random. Put
embed_tokensandlm_headinmodules_to_save, accepting that the adapter grows by those full matrices. - Learning rate copied across methods. Prompt tuning at a LoRA learning rate looks like "PEFT does not work".
- Adapter on the wrong base. An adapter is a delta against one exact checkpoint. Loading it on a different revision or a re-quantized base silently degrades output. Pin the base revision in the adapter's metadata and check it at load.
- Template drift. Training with one chat template and serving with another removes most of the gain. Freeze the template with the adapter version.
- Only measuring the target task. Small models forget more readily. Keep a short general-capability evaluation and run it on every adapter.
What to do next
- Write down your task's evaluation set and metric before choosing a method; 300 to 1,000 held-out examples is enough to rank methods.
- Run the base model zero-shot and few-shot on it, so you know what fine-tuning must beat.
- Train LoRA r = 16 on all linear layers as the reference run, and check the trainable count against the arithmetic above.
- Try IA3, or DoRA if LoRA falls short, on the same data and seed; keep the cheapest method within your accepted gap.
- Turn on gradient checkpointing and size batch and sequence length for activations, which PEFT does not shrink.
- Version each adapter with its base revision, template and evaluation scores; merge for single-task serving, keep separate for multi-tenant serving as in serving LoRA adapters.