Low-Rank Adaptation, introduced by Hu and colleagues in 2021, fine-tunes a large model by freezing every original weight and learning a small correction to selected matrices. Instead of updating a weight matrix W0 directly, LoRA learns two thin matrices, B and A, whose product has rank at most r, and adds a scaled B times A to W0. On GPT-3 175B the paper reported roughly 10,000 times fewer trainable parameters and about a third of the GPU memory of full fine-tuning, with comparable quality on the tasks it tested.

This site already covers the mathematics and parameter counting and a full training recipe for small models. This article is about LoRA as code and as an artifact. You will build the layer yourself, inject it into a model, see exactly what happens on the first training step, do the same with the PEFT library, and then learn to open a saved adapter and answer the questions that matter in production: did it train, does it match its base model, and does merging preserve its behaviour?

Advertisement

The layer, from scratch

A LoRA layer: the frozen path and the trainable low-rank path add up; only A and B are savedinput xd_inFrozen W0d_out x d_in, no gradAr x d_in, trainableBd_out x r, starts at 0x scalealpha / rsum = hW0 x + s B A xadapter_model.safetensorsonly lora_A / lora_B tensorsMerge: W = W0 + s B Asame output, no extra matmulsBecause B starts at zero, h equals W0 x at step 0: the adapted model is exactly the base model until training moves B
A LoRA-wrapped linear layer. The frozen weight and the low-rank path run in parallel; merging folds the product into the weight.

A LoRA layer wraps an existing linear layer. The forward pass computes the original output and adds the low-rank correction, scaled by alpha divided by r. A is initialised randomly and B to zero. The whole idea fits in about thirty lines of PyTorch.

import math
import torch
import torch.nn as nn

class LoRALinear(nn.Module):
    def __init__(self, base: nn.Linear, r=8, alpha=16, dropout=0.0):
        super().__init__()
        self.base = base
        self.base.weight.requires_grad_(False)
        if self.base.bias is not None:
            self.base.bias.requires_grad_(False)
        self.r, self.scale = r, alpha / r
        dev = base.weight.device
        self.lora_A = nn.Parameter(torch.empty(r, base.in_features, device=dev))
        self.lora_B = nn.Parameter(torch.zeros(base.out_features, r, device=dev))
        nn.init.kaiming_uniform_(self.lora_A, a=math.sqrt(5))
        self.dropout = nn.Dropout(dropout)
        self.merged = False

    def forward(self, x):
        out = self.base(x)
        if not self.merged:
            out = out + (self.dropout(x) @ self.lora_A.T @ self.lora_B.T) * self.scale
        return out

    @torch.no_grad()
    def merge(self):
        if not self.merged:
            self.base.weight += (self.lora_B @ self.lora_A) * self.scale
            self.merged = True

    @torch.no_grad()
    def unmerge(self):
        if self.merged:
            self.base.weight -= (self.lora_B @ self.lora_A) * self.scale
            self.merged = False

Notice the order of the multiplications: x @ A.T @ B.T never materialises the full d_out by d_in product during training, which is where the compute saving comes from. The merged flag matters more than it looks. Merging twice adds the update twice, and running the low-rank path on top of a merged weight also doubles it. Both bugs produce a model that is subtly wrong rather than broken, which is the hardest kind to catch.

Injecting it into a model

To adapt a model you replace selected linear layers by name and freeze everything else. Module names differ between architectures, so the matching rule is the most fragile part of any LoRA setup.

def inject_lora(model, targets=("q_proj", "v_proj"), r=8, alpha=16):
    for p_ in model.parameters():
        p_.requires_grad_(False)
    replaced = []
    for name, module in list(model.named_modules()):
        for child_name, child in list(module.named_children()):
            if isinstance(child, nn.Linear) and child_name in targets:
                setattr(module, child_name, LoRALinear(child, r=r, alpha=alpha))
                replaced.append(f"{name}.{child_name}")
    if not replaced:
        raise RuntimeError(f"no modules matched {targets}; check model.named_modules()")
    trainable = sum(p_.numel() for p_ in model.parameters() if p_.requires_grad)
    total = sum(p_.numel() for p_ in model.parameters())
    print(f"wrapped {len(replaced)} layers; trainable {trainable:,} of {total:,}")
    return replaced

The RuntimeError on zero matches is not decoration. If your target names do not exist in the model, the optimizer receives no parameters or only unrelated ones, training runs, the loss may even drift because of other trainable pieces, and the saved adapter is empty. Fail loudly instead.

Advertisement

Where the memory goes

The saving is not in the weights, which stay in memory at full size, but in everything the optimizer keeps per trainable parameter. With Adam in mixed precision a full fine-tune holds, for every weight, a gradient, two moment estimates and often a float32 master copy, roughly 12 to 16 bytes beyond the weight itself. LoRA pays that only for A and B. You can measure the split directly instead of estimating it.

def memory_split(model, bytes_per_trainable=16):
    frozen = sum(p_.numel() * p_.element_size() for p_ in model.parameters() if not p_.requires_grad)
    trainable = [p_ for p_ in model.parameters() if p_.requires_grad]
    n = sum(p_.numel() for p_ in trainable)
    print(f"frozen weights:      {frozen / 2**30:6.2f} GiB")
    print(f"trainable params:    {n:,}")
    print(f"optimizer + grads:   {n * bytes_per_trainable / 2**30:6.3f} GiB (estimate)")

For a typical 1.5-billion-parameter model with rank 16 on the four attention projections, the trainable count lands in the low millions, so optimizer state drops from tens of gigabytes to tens of megabytes. Activations do not shrink, because the forward pass still runs through the full network; long sequences and large batches remain the limit, which is why gradient checkpointing is still common with LoRA.

What happens on step zero

Because B is all zeros, B times A is zero and the wrapped model's output is identical to the base model's. That is deliberate: training starts from the pretrained behaviour, not from a randomly perturbed one. You can and should test it.

torch.manual_seed(0)
x = torch.randn(4, 512)
base = nn.Linear(512, 512)
ref = base(x).detach().clone()
lora = LoRALinear(base, r=8, alpha=16)
assert torch.equal(lora(x), ref)          # exactly equal at initialisation

loss = lora(x).pow(2).mean()
loss.backward()
print(lora.lora_A.grad.abs().max())       # 0: gradient of A flows through B, which is zero
print(lora.lora_B.grad.abs().max())       # > 0: B learns first

The gradient of the loss with respect to A is proportional to B transposed, so on the very first step only B moves. From the second step on, B is non-zero and both matrices train. This asymmetry is harmless with a normal optimizer, but it explains a common confusion: logging the norm of A after one step and concluding the adapter is frozen. It also explains why the scale matters. With the original alpha over r scaling, the size of the update shrinks as r grows, which is why the rank-stabilised variant divides by the square root of r instead; the reasoning is in LoRA theory.

The same thing with PEFT

In practice you will use Hugging Face PEFT, which implements the same layer with more options. Its LoraConfig defaults to r=8 and lora_alpha=8, and its default initialisation sets B to zero so the adapter is a no-op before training. target_modules="all-linear" wraps every linear layer except the output layer. use_rslora=True switches the scale to alpha divided by the square root of r, and use_dora=True enables DoRA, which learns a magnitude vector for each adapted weight matrix and uses LoRA for its direction.

from transformers import AutoModelForCausalLM
from peft import LoraConfig, get_peft_model

base = AutoModelForCausalLM.from_pretrained(BASE_ID, revision=BASE_REVISION, torch_dtype="bfloat16")
config = LoraConfig(
    r=16, lora_alpha=32, lora_dropout=0.05,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
    task_type="CAUSAL_LM",
)
model = get_peft_model(base, config)
model.print_trainable_parameters()

# ... train ...
model.save_pretrained("adapters/support-v3")   # adapter_config.json + adapter_model.safetensors

Pin the base model revision, not just its name. An adapter is a set of corrections to specific weights; loaded onto a different revision of the same model, it applies those corrections to weights it never saw. PEFT records the base model name in adapter_config.json but will not stop you loading the adapter onto something else.

Worked example: did the adapter actually train?

A team fine-tuned a 1.5-billion-parameter model for support-ticket triage. Evaluation showed no improvement over the base model at all, to three decimal places. Identical scores are a signal, not bad luck: a real adapter changes outputs at least slightly. They opened the artifact.

import json
from safetensors.torch import load_file

cfg = json.load(open("adapters/triage-v1/adapter_config.json"))
print(cfg["base_model_name_or_path"], cfg["r"], cfg["lora_alpha"], cfg["target_modules"])

tensors = load_file("adapters/triage-v1/adapter_model.safetensors")
print(len(tensors), "tensors")
scale = cfg["lora_alpha"] / cfg["r"]
for key in sorted(k for k in tensors if "lora_B" in k)[:6]:
    B = tensors[key].float()
    A = tensors[key.replace("lora_B", "lora_A")].float()
    delta = (B @ A) * scale
    print(f"{key:70s} |B|={B.norm():.4f}  |dW|={delta.norm():.4f}")

Every B norm printed 0.0000. B had never moved, so every update was zero. The cause was in the training script: the optimizer had been built from base.parameters() before get_peft_model wrapped the model, so it held references to frozen tensors and none of the new LoRA parameters. The loss curve had looked plausible because of noise and the learning-rate schedule. Building the optimizer after wrapping, from parameters that require gradients, fixed it, and the update norms then varied by layer as expected, larger in later layers.

Two habits came out of the incident. The training script now asserts that at least one lora_B tensor has a non-zero norm after the first hundred steps. And the release pipeline prints the per-layer update norms for every adapter, so an all-zero or wildly large layer is visible before evaluation runs.

Verifying a merge

Merging, model.merge_and_unload() in PEFT, folds the update into the base weights so inference needs no extra matrix multiplications. It is not an in-place operation: assign the returned model and use it. If you need to switch adapters later, merge_adapter() and unmerge_adapter() keep a path back. Whichever you use, check the merge numerically rather than trusting it.

ids = tokenizer(probe_prompts, return_tensors="pt", padding=True).input_ids
with torch.no_grad():
    unmerged = peft_model(ids).logits.float()
    merged_model = peft_model.merge_and_unload()
    merged = merged_model(ids).logits.float()
print("max |diff|:", (merged - unmerged).abs().max().item())

In float32 the difference should be tiny. In bfloat16 it will be larger, because adding a small update to a large weight in a format with only 8 bits of precision rounds some of the update away. Agree a threshold by measuring task metrics, not just logits. Merging into a quantized base is lossier still: the sum must be re-quantized, and the update can fall between quantization levels. The usual order is to merge into the full-precision base and quantize afterwards, as described in QLoRA theory.

Several adapters on one base

Because an adapter is small, one base can host many. PEFT can load several named adapters and switch with set_adapter, and with model.disable_adapter(): runs the base model alone, which is the cleanest A/B comparison you can make. add_weighted_adapter combines adapters into a new one; treat the result as a new model that needs its own evaluation, because combining corrections trained separately does not combine their behaviours in any guaranteed way. For serving many adapters at once without merging, see serving LoRA adapters.

Failure modes

SymptomCauseCheck or fix
Scores identical to baseB never trained; optimizer built before wrappingPer-layer B norms; build optimizer after injection
Wrapped zero or few layersTarget names do not match this architectureFail on zero matches; print named_modules()
Behaviour doubles or overshootsMerged twice, or low-rank path run on merged weightsTrack merge state; merge exactly once per export
Quality drops after mergebfloat16 or quantized merge roundingMerge in float32, quantize afterwards, compare metrics
Adapter fine in dev, odd in prodDifferent base revision or tokenizerPin revision; store base hash with the adapter
Different results with same filesScale changed: alpha, r or rsLoRA flag editedNever edit adapter_config.json after training
Huge checkpointBase weights saved or modules_to_save too broadInspect tensor keys; save adapter only

What to do next

  1. Implement the from-scratch layer above once and run the step-zero test, so the mechanics are yours rather than the library's.
  2. Make injection fail loudly when no target module matches, and print the trainable parameter count at the start of every run.
  3. Build the optimizer only after wrapping, from parameters that require gradients.
  4. Add a release check that loads each adapter, prints per-layer update norms, and rejects all-zero or outlier layers.
  5. Store the base model revision and a hash alongside every adapter, and refuse to load onto anything else.
  6. Compare merged and unmerged outputs on a fixed probe set before shipping a merged model, merging in float32 and quantizing last.
Key takeaway: LoRA is a frozen linear layer plus a scaled low-rank path whose B starts at zero, so the adapted model begins exactly as the base model and only B moves on the first step. Writing the layer yourself makes the failure modes obvious: targets that match nothing, optimizers that hold the wrong parameters, merges applied twice, and adapters loaded onto the wrong base. Treat each adapter as an artifact with a pinned base, inspect its update norms before evaluation, and verify every merge numerically.