LoRA freezes a pretrained weight and learns a low-rank correction B A. It is cheap and usually good enough, but at small ranks it often trails full fine-tuning. DoRA (Weight-Decomposed Low-Rank Adaptation, Liu and colleagues at NVIDIA, ICML 2024) closes part of that gap with one change: it splits every adapted weight into a magnitude vector and a direction matrix, trains the magnitude directly, and lets LoRA update only the direction.

This article explains that change from first principles, builds a DoRA layer from scratch in PyTorch, walks through how the Hugging Face PEFT library implements it and what each line costs on a GPU, sizes it for an 8B model, and ends with when to choose it over plain LoRA. It assumes you know how LoRA's training step works; if not, start with LoRA fine-tuning on GPUs.

From LoRA to DoRA: the decomposition

DoRA reparameterises each adapted weight as magnitude times unit directionW0 (frozen)out x in, pretrainedBout x r, init 0Ar x in, randomV = W0 + s B Adirection before scalings = alpha / rrow norms ||V||one per output, no gradientm (trainable)out values, init ||W0||W' = (m / ||V||) * Vrow-wise scale; equals W0 at step 0Trainable: A, B, mFrozen: W0After training, W' is computed once and written over W0: the served model is a plain linear layer.
The magnitude m and the low-rank pair B, A are the only trainable tensors; the norm is recomputed every step but carries no gradient.

Take a linear layer with weight W0 of shape out by in, as PyTorch stores it. Each row produces one output feature, so the layer can be written as a length per row times a unit vector per row. DoRA makes that split explicit and trainable:

V  = W0 + s * (B @ A)                 # LoRA-updated direction, s = lora_alpha / r
W' = m[:, None] * V / ||V||_row       # ||V||_row: L2 norm of each row, shape [out]

There are three trainable tensors: A (r by in), B (out by r) and the magnitude m (length out). B starts at zero and m starts at the row norms of W0, so at step zero V equals W0 and W' equals W0 exactly. Training starts from the pretrained model, as with LoRA.

The paper writes the norm as column-wise because it stores weights as in by out; PEFT computes the same thing as torch.linalg.norm(weight, dim=1) over the PyTorch layout, one magnitude per output feature. Read your framework's layout before porting the formula, or you will normalise the wrong axis and still get a model that trains, just worse.

Why separate magnitude and direction

The authors analysed how weights change during full fine-tuning and LoRA by measuring, per layer, how much the magnitude and the direction moved. In full fine-tuning the two changes were negatively correlated: some layers changed direction a lot with little magnitude change, others the reverse. LoRA showed a positive correlation, because one low-rank update has to move both together. DoRA's decomposition lets them move independently, and its pattern of updates looked more like full fine-tuning.

There is also a simpler intuition. A rank-r update is a poor tool for rescaling a whole 4,096-wide row; it would spend rank on something a single scalar does exactly. Giving each row its own scalar frees the low-rank update to change what the row points at.

The reported gains are modest and consistent: on commonsense reasoning, DoRA beat LoRA's average accuracy by 3.7 points on LLaMA-7B, about 1 on LLaMA-13B, 2.1 on LLaMA 2 7B and 4.4 on LLaMA 3 8B, and a half-rank DoRA still beat full-rank LoRA on all four. Treat these as evidence it is worth a trial on your data, not as a guarantee.

A DoRA layer from scratch

The quickest way to understand DoRA is to write it. This layer wraps a frozen nn.Linear, follows the paper's recommendation to treat the norm as a constant in the backward pass, and can produce a merged weight:

import math
import torch
import torch.nn as nn
import torch.nn.functional as F

class DoRALinear(nn.Module):
    def __init__(self, base: nn.Linear, r=16, alpha=32):
        super().__init__()
        self.base = base.requires_grad_(False)
        out_f, in_f = base.weight.shape
        self.s = alpha / r
        self.A = nn.Parameter(torch.empty(r, in_f))
        nn.init.kaiming_uniform_(self.A, a=math.sqrt(5))
        self.B = nn.Parameter(torch.zeros(out_f, r))
        self.m = nn.Parameter(base.weight.detach().float().norm(dim=1))

    def forward(self, x):
        V = self.base.weight + self.s * (self.B @ self.A)
        norm = V.float().norm(dim=1).detach()          # constant for backward (paper, sec. 4.3)
        scale = (self.m / norm).to(x.dtype)            # one factor per output feature
        y = F.linear(x, self.base.weight) + self.s * F.linear(F.linear(x, self.A), self.B)
        y = y * scale
        return y if self.base.bias is None else y + self.base.bias

    @torch.no_grad()
    def merged_weight(self):
        V = self.base.weight + self.s * (self.B @ self.A)
        return V * (self.m / V.float().norm(dim=1)).to(V.dtype)[:, None]

layer = DoRALinear(nn.Linear(64, 32))
x = torch.randn(8, 64)
assert torch.allclose(layer(x), F.linear(x, layer.merged_weight(), layer.base.bias), atol=1e-5)

Note that the forward pass never multiplies x by the full matrix V: it runs the frozen GEMM and the two thin LoRA GEMMs as usual, then scales each output column. V is built only to measure its row norms. Detaching the norm means autograd does not keep V for the backward pass and does not differentiate through the square root, which the authors report cuts training memory by about 24 percent when fine-tuning LLaMA, at a negligible accuracy cost. Gradients still reach A and B through the LoRA branch, and reach m through the scale.

DoRA in PEFT, line by line

In PEFT you enable DoRA with one flag on the usual LoRA configuration. The documentation states that DoRA currently supports linear and Conv2D layers, adds more overhead than plain LoRA, and should be merged for inference:

from peft import LoraConfig, get_peft_model

cfg = LoraConfig(
    r=16, lora_alpha=32, lora_dropout=0.05, task_type="CAUSAL_LM",
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
                    "gate_proj", "up_proj", "down_proj"],
    use_dora=True,
)
model = get_peft_model(base_model, cfg)
model.print_trainable_parameters()     # LoRA A/B plus the magnitude vectors

# ... train exactly as for LoRA ...

merged = model.merge_and_unload()      # bakes m / ||V|| into the weights
merged.save_pretrained("out/merged")

Internally, PEFT's DoRA layer computes the row norm of W0 + s * B @ A, detaches it, forms mag_norm_scale = magnitude / weight_norm and combines the outputs as (mag_norm_scale - 1) * base_result + mag_norm_scale * lora_result * scaling. The minus one is there because the base layer's output has already been added to the result; it is the same algebra as the scratch layer above. The magnitude is stored as a parameter initialised from the weight norm, so a freshly wrapped model reproduces the base model exactly; check that with a single forward pass before training. Training hyperparameters carry over from LoRA as a starting point; the original authors report DoRA is somewhat less sensitive to rank.

Worked example: Llama 3 8B at rank 16

Size it for a Llama 3 8B style model: 32 layers, hidden size 4,096, MLP size 14,336, and grouped-query attention with 1,024-wide key and value projections. Adapt all seven projections at rank 16. LoRA adds r times (in plus out) parameters per matrix; DoRA adds out more:

ProjectionShape (out x in)LoRA params, r=16DoRA magnitude
q_proj, o_proj4096 x 4096131,072 each4,096 each
k_proj, v_proj1024 x 409681,920 each1,024 each
gate_proj, up_proj14336 x 4096294,912 each14,336 each
down_proj4096 x 14336294,9124,096
Per layer1,310,72043,008
32 layers41.9 M1.38 M

The magnitude vectors add 3.3 percent to the trainable parameters, about 22 MB of extra weights, gradients and Adam state in fp32. That is not where the cost is. The cost is the norm: every forward pass of every adapted module builds a full out by in temporary. The largest here is 14,336 by 4,096, about 58.7 million elements or 117 MB in bf16, plus its fp32 squares if the norm is upcast. Those temporaries are freed module by module, but they raise peak memory and add memory traffic, and with gradient checkpointing they are rebuilt again during the backward recompute.

The arithmetic for the product B @ A is about 2 times 58.7 million times 16, roughly 1.9 GFLOP for that module, independent of how many tokens are in the micro-batch. The base GEMM for 4,096 tokens is about 2 times 4,096 times 58.7 million, roughly 480 GFLOP. So the fixed DoRA overhead is small for large micro-batches and dominant for tiny ones. If you train with micro-batches of a few hundred tokens, measure step time with and without use_dora before committing.

What it means on the GPU

Four practical consequences follow from that structure.

  • Bigger micro-batches amortise DoRA. The norm work happens once per module per forward, so packing more tokens per step reduces the relative overhead. Gradient accumulation does not help, because each micro-batch recomputes the norm.
  • Compute the norm in fp32. Sums of 4,096 or 14,336 squared bf16 values lose precision. The scratch layer upcasts; check what your PEFT version does under autocast.
  • Quantised bases need a dequantise. With a 4-bit base the norm requires a dequantised weight every step, which costs more than in bf16. Recent PEFT releases support DoRA on some quantised layer types; check the release notes for your version and measure memory.
  • Sharding works, at a cost. Under FSDP the full weight is gathered for the base GEMM anyway, so the norm reuses it, but the extra temporary still lands on each rank's peak. See full fine-tuning on GPUs for how sharding shapes memory.

Merging and serving

After training, merge. The merged weight is (m / ||W0 + s B A||) * (W0 + s B A), row by row, and the result is an ordinary linear layer with no inference overhead at all. This is the main reason DoRA is attractive: its training cost is paid once, and the served model is indistinguishable in shape from the base.

Unmerged DoRA is a different story. It cannot be stored as just two low-rank matrices, because the scale depends on the full updated weight. That matters for multi-adapter serving, where systems keep one base model and swap many small adapters per request; the multi-LoRA serving design assumes the adapter is purely additive. PEFT's documentation notes that mixed-adapter batches via adapter_names do not work with DoRA. Before planning to serve many DoRA adapters on one base, confirm that your serving engine supports DoRA adapters at all; otherwise merge and serve each as its own model, or use plain LoRA.

Merging into a quantised model is also lossy: you merge in bf16 and re-quantise, and the merged weights quantise differently from the originals. Evaluate the re-quantised model, not the bf16 one you merged.

Failure modes

  • Normalising the wrong axis when porting the formula between layouts. The model trains but underperforms. Test that a wrapped model reproduces the base outputs at step zero.
  • Gradient through the norm. Forgetting to detach keeps the full-size temporary alive for backward and gives up the paper's memory saving. Keep it detached unless you have measured a reason not to.
  • Comparing at different budgets. DoRA steps are slower; compare against LoRA at equal wall-clock time or equal tokens, and say which.
  • Magnitude drift. With a high learning rate, m can move a lot and change activation scales. Log the ratio of m to the original row norms per layer and look for outliers.
  • Evaluating the unmerged model only. Merge, reload and evaluate the artifact you will ship; the outputs should match the adapter model within bf16 tolerance.

Trade-offs: DoRA, LoRA or full fine-tuning

MethodTrainable params (8B, r=16)Training costServing
LoRAabout 42 Mlowest; thin GEMMs onlymerge, or hot-swap many adapters
DoRAabout 43 MLoRA plus a per-module weight norm each stepmerge; multi-adapter support varies
Full fine-tune8 Bweights, gradients and optimizer state for allone full model

Choose DoRA when you will merge anyway, you want low ranks to behave more like high ones, and your micro-batches are large enough to hide the norm cost. Choose plain LoRA when you need hundreds of adapters on one base or your steps are already memory-bound. For the operational side of either, see fine-tuning operations.

What to do next

  1. Run the scratch layer and its merge assertion; then remove the detach and measure the change in peak memory with torch.cuda.max_memory_allocated.
  2. Wrap your model with use_dora=True and confirm the step-zero outputs equal the base model's.
  3. Train LoRA and DoRA at r=8 and r=32 on the same data and token budget; record quality, step time and peak memory.
  4. Try two micro-batch sizes and see how the DoRA overhead changes.
  5. Merge, reload, evaluate the shipped artifact, and confirm your serving stack accepts it.
Key takeaway: DoRA writes each adapted weight as a trainable per-row magnitude times a unit direction, and uses LoRA only for the direction. It adds about three percent more trainable parameters, but every step rebuilds a full-size temporary to compute row norms, so it costs most with small micro-batches. Detach the norm and compute it in fp32. Merge before serving, which removes all overhead, and compare it with LoRA at equal budgets on your own data.