QLoRA is the reason fine-tuning a 7B or 8B model on a single consumer or workstation GPU stopped being a stunt. The idea, from Dettmers et al.'s 2023 paper "QLoRA: Efficient Finetuning of Quantized LLMs", is simple to state: store the pretrained weights in 4 bits and freeze them, attach small trainable low-rank adapters in 16 or 32 bits, and backpropagate through the quantized model into the adapters. The paper fine-tuned a 65B model on one 48 GB GPU.
This site already explains the NF4 data type itself in NF4 and QLoRA and the arithmetic of double quantization, and LoRA from scratch in LoRA, low-rank adaptation. This article does not re-derive the codebook. It follows the training step: what is computed in which precision, why gradients reach the input but never the base weights, where every gigabyte goes for an 8B model, what paged optimizers actually do, and the trap at the end of the process, merging adapters that were trained against quantized weights back into a model that is not quantized.
What is frozen, what is trained
Start with one linear layer, y = x W^T. Full fine-tuning stores W, its gradient and optimizer state for every element. For Adam with mixed precision that is commonly estimated at around 16 bytes per parameter before activations, so an 8B model needs well over 100 GB just for weights and state.
LoRA replaces the update with a low-rank one. W stays frozen and the layer computes y = x W^T + s (x A^T) B^T, where A is r by in, B is out by r, the rank r is small (8 to 64), and s = lora_alpha / r. B starts at zero so the model is unchanged at step zero. Only A and B get gradients and optimizer state, and for a 4096 by 4096 projection with r = 16 that is 131,072 trainable numbers instead of 16.8 million.
QLoRA makes one further change: the frozen W is stored in 4-bit NF4, in blocks of 64 weights that each share one scale (the block's absolute maximum), and those scales are themselves quantized to 8 bits in groups of 256. That is double quantization, and it brings the storage cost to about 4.127 bits per weight (4 + 8/64 + 32/(64 x 256)). The adapters are not quantized. They are ordinary floating-point tensors, which is what makes training them with ordinary optimizers possible.
One training step, numerically
In the forward pass, bitsandbytes dequantizes each block of W into the compute dtype you chose (bfloat16 is the usual choice), runs a normal matmul with the bf16 activations, and discards the dequantized copy. The adapter path is computed alongside and added. Storage is 4-bit; arithmetic is 16-bit.
The backward pass is where people's intuition fails. The base weights are frozen, so there is no gradient with respect to W. But the layer below still needs the gradient with respect to its output, which is this layer's input x, and computing grad_x requires multiplying by W again. So W is dequantized a second time in the backward pass. Gradients therefore flow through every frozen 4-bit layer, all the way down to the first adapter, and that is what lets adapters in early layers learn at all.
# What one QLoRA linear layer does, written out (pseudocode, not the bitsandbytes kernel)
def forward(x, Wq, consts, A, B, scale):
W = dequantize_nf4(Wq, consts, dtype=bf16) # per 64-weight block: code[idx] * absmax
base = x @ W.T # W is discarded after use, never stored
lora = (dropout(x) @ A.T) @ B.T # rank-r detour, tiny matmuls
return base + scale * lora # scale = lora_alpha / r
def backward(grad_y, x, Wq, consts, A, B, scale):
W = dequantize_nf4(Wq, consts, dtype=bf16) # dequantize again: nothing was cached
grad_x = grad_y @ W + scale * (grad_y @ B) @ A # needed by the layer below
grad_B = scale * grad_y.T @ (x @ A.T)
grad_A = scale * (grad_y @ B).T @ x
# no grad_W: the base weights are frozen and stay 4-bit
return grad_x, grad_A, grad_BTwo consequences follow. First, a QLoRA step does more work than a bf16 LoRA step on the same model: every linear layer dequantizes twice per step, and three times if gradient checkpointing recomputes the forward. Expect lower throughput than bf16 LoRA; measure it. Second, the adapters learn to correct the quantized model, not the original one. The function being trained is dequant(Q(W)) + s B A. Remember that; it matters when you merge.
Where the memory goes: an 8B model, byte by byte
Take a Llama-3-8B-shaped model: 32 layers, hidden size 4096, MLP size 14336, grouped-query attention with 8 key-value heads (so the K and V projections are 4096 by 1024), and a 128,256-token vocabulary. Each layer's seven linear projections hold 218.1 million weights, so the 32 layers hold about 6.98 billion. The input embedding and output head add about 1.05 billion more. Total, about 8.03 billion.
The transformers bitsandbytes integration quantizes nn.Linear modules and, by default, leaves the output head in higher precision; the embedding is not a linear layer and stays in bf16 as well. Check what your version did by printing the model: quantized layers show up as Linear4bit.
| Item | How it is computed | Approximate size |
|---|---|---|
| Base linear layers, NF4 + double quant | 6.98 B x 4.127 bits / 8 | 3.6 GB |
| Embedding + output head | 1.05 B x 2 bytes in bf16; x 4 once prepare_model_for_kbit_training upcasts them | 2.1 GB, then 4.2 GB |
| LoRA adapters, r = 16 on all seven projections | 1.31 M per layer x 32 = 41.9 M params, fp32 | 0.17 GB |
| Adapter gradients, fp32 | same count | 0.17 GB |
| AdamW state (two moments), fp32 | 2 x 41.9 M x 4 bytes | 0.34 GB |
| Activations, checkpointed, batch 1 x 4096 tokens | layer inputs kept, one layer recomputed at a time | roughly 1 to 3 GB |
| Logits for the loss, 4096 x 128,256 | bf16, more if upcast to fp32 | 1 to 2 GB |
Two lessons sit in that table. The quantized weights are not the whole story; on a large-vocabulary model the unquantized embedding and head are more than half the weight memory of the 4-bit body. And the trainable state is small: the adapters, their gradients and Adam's moments together are under 0.7 GB. What decides whether a run fits is activations and logits, which scale with sequence length times batch size. That is why gradient checkpointing is effectively mandatory.
Paged optimizers: what they buy and what they do not
The third QLoRA ingredient is paged optimizer state. Optimizer tensors are allocated in NVIDIA unified memory, so when the GPU runs short the driver can evict pages to CPU RAM and bring them back when the optimizer step touches them. The motivation in the paper is memory spikes: a batch with an unusually long sequence can briefly push activation memory over the limit, and without paging that kills the run.
Paging is a shock absorber, not extra capacity. If the steady-state working set does not fit, every step pages and throughput collapses. In the Hugging Face trainer the options include paged_adamw_8bit and paged_adamw_32bit. If you see constant paging, shorten sequences and raise gradient accumulation instead.
A working setup
The standard stack is transformers for the model, bitsandbytes for the 4-bit kernels and PEFT for the adapters. The configuration below is the one most QLoRA runs use; each line is commented with what it controls.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training
model_id = "meta-llama/Meta-Llama-3-8B" # any causal LM you are licensed to use
bnb = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4", # NF4 codebook (the alternative is "fp4")
bnb_4bit_use_double_quant=True, # quantize the per-block absmax constants too
bnb_4bit_compute_dtype=torch.bfloat16 # dtype the blocks are dequantized into for matmuls
)
model = AutoModelForCausalLM.from_pretrained(model_id, quantization_config=bnb, device_map={"": 0})
tok = AutoTokenizer.from_pretrained(model_id)
# Casts every non-4-bit fp16/bf16 parameter (norms, embedding, head) to fp32,
# enables input gradients and gradient checkpointing.
model = prepare_model_for_kbit_training(model, use_gradient_checkpointing=True)
lora = LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05,
target_modules="all-linear", # every linear layer except the output head
task_type="CAUSAL_LM",
)
model = get_peft_model(model, lora)
model.print_trainable_parameters() # sanity check: should be well under 1% of the totalThree details matter. Use bfloat16 compute where the GPU supports it; float16 is more prone to overflow. prepare_model_for_kbit_training is not optional decoration: it upcasts every parameter that is not 4-bit (norms, embedding, head) to fp32 for stability and turns on input gradients so the checkpointed graph connects. And target_modules="all-linear" follows the paper's finding that adapting only the attention query and value projections, as the original LoRA paper did, was not enough to match 16-bit full fine-tuning; adapting every linear layer was.
Choosing rank, alpha, targets and learning rate
Rank r sets capacity. For instruction tuning and style adaptation, r = 8 to 16 is usually enough; for teaching new domain knowledge, r = 32 to 64 is a common starting point. The QLoRA paper's main runs used r = 64 with alpha = 16. The effective update is scaled by alpha / r, so if you change r and want the same update magnitude, change alpha with it; many practitioners simply keep alpha = 2r.
Adapter learning rates are far higher than full fine-tuning ones: around 1e-4 to 2e-4 with AdamW and a short warmup, against 1e-5 or below. Too high shows up as an early loss spike or quickly lost general ability.
The merge mismatch
After training you have a 4-bit base and an adapter. You can serve them unmerged, with the adapter applied at runtime, which is how multi-adapter servers work (see serving many LoRA adapters). Or you can merge, folding s B A into the base weights to get a single dense model with no adapter overhead.
Here is the subtlety. The adapter was trained so that dequant(Q(W)) + s B A behaves well. The common merge recipe loads the original bf16 W and computes W + s B A. Those differ by exactly the quantization error W - dequant(Q(W)), which the adapter may have partly learned to compensate. The difference is usually small but not zero.
import torch
from transformers import AutoModelForCausalLM
from peft import PeftModel
# Option A (common): load the ORIGINAL bf16 base and merge. Produces W + s*B*A.
base = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16)
merged = PeftModel.from_pretrained(base, "out/adapter").merge_and_unload()
merged.save_pretrained("out/merged-bf16")
# Option B (matches training exactly): merge into dequantize(Q(W)) instead of W,
# i.e. dequantize each Linear4bit weight to bf16 first. Evaluate either result
# against the unmerged adapter on the same prompts before shipping.You cannot merge into the 4-bit weights and stay 4-bit without re-quantizing, because W + s B A is not on the NF4 grid. If you want a 4-bit deployment, merge into bf16 first and then quantize the merged model with a method built for inference, such as GPTQ, AWQ or GGUF (the options are compared in bitsandbytes vs AutoAWQ vs AutoGPTQ). That is a second quantization with its own error, so evaluate after it as well.
Initialisation schemes such as LoftQ attack the same mismatch from the other end, choosing the quantized base and initial adapters jointly so that Q(W) + B A approximates W from step zero. PEFT exposes it as an option; measure before adopting it.
Failure modes and how to recognise them
- Out of memory after loading fine. Activations and logits, not weights. Enable gradient checkpointing, cut max sequence length or micro-batch size, and accumulate gradients.
- Loss does not move. Adapters not attached to anything, or gradients not connected.
print_trainable_parameters()reporting zero, or a checkpointed model without input gradients, are the usual causes. - NaN or inf loss. float16 compute dtype overflowing, or a learning rate too high for the rank and alpha. Switch to bfloat16 and halve the LR.
- Merged model worse than adapter. The merge mismatch above, or merging into a different base revision than you trained on. Compare against the unmerged adapter on fixed prompts.
Trade-offs: QLoRA, bf16 LoRA and full fine-tuning
| QLoRA | LoRA on bf16 base | Full fine-tune | |
|---|---|---|---|
| Weight memory (8B) | about 5.7 GB at load, 7.8 GB after the fp32 upcast | about 16 GB | about 16 GB plus gradients |
| Trainable state | under 1 GB | under 1 GB | 100 GB-plus with AdamW |
| Step speed | slowest of the three per token on the same GPU | faster | fastest per token if it fits |
| Quality | close to 16-bit in the paper's tests | reference for adapters | highest ceiling, highest forgetting risk |
| Deployment | adapter, or merge then re-quantize | adapter or merge | a new full checkpoint |
Pick QLoRA when memory is the binding constraint. If the bf16 model fits with room for activations, bf16 LoRA trains faster and avoids the merge mismatch entirely. Full fine-tuning is for when you have the hardware and need the model's behaviour to move a long way.
What to do next
- Load your base model in 4-bit with the configuration above and print it; confirm which modules became Linear4bit and which stayed bf16.
- Fill in the memory table for your model and your sequence length before buying or renting a GPU; the logits row is the one people forget.
- Run 200 steps with r = 16, alpha = 32, LR 2e-4 and gradient checkpointing; check that loss falls and that
print_trainable_parameters()is non-zero. - Build a fixed evaluation set of production-formatted prompts, and score base, adapter and merged models on it.
- Merge into bf16, re-evaluate, and only then quantize for serving; evaluate once more after that step.
- If memory is not tight, rerun the same recipe as bf16 LoRA and compare quality and wall-clock time.