QLoRA, introduced by Tim Dettmers and colleagues in 2023, fine-tunes a language model whose frozen base weights are stored in 4 bits while small low-rank adapters are trained in 16-bit precision. Gradients flow through the quantized base, but only the adapters change. The paper's headline was fitting a 65-billion-parameter model onto a single 48 GB GPU.
Small language models change the question. A 1B or 3B model already fits comfortably in 16-bit precision on most training GPUs, so the case for QLoRA is no longer 'it is the only way to fit'. It becomes a trade: less memory for weights against slower steps, a small quality cost, and a more awkward path to a deployable artifact. This article works through that trade for models from 0.5B to 8B parameters. It covers where the memory actually goes, when QLoRA wins, a complete training run on a 12 GB-class GPU, and how to merge and re-quantize the result for on-device serving without losing the quality you trained for.
The mechanism in brief
Each linear layer's weight matrix W is stored in NF4, a 4-bit data type whose 16 levels are placed at quantiles of a normal distribution, which suits the roughly normal distribution of pretrained weights. W is split into blocks of 64 values, and each block has its own scale. Double quantization then stores those scales in 8 bits, with a second-level 32-bit scale for every 256 blocks. The result costs about 4.127 bits per weight instead of 4.5 without double quantization.
On every forward pass the layer dequantizes W to the compute dtype, usually bf16, and computes W x plus the adapter path (alpha / r) times B A x. A and B are the LoRA matrices, with rank r far smaller than the layer width. Backward computes gradients for A and B and for the activations, but never for W. The derivation, the NF4 construction and a full memory proof are in QLoRA theory; this page is about applying it to small models.
Where the memory goes in a small model
The bitsandbytes integration replaces linear layers with 4-bit versions. It does not quantize the token embedding, the norms, or the output head. On a 70B model that hardly matters. On a small model with a large vocabulary it matters a great deal, because the embedding matrix is vocabulary size times hidden width, and modern tokenizers have vocabularies of 128K to 152K entries.
| Model shape | Parameters | Embedding and head params | bf16 weights | QLoRA weights | Embedding share of QLoRA |
|---|---|---|---|---|---|
| Qwen2.5-0.5B class, 151,936 x 896, tied | 0.49B | 0.14B | 0.99 GB | 0.46 GB | about 60 percent |
| Llama-3.2-1B class, 128,256 x 2,048, tied | 1.24B | 0.26B | 2.47 GB | 1.03 GB | about 51 percent |
| Llama-3.2-3B class, 128,256 x 3,072, tied | 3.21B | 0.39B | 6.43 GB | 2.24 GB | about 35 percent |
| Llama-3.1-8B class, untied embedding and head | 8.03B | 1.05B | 16.06 GB | 5.70 GB | about 37 percent |
The QLoRA column assumes 4.127 bits for every parameter outside the embedding and head, and 16 bits for those. The pattern is clear. At 1B, half of the quantized model's weight memory is an embedding that QLoRA never touched, and the total saving over bf16 is about 1.4 GB. At 8B the saving is over 10 GB, which is the difference between fitting on a 16 GB card and not.
Weights are also not the whole bill. Activations grow with batch size and sequence length, and for small models with big vocabularies the logits tensor is often the largest single allocation: 2,048 tokens times 128,256 vocabulary entries in fp32 is 0.98 GiB per sequence. TRL's default chunked_nll loss avoids materialising the full tensor, and gradient checkpointing, on by default in SFTConfig, trades recomputation for activation memory. PEFT's prepare_model_for_kbit_training, when it is applied, also upcasts the non-quantized parameters to fp32 for stability, which on a 1B model adds roughly another half gigabyte for the embedding alone.
The adapters are small. With rank 16 on every linear layer of a 3B model with 28 layers, a 3,072 hidden width, 8 key-value heads and an 8,192 MLP, LoRA adds 24.3 million parameters, under 1 percent of the model. Their fp32 weights and optimizer states come to a few hundred megabytes at most.
A decision rule: QLoRA or plain LoRA
Because the saving shrinks with model size, the decision depends on the ratio of model to GPU. The table is a starting rule of thumb, not a law. Measure your own peak memory with the sequence length you actually need.
| Model | GPU memory | Recommendation |
|---|---|---|
| 0.5B to 1B | Any training GPU with 8 GB or more | Plain LoRA in bf16, or full fine-tuning. QLoRA saves little and costs speed. |
| 3B to 4B | 24 GB or more | Plain LoRA in bf16. |
| 3B to 4B | 8 to 16 GB, or long contexts | QLoRA. The weight saving buys batch size or sequence length. |
| 7B to 8B | 24 GB or less | QLoRA. |
| 7B to 8B | 40 GB or more | Plain LoRA, unless you need very long sequences. |
Speed is the other side of the trade. Dequantizing every linear weight on every forward pass, and again during recomputation under gradient checkpointing, adds work that plain LoRA does not do, so QLoRA steps are noticeably slower on the same hardware. Quality is usually close to 16-bit LoRA, and the paper reported matching 16-bit fine-tuning on its benchmarks, but the base the adapters see is slightly perturbed, which matters later when you merge.
A complete run on a 12 GB-class GPU
The worked example fine-tunes a 3B instruct model to answer support tickets in a house format. The data is a JSONL file of prompt-completion pairs in chat form, so TRL computes the loss on completions only by default. By the ledger above, the 4-bit weights take about 2.2 GB, leaving room for activations at 2,048 tokens with a per-device batch of 4 and checkpointing on.
import torch
from datasets import load_dataset
from peft import LoraConfig
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from trl import SFTConfig, SFTTrainer
MODEL = "meta-llama/Llama-3.2-3B-Instruct"
bnb = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16, # use float16 on GPUs without bf16
)
model = AutoModelForCausalLM.from_pretrained(
MODEL, quantization_config=bnb, dtype=torch.bfloat16, device_map={"": 0})
tok = AutoTokenizer.from_pretrained(MODEL)
peft_cfg = LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05,
target_modules="all-linear", # every linear layer except the output head
task_type="CAUSAL_LM",
)
# prompt-completion rows: {"prompt": [...messages], "completion": [...messages]}
data = load_dataset("json", data_files="tickets.jsonl", split="train").train_test_split(0.05, seed=0)
args = SFTConfig(
output_dir="out/qlora-3b",
per_device_train_batch_size=4,
gradient_accumulation_steps=4, # 16 sequences per optimizer step
learning_rate=2e-4, # adapters need far more than the 2e-5 default
lr_scheduler_type="cosine",
warmup_steps=30,
num_train_epochs=2,
max_length=2048,
optim="paged_adamw_8bit",
bf16=True,
logging_steps=10,
eval_strategy="steps", eval_steps=100, save_steps=100,
)
trainer = SFTTrainer(model=model, args=args, processing_class=tok,
train_dataset=data["train"], eval_dataset=data["test"],
peft_config=peft_cfg)
trainer.train()
trainer.save_model("out/qlora-3b/adapter") # adapter weights only, tens of megabytesA few choices deserve explanation. The compute dtype is bf16. GPUs without bf16 support, such as the T4, need torch.float16 instead, with the loss-scaling care that fp16 brings. target_modules="all-linear" puts adapters on every attention and MLP projection. The QLoRA paper found that adapting all linear layers was necessary to match full fine-tuning, and that attention-only adapters fell short. The paged 8-bit AdamW optimizer keeps its states in 8 bits and uses unified memory so that a transient spike pages optimizer state to CPU memory instead of crashing. With adapters this small the optimizer is rarely the problem, but the paging protects against long-sequence spikes.
Hyperparameters that matter more at small scale
- Learning rate. Adapters need a much higher rate than full fine-tuning. The paper used 2e-4 for its 7B and 13B runs, and values from 1e-4 to 2e-4 are a sensible start. Watch the eval loss for the first few hundred steps and lower the rate if it rises.
- Rank and alpha. Rank 8 to 32 covers most task adaptation. Raising rank beyond that rarely helps a small model and increases overfitting. Keep alpha proportional to rank so the effective scale does not change when you vary r.
- Epochs. Small models memorise quickly. One to three epochs over a few thousand good examples is typical, and a rising eval loss while train loss falls is the signal to stop.
- Chat template and end-of-turn token. Train with the template the model will be served with, and make sure completions end with the end-of-turn token. A model trained without it learns to keep talking.
- Loss masking. Compute loss on completions only. Training on prompts teaches the model to generate your prompts.
Shipping: merge, re-quantize, verify
Training produced an adapter that was fitted against the dequantized NF4 weights, call them W', not against the original weights W. You now have three ways to serve it, and each one sees slightly different weights.
| Option | Weights at inference | Trade-off |
|---|---|---|
| Keep the 4-bit base plus a separate adapter | Exactly what training saw | Needs a runtime that supports this combination; slower than a merged model |
| Merge into the 16-bit original W | W + BA instead of W' + BA | Standard path; a small mismatch that is usually harmless but must be measured |
| Merge into the 4-bit model | Dequantize, add BA, re-quantize to 4 bits | PEFT supports it, but rounding the sum adds error; avoid for production |
For on-device serving the usual path is to merge into the 16-bit original, convert to GGUF, and quantize with a llama.cpp scheme such as Q4_K_M. That applies a second, different quantization to weights that were trained against the first. Most of the time the result is fine. Sometimes a fine-tune that looked strong in training loses a few points after re-quantization, especially when the task depends on exact formats. The only defence is to evaluate the artifact you ship.
# Merge into the 16-bit original weights, then build a GGUF for on-device serving.
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL = "meta-llama/Llama-3.2-3B-Instruct"
base = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16)
merged = PeftModel.from_pretrained(base, "out/qlora-3b/adapter").merge_and_unload()
merged.save_pretrained("out/merged-3b")
AutoTokenizer.from_pretrained(MODEL).save_pretrained("out/merged-3b")
# then, in a llama.cpp checkout:
# python convert_hf_to_gguf.py out/merged-3b --outfile merged-3b-f16.gguf --outtype f16
# ./llama-quantize merged-3b-f16.gguf merged-3b-Q4_K_M.gguf Q4_K_MRun the same held-out set against three artifacts: the adapter on the 4-bit base, the merged bf16 model, and the final GGUF. The first number is what training achieved. The gap to the last is what shipping cost you. SLM evaluation covers building the regression gate, and the GGUF runtime explains what the quantized file does on the device.
Failure modes
- bf16 on hardware without it. Setting bf16 compute on a GPU that lacks bf16 support fails or runs very slowly. Check before a long run.
- Out of memory at step one. Usually logits or activations, not weights. Reduce sequence length or per-device batch, keep checkpointing on, and keep the chunked loss.
- Attention-only adapters. Copying an old configuration with
q_projandv_projonly gives a weaker model. Use all linear layers. - Run-on generations. The end-of-turn token was missing from training completions, or the serving template differs from the training template.
- Quality lost at export. The merged or re-quantized model scores lower than the adapter did. Evaluate each artifact and consider a less aggressive GGUF quantization such as Q5_K_M or Q8_0 when the gap matters.
- Paging thrash. When the paged optimizer pages constantly, steps slow sharply. That signals the configuration is too close to the memory limit.
- Wrong conclusion from a small model. Choosing QLoRA for a 1B model and blaming the method for slow steps. At that size plain LoRA is usually the right tool.
Trade-offs and alternatives
QLoRA buys memory with speed and a small amount of fidelity. For 7B and 8B models on consumer or mid-range GPUs that is an excellent trade, and often the only one. For models of 3B and below with a reasonable GPU, plain LoRA in bf16 trains faster, avoids the train-serve weight mismatch and is simpler to export, as covered in LoRA for small models. For the smallest models, full fine-tuning is affordable and sometimes better, and SLM fine-tuning compares all the options with their memory math. When the base is already quantized for a device, remember that training quantization and serving quantization are separate decisions. Edge quantization covers the serving side.
What to do next
- Compute the weight ledger for your model, including the embedding and head at 16 bits, and compare the saving with your GPU's memory.
- If the model is 3B or smaller and you have 16 GB or more, try plain bf16 LoRA first and keep QLoRA as the fallback.
- Build a prompt-completion dataset with your serving chat template, and confirm every completion ends with the end-of-turn token.
- Train with rank 16 on all linear layers at 1e-4 to 2e-4, evaluate every hundred steps, and stop when eval loss turns up.
- Merge into the 16-bit original, convert to GGUF, and quantize.
- Evaluate the 4-bit adapter, the merged model and the GGUF on the same held-out set, and ship only if the final artifact passes your gate.