QLoRA fine-tunes a model whose frozen weights are stored in 4 bits while small low-rank adapters train in higher precision. This page builds the storage format with NumPy, checks it against the bitsandbytes source, measures the error on a real-sized matrix, and shows how to audit a loaded model.

Other pages own the rest of the subject. The theory, including the derivation of NF4, the memory proof and paged optimizers, is in QLoRA theory. A full training setup for larger models is in QLoRA fine-tuning explained, and when QLoRA is worth it for small models, with a 12 GB run, is in QLoRA for small language models. The NumPy experiments here were run with NumPy 2.4 and SciPy 1.17; the PyTorch code at the end is a sketch and audit tool that was not run for this article, and is labelled as such.

Advertisement

What a QLoRA layer stores

Take one linear layer with weight matrix W of shape out by in. Plain LoRA keeps W in 16-bit precision, frozen, and adds a trainable path B A x with a small rank r. QLoRA changes only how W is stored. In bitsandbytes, the library QLoRA uses through Hugging Face, the weight becomes a packed byte tensor holding two 4-bit indices per byte, and a QuantState object holds what is needed to turn the indices back into numbers. Reading its source, the fields are absmax (one scale per block), shape, code (the 16-entry lookup table), dtype (the original dtype to restore), blocksize, quant_type (nf4 or fp4), and, when nested statistics are on, offset and state2, the state for the quantized scales.

What a QLoRA linear layer stores, and what each forward pass doesFrozen base (bitsandbytes)packed uint8 weighttwo 4-bit indices per byteabsmax per 64 values8-bit codes if nestedquant_map, offset, state2NF4 table, DQ statedequantizeW in compute dtypetemporary, bf16x W^Tbase outLoRA A, Btrainable, 16/32-bit(alpha/r) x A^T B^Tadapter outsumBackward: the gradient flows to x and to A, B. The base is dequantized again; it never gets a gradient.Storage per base weight at block 64 with nested statistics: 4 + 8/64 + 32/(64 x 256) = 4.127 bits.
The base weight lives as packed 4-bit indices plus per-block scales. Each forward pass rebuilds a temporary weight in the compute dtype; only the adapters A and B receive gradients.

So the stored format is not one the GPU computes in: every forward pass dequantizes to the compute dtype, usually bf16. The scales cost memory too, which double quantization reduces. And the base never needs a gradient or optimizer state.

Step 1: the NF4 codebook

NF4 is a lookup table of 16 values in the range minus one to one. An index is a position in that table, not a sign, exponent and mantissa. The values are quantiles of a standard normal distribution, chosen so that if weights are roughly normal, each of the 16 levels is used about equally often. bitsandbytes generates them with a function called create_normal_map: 8 positive values from evenly spaced probabilities between 0.9677083 and 0.5, 7 negative values the same way, an exact zero, then everything divided by the largest value. The asymmetry, 8 positive and 7 negative, is the price of having an exact zero in 16 slots. The constant 0.9677083 is the outermost quantile; per the library's docstring it was tuned empirically.

import numpy as np
from scipy.stats import norm

def nf4_codebook(offset=0.9677083):
    pos = norm.ppf(np.linspace(offset, 0.5, 9)[:-1])     # 8 positive levels
    neg = -norm.ppf(np.linspace(offset, 0.5, 8)[:-1])    # 7 negative levels
    v = np.sort(np.concatenate([neg, [0.0], pos]))
    return (v / v.max()).astype(np.float32)

CODE = nf4_codebook()
MID = (CODE[1:] + CODE[:-1]) / 2       # 15 decision boundaries between levels

Running that and comparing with the 16 values hard-coded in bitsandbytes' functional.py gives a maximum difference of 1.19e-07, which is one unit of float32 rounding: the derivation is the table. The levels are dense near zero, where most weights are, and sparse near the edges, as the codebook line in the output shows.

Advertisement

Step 2: quantize in blocks and pack

A single scale for a whole matrix would let one large weight stretch the grid for everyone. QLoRA uses blocks instead: the flattened weight is cut into groups of 64 consecutive values, and each block is divided by its own absolute maximum so it fits in minus one to one. Each scaled value is then rounded to the nearest codebook level. bitsandbytes' reference backend does this with the 15 midpoints between adjacent levels and a bucketize call; searchsorted is the NumPy equivalent. Two indices are packed per byte, and the reference backend puts the first value of each pair in the high nibble.

def quantize_nf4(w, block=64):
    flat = w.reshape(-1, block).astype(np.float32)
    absmax = np.abs(flat).max(axis=1, keepdims=True)
    idx = np.searchsorted(MID, flat / absmax).astype(np.uint8)     # nearest level
    packed = (idx.reshape(-1)[0::2] << 4) | idx.reshape(-1)[1::2]  # first value in high nibble
    return packed, absmax.reshape(-1)

def dequantize_nf4(packed, absmax, shape, block=64):
    idx = np.empty(packed.size * 2, dtype=np.uint8)
    idx[0::2], idx[1::2] = packed >> 4, packed & 0xF
    return (CODE[idx].reshape(-1, block) * absmax[:, None]).reshape(shape)

rng = np.random.default_rng(0)
W = rng.normal(0, 0.02, size=(4096, 4096)).astype(np.float32)
packed, absmax = quantize_nf4(W)
Wq = dequantize_nf4(packed, absmax, W.shape)

For a 4096 by 4096 matrix of normally distributed weights with standard deviation 0.02, typical of a pretrained layer, the 16.8 million weights pack into 8,388,608 bytes, half a byte each, plus 262,144 float32 scales taking 1,048,576 bytes. Without double quantization the scales add 32 / 64 = 0.5 bits per weight, so storage is 4.5 bits per weight, not 4.

Step 3: measure the error

The full output of the experiment is below. The relative error, the Frobenius norm of the difference divided by the norm of W, is 0.092 for NF4. A plain 4-bit integer grid with the same blocks and absmax scaling, levels minus seven to seven, gives 0.108, so the normal-quantile grid is about 15 percent better on normally distributed data. The error in a matrix product with random inputs is the same 0.092, because the rounding errors are close to independent and do not cancel or compound much in one product.

max |derived - bnb| = 1.1920928955078125e-07
codebook: -1.0000 -0.6962 -0.5251 -0.3949 -0.2844 -0.1848 -0.0910 +0.0000 +0.0796 +0.1609 +0.2461 +0.3379 +0.4407 +0.5626 +0.7230 +1.0000
packed bytes: 8388608 absmax fp32 bytes: 1048576
NF4 rel error:       0.0920
int4 absmax rel err: 0.1076
NF4 + DQ rel error:  0.0920
matmul rel error:    0.0918
level usage %: 1.9 4.4 5.8 7.1 8.1 8.8 9.2 8.8 8.1 7.8 7.3 6.7 5.8 4.8 3.6 1.7
with outliers rel error: 0.1020
non-outlier error inside outlier blocks: 0.6101
zeroed fraction inside outlier blocks: 0.713

The level-usage line shows the levels are not used equally: the extremes get under 2 percent each, the middle about 9. Per-block absmax scaling maps each block's largest value exactly to plus or minus one, so the outer levels serve only a block's maximum and its few neighbours. A 9 percent perturbation sounds large, but the adapters are trained on the quantized model's actual outputs and compensate for the error that matters to the task, which is also why the merge is delicate.

Block size: the dial nobody turns

Block sizeRelative errorBits per weight, fp32 scalesBits per weight, nested scales
320.08735.0004.254
64 (default)0.09204.5004.127
1280.09564.2504.063
2560.09874.1254.032

Errors come from the same NumPy run; bit counts are arithmetic. Smaller blocks track local scale better but need more scales. bitsandbytes accepts block sizes from 32 to 4096. The differences between 32 and 128 are small next to the effect of outliers.

Outliers: one value wrecks its block

Pretrained transformers contain a small number of weights far larger than the rest. The experiment simulates this by setting 2,000 random weights in the matrix to 0.5, about 25 standard deviations. The overall error rises only from 0.092 to 0.102, which looks harmless. Inside the blocks that contain an outlier, it is a different picture: the error on the ordinary weights is 0.61, and 71 percent of them are rounded to exactly zero.

The outlier sets the block's absmax to 0.5. After scaling, zero's bucket runs from -0.0455 to 0.0398, the midpoints to its neighbours, so any weight between -0.0228 and 0.0199 rounds to zero: about one standard deviation. For weights drawn with standard deviation 0.02 that predicts 71.3 percent, matching the measurement. Real models have milder outliers than this test, but if a layer behaves badly after quantization, look for large-magnitude weights first. Options include a smaller block size for that layer, leaving the layer unquantized, or methods that handle outliers explicitly, compared in NF4 quantization.

Double quantization, as implemented

Nested statistics, the bnb_4bit_use_double_quant option, quantize the scales themselves. The QLoRA paper describes this as quantizing the first-level constants to 8 bits with a block size of 256. The bitsandbytes source adds a detail: it subtracts the mean of all absmax values first, stores that mean as offset, then quantizes the centred values in blocks of 256 with its 8-bit dynamic code, keeping one float32 scale per 256 scales in state2. Storage becomes 8 / 64 bits per weight for the scales plus 32 / (64 x 256) for their scales, giving the familiar 4.127 bits in total.

The error cost is negligible: replacing the scales with a linear int8 stand-in for bitsandbytes' dynamic code leaves the relative error at 0.0920 to four places. On a 7B-class model it saves about 0.37 bits per weight, a few hundred megabytes.

The forward and backward pass

The sketch below shows the computational shape of a QLoRA layer in PyTorch. It was not run for this article; real implementations use fused CUDA kernels and handle many more cases. The two decisions that matter are visible. Forward saves the packed 4-bit data, not the dequantized weight, so activation memory does not include a full-precision copy of W; backward dequantizes again. And backward returns a gradient for the input only, so the base is frozen by construction rather than by a flag. B starts at zero, so the model initially computes exactly the quantized base.

# Sketch in PyTorch (not run for this article): the shape of a QLoRA linear layer.
import torch

class NF4MatMul(torch.autograd.Function):
    @staticmethod
    def forward(ctx, x, packed, absmax, shape):
        W = dequantize(packed, absmax, shape).to(x.dtype)   # temporary full-size weight
        ctx.save_for_backward(packed, absmax)               # save 4-bit data, not W
        ctx.shape = shape
        return x @ W.t()

    @staticmethod
    def backward(ctx, grad_out):
        packed, absmax = ctx.saved_tensors
        W = dequantize(packed, absmax, ctx.shape).to(grad_out.dtype)
        return grad_out @ W, None, None, None               # grad for x only, never for W

class QLoRALinear(torch.nn.Module):
    def __init__(self, packed, absmax, shape, r=16, alpha=32):
        super().__init__()
        out_f, in_f = shape
        self.register_buffer("packed", packed)             # buffers: no grad, no optimizer
        self.register_buffer("absmax", absmax)
        self.shape = shape
        self.A = torch.nn.Parameter(torch.randn(r, in_f) / r)   # small random init
        self.B = torch.nn.Parameter(torch.zeros(out_f, r))      # zero: starts as the base
        self.scale = alpha / r

    def forward(self, x):
        base = NF4MatMul.apply(x, self.packed, self.absmax, self.shape)
        return base + self.scale * (x @ self.A.t() @ self.B.t())

So every step dequantizes each base weight at least twice, three times with gradient checkpointing, work that plain bf16 LoRA skips; compare LoRA from scratch.

Auditing a loaded model

Configuration mistakes in QLoRA are silent: the model loads, trains and produces a loss curve whether or not the layers you expected are quantized. Before a long run, audit. The function below walks a model loaded with a 4-bit BitsAndBytesConfig, counts Linear4bit modules by quantization settings, and totals memory by kind. It uses the QuantState fields named in the bitsandbytes source; like the sketch above, it was not run for this article, so treat the output format as illustrative.

# Audit a model loaded with BitsAndBytesConfig(load_in_4bit=True, ...).
# Uses bitsandbytes' Linear4bit and the QuantState fields from its source.
import collections
import torch
import bitsandbytes as bnb

def audit(model):
    seen, bytes_by_kind = collections.Counter(), collections.Counter()
    for name, mod in model.named_modules():
        if isinstance(mod, bnb.nn.Linear4bit):
            qs = mod.weight.quant_state
            seen[(qs.quant_type, qs.blocksize, qs.nested, str(qs.dtype))] += 1
            bytes_by_kind["4-bit packed"] += mod.weight.numel() * mod.weight.element_size()
    for name, p in model.named_parameters():
        if p.dtype not in (torch.uint8,) and p.requires_grad:
            bytes_by_kind["trainable " + str(p.dtype)] += p.numel() * p.element_size()
        elif p.dtype not in (torch.uint8,):
            bytes_by_kind["frozen " + str(p.dtype)] += p.numel() * p.element_size()
    for k, v in seen.items():
        print("Linear4bit", k, "x", v)
    for k, v in bytes_by_kind.items():
        print(f"{k:24s} {v / 2**30:6.2f} GiB")

Look for exactly one settings tuple matching what you asked for, a frozen total that is mostly embedding, head and norms, and a trainable total equal to your calculated LoRA parameter count. A large frozen fp32 total means preparation code upcast non-quantized parameters.

Failure modes

  • Merging into the wrong base. Adapters were trained against the dequantized NF4 weights. Adding B A to the original bf16 weights gives a slightly different model from the one you evaluated. Merge into the dequantized base, then evaluate the merged artifact.
  • Re-quantizing after merge. Quantizing the merged model again, to NF4 or a GGUF format, adds a second rounding. Evaluate the final file, not the training checkpoint.
  • Silent outlier layers. A layer whose blocks contain large weights loses most of its small weights to zero; overall error looks fine. Measure per-layer error before blaming the data.
  • Compute dtype mismatch. float16 compute on bf16-trained models can overflow; prefer bf16.
  • Assuming everything is 4-bit. Embeddings, the head and norms stay in higher precision, and on small models they are a large share of memory.
  • Version drift. Quantized weights serialise with their QuantState; pin library versions and save adapters separately.

What to do next

  1. Run the NumPy codebook and quantizer on one weight matrix from your own model and compare its error with the 0.092 baseline here.
  2. Compute per-block absmax statistics for every linear layer and flag layers whose maximum is many standard deviations above the mean.
  3. Run the audit function after loading, and stop the run if any layer you expected is missing from the 4-bit count.
  4. Keep block size 64 and nested statistics on unless a measured layer error says otherwise.
  5. After training, merge into the dequantized base, then evaluate the exact file you will ship.
  6. Benchmark step time against bf16 LoRA on the same data so you know what the memory saving costs.
Key takeaway: QLoRA stores each frozen base weight as a 4-bit index into a 16-value normal-quantile table, with one scale per block of 64 weights, and dequantizes to bf16 for every matrix multiply. Built from scratch, the NF4 table matches bitsandbytes exactly and gives 9.2 percent relative error on normal weights, better than a 4-bit integer grid. Double quantization brings storage to 4.127 bits per weight at no measurable error cost. Large outliers are the real hazard: they flatten their block's other weights to zero. Audit what a loaded model actually quantized, and evaluate the merged artifact you ship, not the training checkpoint.