Quantizing a weight to 4 bits saves nothing until those 4 bits are stored next to other 4-bit values with no gaps. Memory is addressed in bytes, and the common checkpoint formats store low-bit weights as ordinary uint8 or int32 tensors, so every low-bit format invents its own way to squeeze several values into a byte or a 32-bit word. That is bit packing, and it is the layer where quantized models most often break silently: a loader that unpacks along the wrong axis or in the wrong order produces numbers, just the wrong ones.
This article covers the storage and interoperability side of packing from first principles: the arithmetic, the design choices every format makes, the actual layouts of GPTQ, AWQ, GGUF and NF4 checkpoints, packing widths that do not divide 32, how much a weight really costs once scales and zeros are counted, how to convert between formats and how to prove a conversion is correct. How kernels turn packed values back into fp16 in registers is a separate topic, covered in INT4 kernel architecture.
The packing arithmetic
A b-bit unsigned code q is an integer from 0 to 2b - 1. To put the code for logical element i into a stream of 32-bit words, compute its bit position pos = i * b, its word pos // 32 and its offset pos % 32, then OR the code shifted left by that offset into the word. Unpacking reverses it: shift right by the offset and mask with (1 << b) - 1. When b divides the word size, as 2, 4 and 8 do, every value sits inside one word: eight 4-bit values per int32, two per byte. When it does not, as with 3 or 6 bits, some values straddle two words.
import numpy as np
def pack_bits(q, bits):
"""Reference packer: logical element i occupies bits [i*bits, (i+1)*bits) of a
little-endian 32-bit word stream. len(q) * bits must be a multiple of 32."""
words = [0] * (len(q) * bits // 32)
for i, v in enumerate(q):
w, off = divmod(i * bits, 32)
words[w] |= (int(v) << off) & 0xFFFFFFFF
if off + bits > 32: # value straddles two words
words[w + 1] |= int(v) >> (32 - off)
return np.array(words, dtype=np.uint32)
def unpack_bits(words, bits, n):
mask, out = (1 << bits) - 1, np.empty(n, dtype=np.uint8)
for i in range(n):
w, off = divmod(i * bits, 32)
v = int(words[w]) >> off
if off + bits > 32:
v |= int(words[w + 1]) << (32 - off)
out[i] = v & mask
return outThis loop is slow but unambiguous, which makes it the right oracle for testing fast vectorised packers. Signed weights are stored as unsigned codes with an offset: either a fixed one, such as 8 for 4-bit symmetric codes, or a per-group zero point. The trade-offs between the two are covered in symmetric versus asymmetric quantization.
Five decisions every format makes
| Decision | Options | What it affects |
|---|---|---|
| Container | uint8 bytes, int32 words, fixed-size blocks | Alignment, how a kernel loads data, how the tensor appears in safetensors |
| Pack axis | Along the input dimension K, or the output dimension N | Which values a thread gets in one load; whether groups along K stay contiguous |
| Slot order | Sequential, interleaved, or split-half | Whether a cheap bit trick emits values in the order the kernel needs |
| Offset | Fixed constant, or per-group zero point | Accuracy for skewed weights and extra metadata |
| Metadata layout | Per-group scales and zeros, per-block headers | Effective bits per weight; how many separate tensors a loader must read |
None of these choices is better in general. They follow from the kernel the format was designed for, which is why formats differ and why converting between them is a real job.
Real layouts compared
Four widely used formats show the range. The figure shows one 32-bit word, or one byte, from each.
GPTQ (AutoGPTQ's layout) stores qweight as int32 with shape [in_features * bits / 32, out_features]: packed along the input dimension, eight 4-bit rows per word, row 8r + j at bits 4j. Alongside it sit scales (fp16, [groups, out_features]), qzeros (int32, packed along the output dimension) and g_idx, which maps each input row to its group. AutoGPTQ's pack step subtracts one from the zero points before storing them; this is the difference GPTQModel's gptq_v2 format setting addresses, and reading one convention as the other shifts every weight by one quantization step.
AWQ (the AutoAWQ GEMM layout) stores qweight as int32 with shape [in_features, out_features / 8]: packed along the output dimension, with logical columns placed in slots in the order 0, 2, 4, 6, 1, 3, 5, 7. Readers such as vLLM's AWQ Triton kernel undo it with the reverse order 0, 4, 1, 5, 2, 6, 3, 7, which means shifts of 0, 16, 4, 20, 8, 24, 12 and 28 bits. The interleave suits the dequantization bit tricks in the kernel the format was built for; to any other reader it is simply a permutation to undo.
GGUF Q4_0 uses blocks of 32 weights: one fp16 scale d and 16 bytes of codes. Byte j holds element j in its low nibble and element j + 16 in its high nibble, and a weight is d * (q - 8). The block format carries its own metadata, so no separate scale tensor is needed; GGUF quantization covers the k-quant families built on the same idea.
NF4 in bitsandbytes stores two 4-bit codes per byte, but the codes are indices into a fixed 16-entry table of non-uniform levels, scaled by an absmax per block of weights, typically 64. Because the levels are not evenly spaced, NF4 cannot be converted to an integer format such as GPTQ without requantizing.
Vectorised packers for the two int32 layouts
AWQ_ORDER = [0, 2, 4, 6, 1, 3, 5, 7]
SHIFTS = np.arange(8, dtype=np.uint32) * 4
def pack_gptq_4bit(q): # q: [K, N] codes 0..15, K % 8 == 0
K, N = q.shape
q = q.astype(np.uint32).reshape(K // 8, 8, N)
return np.bitwise_or.reduce(q << SHIFTS[None, :, None], axis=1).view(np.int32)
def pack_awq_4bit(q): # q: [K, N] codes 0..15, N % 8 == 0
K, N = q.shape
q = q.astype(np.uint32).reshape(K, N // 8, 8)[:, :, AWQ_ORDER]
return np.bitwise_or.reduce(q << SHIFTS[None, None, :], axis=2).view(np.int32)
def pack_q4_0_codes(q): # q: 32 codes for one block
return (q[:16] | (q[16:] << 4)).astype(np.uint8)Note the final .view(np.int32): checkpoints store signed int32, so the top slot's high bit becomes the sign bit. Unpacking must therefore shift as unsigned (or mask after an arithmetic shift), or values in the top slot pick up sign-extension bits and come back wrong.
Odd widths: 3-bit, 2-bit and 6-bit
Two-bit codes divide 32 evenly, sixteen per word, and pack exactly like 4-bit. Three bits do not: 32 values occupy 96 bits, exactly three int32 words, and with sequential packing the value at index 10 spans bits 30 to 32, half in the first word and half in the second; the value at index 21 spans the second and third. The reference packer above handles this with its straddle branch, and kernels handle it by loading words in groups of three and stitching the two straddling values.
The alternative is bit-plane splitting: store the low bits of every value in one array and the high bits in another, so each array has a width that divides evenly. GGUF's Q6_K does this, keeping the low 4 bits and the high 2 bits of its 6-bit codes in separate arrays. Splitting costs one more load and a shift-and-OR per value but keeps every access aligned, which is usually faster than straddling on GPUs.
Effective bits per weight, worked through
The headline bit width is never the real cost. Take one 4096 by 4096 projection, 16,777,216 weights, which is 33,554,432 bytes in fp16.
| Scheme | Codes | Metadata | Total bytes | Bits per weight |
|---|---|---|---|---|
| 4-bit, group 128, fp16 scale, 4-bit zero | 8,388,608 | 131,072 groups x 2.5 bytes = 327,680 | 8,716,288 | 4.16 |
| 4-bit, group 128, fp16 scale only | 8,388,608 | 131,072 x 2 = 262,144 | 8,650,752 | 4.125 |
| GGUF Q4_0, blocks of 32 | 8,388,608 | 524,288 blocks x 2 = 1,048,576 | 9,437,184 | 4.5 |
| 4-bit, group 32, fp16 scale and zero | 8,388,608 | 524,288 x 4 = 2,097,152 | 10,485,760 | 5.0 |
Smaller groups track the weights more closely and cost more metadata; the accuracy side of that trade is covered in quantization granularity. When you compare formats or estimate whether a model fits in memory, use the effective figure: at the scale of a 70-billion-parameter model the gap between 4.125 and 5.0 bits per weight is about 7.7 GB.
Repacking between formats
Serving stacks frequently repack at load time: a GPTQ checkpoint is converted into a kernel-specific layout such as Marlin's, or a model is moved between toolchains. The only safe method goes through the logical matrix:
- Unpack codes to a plain
[K, N]uint8 matrix using the source format's axis, order and sign handling. - Unpack zero points and normalise conventions, for example adding back the one AutoGPTQ subtracted, so the logical rule is
w = scale * (q - zero). - Resolve
g_idx. If it is not monotonic, as with act-order checkpoints, groups are scattered across K. A target layout withoutg_idxneeds the K rows permuted so groups become contiguous, and the same permutation applied to the layer's input activations; otherwise the conversion is wrong even though every value is intact. - Check that dequantized weights from the source and target layouts are bit-identical, then pack with the target's axis and order and write its metadata tensors.
Never repack by reinterpreting words directly; it seems to work on symmetric test cases and fails on real checkpoints.
Endianness, views and padding
Every platform in practice is little-endian, so slot 0 of an int32 is in its first byte. That matters when a tool views a packed int32 tensor as bytes, or bytes as int32: the reinterpretation is free only if both sides agree on byte order and alignment. Safetensors stores raw little-endian bytes with a dtype label, so a byte-level copy is safe; a dtype change during save, such as converting int32 to int64, is not.
Dimensions must be multiples of the pack factor and the group size. Real models occasionally have odd sizes, such as vocabulary projections. The usual fix is to pad K or N with zero weights up to the next multiple and record the logical size, then slice outputs after the matmul. Forgetting the slice adds garbage columns; forgetting the padding corrupts the last word.
Verification
Packing bugs produce plausible numbers, so test with inputs where wrong numbers are obvious.
def test_layout(pack, unpack, K=256, N=256):
probe = ((np.arange(K)[:, None] * 7 + np.arange(N)[None, :] * 3) % 16).astype(np.uint8)
assert np.array_equal(unpack(pack(probe)), probe) # round trip
rng = np.random.default_rng(0)
rand = rng.integers(0, 16, size=(K, N), dtype=np.uint8)
assert np.array_equal(unpack(pack(rand)), rand)
assert np.array_equal(unpack(pack(np.full((K, N), 15, np.uint8))), np.full((K, N), 15)) # sign bitThe patterned probe makes wrong axes and orders visible as a scrambled pattern; the all-15 case catches sign-extension bugs in the top slot. Then test end to end: dequantize a real layer with your reader and with the reference library, require exact equality, and compare a matmul output against fp16 on real activations.
Failure modes
- Wrong pack axis. Reading GPTQ as if packed along N, or AWQ as if along K, transposes groups of values. Output is fluent-looking nonsense or a perplexity jump.
- Wrong slot order. Ignoring AWQ's interleave swaps columns within each group of eight.
- Zero-point off by one from mixing GPTQ conventions; a small, consistent accuracy loss that is easy to blame on quantization itself.
- Sign extension on the top slot when unpacking signed int32 with an arithmetic shift.
- Dropped act-order permutation when converting to a layout without
g_idx. - Missing padding or slicing for dimensions that are not multiples of the group or pack size.
Operational guidance and trade-offs
Treat a packed format as an interface with a written specification: axis, order, offset convention, metadata shapes and padding rules. Keep a reference unpacker in plain Python next to every fast one and run the probes above in CI. Record the format and its version in model metadata, since tensor shapes alone do not distinguish conventions. Repack once at build time when you can, rather than on every server start, and keep the original checkpoint so a bug in the repacker can be fixed and rerun.
The design trade-off is between simple sequential layouts, which are easy to read and convert, and kernel-shaped layouts, which are faster but tie the file to one kernel family. Block formats that carry their own metadata are the most portable at the cost of slightly more bits per weight.
What to do next
- Write a one-paragraph layout spec for each format you serve: container, axis, slot order, offset, metadata shapes and padding.
- Implement the reference
pack_bitsandunpack_bitsabove and use them as the oracle for any fast packer. - Run the patterned, random and all-15 probes against every reader and converter you own.
- Compute effective bits per weight, including scales and zeros, for each candidate format before choosing one for a memory budget.
- For any repacking path, convert through the logical matrix, normalise zero-point conventions and handle act-order permutations explicitly.
- Dequantize a real layer with your code and with the reference library and require bit-identical results before shipping.