GGUF is the file format llama.cpp and the tools built on it use to ship a model as a single file: weights, architecture hyperparameters, tokenizer and chat template together. Most people only ever pass a GGUF path to a runtime. Sooner or later something goes wrong, a model loads but answers in gibberish, a conversion produces a file one runtime accepts and another rejects, or a download is silently truncated, and then it pays to know exactly what is in the file.
This article treats GGUF as a container to read and check, not a quantization method. The quantization ideas behind the block types are covered in the GGUF quantization article and the bit-level k-quant layouts in the GGUF deep dive. Here you will write a reader, a validator and a dequantizer, and learn which metadata fields cause which failures. Facts below follow the ggml repository's GGUF specification and source as of October 2026.
The layout in one picture
A GGUF file has two halves. The first half is small and parsed: a fixed header, a list of typed key-value pairs, and one info record per tensor. The second half is the tensor data itself, padded so it begins at an aligned position, with each tensor also starting at an aligned offset. That split is the point of the format. A runtime parses a few megabytes of metadata, then memory-maps the gigabytes of weights and uses them in place without copying or deserializing them, so a model loads as fast as the operating system can page it in.
The header holds the four magic bytes GGUF, a 32-bit version, and two 64-bit counts: tensors and metadata pairs. Version 3 is current. The specification's history is short and worth knowing: version 2 widened most counts and lengths from 32-bit to 64-bit integers, and version 3 added big-endian support. Files are little-endian by default; a big-endian file has every value, including tensor data, in big-endian order, and there is no flag in the header saying so, which is why readers sometimes detect it by noticing an absurd version number.
Primitive encoding: strings, typed values and arrays
Everything in the parsed half is built from a few primitives. A string is a 64-bit byte length followed by UTF-8 bytes, with no terminator. A metadata pair is a string key, a 32-bit value type and the value. Thirteen value types exist:
| Id | Type | Id | Type |
|---|---|---|---|
| 0 | UINT8 | 7 | BOOL |
| 1 | INT8 | 8 | STRING |
| 2 | UINT16 | 9 | ARRAY |
| 3 | INT16 | 10 | UINT64 |
| 4 | UINT32 | 11 | INT64 |
| 5 | INT32 | 12 | FLOAT64 |
| 6 | FLOAT32 |
An array is an element type, a 64-bit element count (elements, not bytes) and the elements, and arrays may nest. Tokenizer vocabularies are arrays of strings with one entry per token, so in a model with a large vocabulary the metadata section alone can run to several megabytes. Keys are dotted names, with general. for model-wide facts, an architecture prefix such as llama. for hyperparameters, and tokenizer. for the tokenizer. The specification requires general.architecture and general.quantization_version, and general.alignment is required whenever the alignment differs from the default.
A complete reader in fifty lines
The fastest way to understand the format is to parse it. This reader handles versions 2 and 3, little-endian, and returns the metadata, tensor infos and the absolute position of the data section. It reads no tensor data, so it runs in milliseconds even on a 40 GB file.
import struct
SCALARS = {0: "<B", 1: "<b", 2: "<H", 3: "<h", 4: "<I", 5: "<i", 6: "<f",
7: "<?", 10: "<Q", 11: "<q", 12: "<d"} # 8 = string, 9 = array
class R:
def __init__(self, f): self.f = f
def unpack(self, fmt):
n = struct.calcsize(fmt)
return struct.unpack(fmt, self.f.read(n))[0]
def string(self):
n = self.unpack("<Q")
return self.f.read(n).decode("utf-8")
def value(self, vtype):
if vtype == 8: return self.string()
if vtype == 9:
etype, count = self.unpack("<I"), self.unpack("<Q")
return [self.value(etype) for _ in range(count)]
return self.unpack(SCALARS[vtype])
def read_gguf(path):
with open(path, "rb") as f:
r = R(f)
if f.read(4) != b"GGUF":
raise ValueError("not a GGUF file")
version = r.unpack("<I")
if version not in (2, 3):
raise ValueError(f"unsupported version {version}")
n_tensors, n_kv = r.unpack("<Q"), r.unpack("<Q")
meta = {}
for _ in range(n_kv):
key = r.string()
meta[key] = r.value(r.unpack("<I"))
tensors = []
for _ in range(n_tensors):
name, nd = r.string(), r.unpack("<I")
dims = [r.unpack("<Q") for _ in range(nd)]
tensors.append(dict(name=name, dims=dims,
type=r.unpack("<I"), offset=r.unpack("<Q")))
align = meta.get("general.alignment", 32)
data_start = (f.tell() + align - 1) // align * align
return version, meta, tensors, data_startTwo lines carry most of the subtlety. The data section starts at the first multiple of the alignment after the last tensor info, so the reader rounds up its position. And the tensor offset is relative to that start, not to the beginning of the file; the specification chose this so writers can lay out tensors before knowing how large the metadata will be. Forgetting to add data_start is the most common bug in hand-written readers. For production, use the gguf Python package from the llama.cpp repository; runtime loading is covered in the GGUF runtime article.
Tensor infos: names, dimension order, types and sizes
Each tensor info holds a name of at most 64 bytes, a dimension count (currently at most four), the dimensions, a ggml type id and the offset. Names follow conventions shared by the converters and the runtime, for example token_embd.weight, blk.0.attn_q.weight, blk.0.ffn_down.weight and output_norm.weight. If an architecture's loader looks for a name that is not there, loading fails with a missing-tensor error, which is how an unsupported architecture or a half-finished conversion usually shows up.
Dimension order surprises everyone once. ggml lists the fastest-varying dimension first, so a weight PyTorch describes as [4096, 11008] (rows, columns) appears in GGUF as [11008, 4096]. The first dimension is the row length in memory, and it is the one that must be a multiple of the block size, because quantization blocks run along rows.
The type id says how the bytes encode values. Float types store one value per element; quantized types store fixed-size blocks, so tensor bytes equal element count divided by block size, times bytes per block. The sizes below come from the static size checks in ggml's source:
| Type (id) | Elements per block | Bytes per block | Bits per weight |
|---|---|---|---|
| F32 (0) | 1 | 4 | 32 |
| F16 (1), BF16 (30) | 1 | 2 | 16 |
| Q4_0 (2) | 32 | 18 (fp16 scale + 16 bytes) | 4.5 |
| Q8_0 (8) | 32 | 34 (fp16 scale + 32 bytes) | 8.5 |
| Q2_K (10) | 256 | 84 | 2.625 |
| Q3_K (11) | 256 | 110 | 3.4375 |
| Q4_K (12) | 256 | 144 | 4.5 |
| Q5_K (13) | 256 | 176 | 5.5 |
| Q6_K (14) | 256 | 210 | 6.5625 |
The ggml type list is longer than this table, with i-quants, integer types and newer formats, and ids are not contiguous. A file mixes types: a Q4_K_M model (file type 15 in general.file_type) is mostly Q4_K but keeps some sensitive tensors at higher precision, and norms usually stay F32. The file-type field is a label; the per-tensor types are the truth.
A validator that catches real breakage
With the reader and the block table you can check a file without loading it into a runtime: every offset aligned, no tensors overlapping, rows divisible by the block size, and the file long enough to hold the last tensor.
import math, os
# ggml_type id -> (elements per block, bytes per block)
BLOCK = {0: (1, 4), 1: (1, 2), 30: (1, 2), 2: (32, 18), 8: (32, 34),
10: (256, 84), 11: (256, 110), 12: (256, 144), 13: (256, 176), 14: (256, 210)}
def validate(path):
version, meta, tensors, data_start = read_gguf(path)
align = meta.get("general.alignment", 32)
problems, end = [], data_start
for t in sorted(tensors, key=lambda t: t["offset"]):
if t["offset"] % align:
problems.append(f"{t['name']}: offset not aligned")
if t["type"] not in BLOCK:
problems.append(f"{t['name']}: type {t['type']} not in table"); continue
per, nbytes = BLOCK[t["type"]]
n = math.prod(t["dims"])
if t["dims"][0] % per:
problems.append(f"{t['name']}: row of {t['dims'][0]} not a multiple of {per}")
start = data_start + t["offset"]
if start < end:
problems.append(f"{t['name']}: overlaps previous tensor")
end = start + n // per * nbytes
if end > os.path.getsize(path):
problems.append(f"truncated: need {end} bytes")
for key in ("general.architecture", "tokenizer.ggml.model"):
if key not in meta:
problems.append(f"missing {key}")
return problemsThe truncation check alone pays for the script: an interrupted download parses perfectly, because the header and metadata are at the front, and only fails when the runtime touches the last layers, sometimes mid-generation. Run the validator in your model registry's ingestion step and refuse files that fail. Unknown types are reported rather than guessed, so extend the table from ggml's source when you meet new ones instead of assuming a size.
Worked example: dequantizing a block by hand
To confirm you understand a quantized tensor, decode one block and compare with the original weights. Q8_0 and Q4_0 are the simplest. A Q8_0 block is a half-precision scale followed by 32 signed bytes, and each weight is byte times scale. A Q4_0 block is a half-precision scale followed by 16 bytes holding 32 four-bit values; each value is stored with an offset of 8, so weight equals (nibble minus 8) times scale. The order matters: in ggml's reference dequantizer, the low nibbles of the 16 bytes are elements 0 to 15 and the high nibbles are elements 16 to 31, not interleaved.
import numpy as np
def dequant_q4_0(block: bytes) -> np.ndarray:
d = np.frombuffer(block[:2], dtype="<f2")[0].astype(np.float32)
qs = np.frombuffer(block[2:18], dtype=np.uint8)
lo = (qs & 0x0F).astype(np.int8) - 8 # elements 0..15
hi = (qs >> 4).astype(np.int8) - 8 # elements 16..31
return np.concatenate([lo, hi]).astype(np.float32) * d
def dequant_q8_0(block: bytes) -> np.ndarray:
d = np.frombuffer(block[:2], dtype="<f2")[0].astype(np.float32)
return np.frombuffer(block[2:34], dtype=np.int8).astype(np.float32) * dTake a Q4_0 block whose scale is 0.01 and whose first byte is 0x4B. The low nibble is 0xB, 11, so element 0 is (11 - 8) times 0.01, which is 0.03; the high nibble is 4, so element 16 is (4 - 8) times 0.01, which is -0.04. Seek to the tensor's absolute position, read one block, decode it and compare against the same 32 weights from the original checkpoint, remembering the dimension reversal. Errors of about half a quantization step are expected; errors of whole steps or wrong signs mean you have the layout wrong, or the converter does. The k-quants add per-sub-block scales and minimums packed into 6-bit fields; their unpacking is explained in the deep dive linked above.
Metadata that breaks inference when it is wrong
Weights are only half of a model. The runtime builds the computation from metadata, and wrong values produce a model that loads cleanly and then misbehaves.
| Key (llama example) | What it controls | Symptom when wrong |
|---|---|---|
llama.block_count | Number of transformer layers | Missing-tensor error or ignored layers |
llama.attention.head_count_kv | Grouped-query attention groups | Shape errors or garbage output |
llama.rope.freq_base | Rotary position frequency | Fine short prompts, degrading long ones |
llama.context_length | Trained context size | Runtimes default to an unsupported length |
tokenizer.ggml.model and tokens | Tokenizer type and vocabulary | Gibberish or broken whitespace |
tokenizer.ggml.eos_token_id | When generation stops | Model never stops, or stops at once |
tokenizer.chat_template | Prompt format for chat | Model ignores roles, leaks markup, rambles |
The chat template is a Jinja template stored as a string. It is the most common reason a model seems worse in GGUF than upstream, so diff it against the upstream tokenizer configuration. Treat it as untrusted input too: it should be rendered in a sandboxed template engine, not every binding has always done so, and a file from an unknown publisher deserves no more trust than other downloaded code.
Producing and changing files
The usual path is a Hugging Face checkpoint converted with llama.cpp's convert_hf_to_gguf.py to an F16 or BF16 GGUF, then llama-quantize to the target type, optionally with an importance matrix from llama-imatrix to weight the error toward activations that matter. Keep the full-precision GGUF: re-quantizing from it is reproducible, while quantizing an already quantized file compounds error. To fix metadata such as a chat template or end-of-sequence id, rewrite the file with the gguf Python package rather than patching bytes, because changing a string length shifts everything after it and the data section must stay aligned. Very large models can be sharded into several numbered files with llama.cpp's split tool, and loaders open them from the first shard; validate each shard. Measure quality after every step with the methods in the quantization evaluation guide.
Failure modes and trade-offs
| Failure | Cause | Detection |
|---|---|---|
| Load fails late or mid-run | Truncated download | File size against last tensor end |
| Unknown tensor type | Runtime older than the file | Type ids against the runtime's table |
| Unknown architecture | Converter newer than runtime | general.architecture against supported list |
| Plausible but poor answers | Wrong chat template or rope settings | Diff metadata against upstream config |
| Wrong values after custom conversion | Dimension order or nibble order | Hand-dequantize one block and compare |
Against safetensors, GGUF trades generality for deployment convenience: one file carries tokenizer and template, quantized blocks are native, and the layout is built for memory mapping, but it is tied to ggml's types and architecture naming, so other frameworks need converters. For llama.cpp-family runtimes on CPUs, Apple silicon and consumer GPUs it is the natural choice; for training and server frameworks built on PyTorch, keep safetensors as the source of truth.
What to do next
- Run the reader on a model you use and print its metadata keys, tensor names, types and dimensions.
- Add the validator to your model download or registry step and reject truncated or misaligned files.
- Diff the chat template, end-of-sequence id and rope settings against the upstream model configuration.
- Hand-dequantize one Q8_0 or Q4_0 block and compare it with the original weights to prove your layout understanding.
- Keep the F16 or BF16 GGUF alongside quantized variants and record the llama.cpp version used for each conversion.
- Evaluate each quantized variant on your own tasks before deploying it.