GGUF is the file format llama.cpp and the tools built on it use to ship a model as a single file: weights, architecture hyperparameters, tokenizer and chat template together. Most people only ever pass a GGUF path to a runtime. Sooner or later something goes wrong, a model loads but answers in gibberish, a conversion produces a file one runtime accepts and another rejects, or a download is silently truncated, and then it pays to know exactly what is in the file.

This article treats GGUF as a container to read and check, not a quantization method. The quantization ideas behind the block types are covered in the GGUF quantization article and the bit-level k-quant layouts in the GGUF deep dive. Here you will write a reader, a validator and a dequantizer, and learn which metadata fields cause which failures. Facts below follow the ggml repository's GGUF specification and source as of October 2026.

Advertisement

The layout in one picture

A GGUF file: a self-describing header, then aligned tensor data you can mmapHeadermagic GGUF | version uint32 | tensor_count uint64 | metadata_kv_count uint64Metadata key-value pairsgeneral.architecture, llama.block_count, tokenizer.ggml.tokens, tokenizer.chat_template ...Tensor infosname | n_dims | dims (fastest first) | ggml_type | offset into data sectionPadding to general.alignment (default 32)Tensor datatoken_embd.weight | blk.0.attn_q.weight | ... each starting at an aligned offsetoffsetEverything before the data section is parsed; the data section is mapped and used in place.
Header, metadata, tensor infos, padding, data. Offsets in tensor infos are relative to the start of the data section.

A GGUF file has two halves. The first half is small and parsed: a fixed header, a list of typed key-value pairs, and one info record per tensor. The second half is the tensor data itself, padded so it begins at an aligned position, with each tensor also starting at an aligned offset. That split is the point of the format. A runtime parses a few megabytes of metadata, then memory-maps the gigabytes of weights and uses them in place without copying or deserializing them, so a model loads as fast as the operating system can page it in.

The header holds the four magic bytes GGUF, a 32-bit version, and two 64-bit counts: tensors and metadata pairs. Version 3 is current. The specification's history is short and worth knowing: version 2 widened most counts and lengths from 32-bit to 64-bit integers, and version 3 added big-endian support. Files are little-endian by default; a big-endian file has every value, including tensor data, in big-endian order, and there is no flag in the header saying so, which is why readers sometimes detect it by noticing an absurd version number.

Primitive encoding: strings, typed values and arrays

Everything in the parsed half is built from a few primitives. A string is a 64-bit byte length followed by UTF-8 bytes, with no terminator. A metadata pair is a string key, a 32-bit value type and the value. Thirteen value types exist:

IdTypeIdType
0UINT87BOOL
1INT88STRING
2UINT169ARRAY
3INT1610UINT64
4UINT3211INT64
5INT3212FLOAT64
6FLOAT32

An array is an element type, a 64-bit element count (elements, not bytes) and the elements, and arrays may nest. Tokenizer vocabularies are arrays of strings with one entry per token, so in a model with a large vocabulary the metadata section alone can run to several megabytes. Keys are dotted names, with general. for model-wide facts, an architecture prefix such as llama. for hyperparameters, and tokenizer. for the tokenizer. The specification requires general.architecture and general.quantization_version, and general.alignment is required whenever the alignment differs from the default.

Advertisement

A complete reader in fifty lines

The fastest way to understand the format is to parse it. This reader handles versions 2 and 3, little-endian, and returns the metadata, tensor infos and the absolute position of the data section. It reads no tensor data, so it runs in milliseconds even on a 40 GB file.

import struct

SCALARS = {0: "<B", 1: "<b", 2: "<H", 3: "<h", 4: "<I", 5: "<i", 6: "<f",
           7: "<?", 10: "<Q", 11: "<q", 12: "<d"}       # 8 = string, 9 = array

class R:
    def __init__(self, f): self.f = f
    def unpack(self, fmt):
        n = struct.calcsize(fmt)
        return struct.unpack(fmt, self.f.read(n))[0]
    def string(self):
        n = self.unpack("<Q")
        return self.f.read(n).decode("utf-8")
    def value(self, vtype):
        if vtype == 8: return self.string()
        if vtype == 9:
            etype, count = self.unpack("<I"), self.unpack("<Q")
            return [self.value(etype) for _ in range(count)]
        return self.unpack(SCALARS[vtype])

def read_gguf(path):
    with open(path, "rb") as f:
        r = R(f)
        if f.read(4) != b"GGUF":
            raise ValueError("not a GGUF file")
        version = r.unpack("<I")
        if version not in (2, 3):
            raise ValueError(f"unsupported version {version}")
        n_tensors, n_kv = r.unpack("<Q"), r.unpack("<Q")
        meta = {}
        for _ in range(n_kv):
            key = r.string()
            meta[key] = r.value(r.unpack("<I"))
        tensors = []
        for _ in range(n_tensors):
            name, nd = r.string(), r.unpack("<I")
            dims = [r.unpack("<Q") for _ in range(nd)]
            tensors.append(dict(name=name, dims=dims,
                                type=r.unpack("<I"), offset=r.unpack("<Q")))
        align = meta.get("general.alignment", 32)
        data_start = (f.tell() + align - 1) // align * align
    return version, meta, tensors, data_start

Two lines carry most of the subtlety. The data section starts at the first multiple of the alignment after the last tensor info, so the reader rounds up its position. And the tensor offset is relative to that start, not to the beginning of the file; the specification chose this so writers can lay out tensors before knowing how large the metadata will be. Forgetting to add data_start is the most common bug in hand-written readers. For production, use the gguf Python package from the llama.cpp repository; runtime loading is covered in the GGUF runtime article.

Tensor infos: names, dimension order, types and sizes

Each tensor info holds a name of at most 64 bytes, a dimension count (currently at most four), the dimensions, a ggml type id and the offset. Names follow conventions shared by the converters and the runtime, for example token_embd.weight, blk.0.attn_q.weight, blk.0.ffn_down.weight and output_norm.weight. If an architecture's loader looks for a name that is not there, loading fails with a missing-tensor error, which is how an unsupported architecture or a half-finished conversion usually shows up.

Dimension order surprises everyone once. ggml lists the fastest-varying dimension first, so a weight PyTorch describes as [4096, 11008] (rows, columns) appears in GGUF as [11008, 4096]. The first dimension is the row length in memory, and it is the one that must be a multiple of the block size, because quantization blocks run along rows.

The type id says how the bytes encode values. Float types store one value per element; quantized types store fixed-size blocks, so tensor bytes equal element count divided by block size, times bytes per block. The sizes below come from the static size checks in ggml's source:

Type (id)Elements per blockBytes per blockBits per weight
F32 (0)1432
F16 (1), BF16 (30)1216
Q4_0 (2)3218 (fp16 scale + 16 bytes)4.5
Q8_0 (8)3234 (fp16 scale + 32 bytes)8.5
Q2_K (10)256842.625
Q3_K (11)2561103.4375
Q4_K (12)2561444.5
Q5_K (13)2561765.5
Q6_K (14)2562106.5625

The ggml type list is longer than this table, with i-quants, integer types and newer formats, and ids are not contiguous. A file mixes types: a Q4_K_M model (file type 15 in general.file_type) is mostly Q4_K but keeps some sensitive tensors at higher precision, and norms usually stay F32. The file-type field is a label; the per-tensor types are the truth.

A validator that catches real breakage

With the reader and the block table you can check a file without loading it into a runtime: every offset aligned, no tensors overlapping, rows divisible by the block size, and the file long enough to hold the last tensor.

import math, os

# ggml_type id -> (elements per block, bytes per block)
BLOCK = {0: (1, 4), 1: (1, 2), 30: (1, 2), 2: (32, 18), 8: (32, 34),
         10: (256, 84), 11: (256, 110), 12: (256, 144), 13: (256, 176), 14: (256, 210)}

def validate(path):
    version, meta, tensors, data_start = read_gguf(path)
    align = meta.get("general.alignment", 32)
    problems, end = [], data_start
    for t in sorted(tensors, key=lambda t: t["offset"]):
        if t["offset"] % align:
            problems.append(f"{t['name']}: offset not aligned")
        if t["type"] not in BLOCK:
            problems.append(f"{t['name']}: type {t['type']} not in table"); continue
        per, nbytes = BLOCK[t["type"]]
        n = math.prod(t["dims"])
        if t["dims"][0] % per:
            problems.append(f"{t['name']}: row of {t['dims'][0]} not a multiple of {per}")
        start = data_start + t["offset"]
        if start < end:
            problems.append(f"{t['name']}: overlaps previous tensor")
        end = start + n // per * nbytes
    if end > os.path.getsize(path):
        problems.append(f"truncated: need {end} bytes")
    for key in ("general.architecture", "tokenizer.ggml.model"):
        if key not in meta:
            problems.append(f"missing {key}")
    return problems

The truncation check alone pays for the script: an interrupted download parses perfectly, because the header and metadata are at the front, and only fails when the runtime touches the last layers, sometimes mid-generation. Run the validator in your model registry's ingestion step and refuse files that fail. Unknown types are reported rather than guessed, so extend the table from ggml's source when you meet new ones instead of assuming a size.

Worked example: dequantizing a block by hand

To confirm you understand a quantized tensor, decode one block and compare with the original weights. Q8_0 and Q4_0 are the simplest. A Q8_0 block is a half-precision scale followed by 32 signed bytes, and each weight is byte times scale. A Q4_0 block is a half-precision scale followed by 16 bytes holding 32 four-bit values; each value is stored with an offset of 8, so weight equals (nibble minus 8) times scale. The order matters: in ggml's reference dequantizer, the low nibbles of the 16 bytes are elements 0 to 15 and the high nibbles are elements 16 to 31, not interleaved.

import numpy as np

def dequant_q4_0(block: bytes) -> np.ndarray:
    d = np.frombuffer(block[:2], dtype="<f2")[0].astype(np.float32)
    qs = np.frombuffer(block[2:18], dtype=np.uint8)
    lo = (qs & 0x0F).astype(np.int8) - 8        # elements 0..15
    hi = (qs >> 4).astype(np.int8) - 8          # elements 16..31
    return np.concatenate([lo, hi]).astype(np.float32) * d

def dequant_q8_0(block: bytes) -> np.ndarray:
    d = np.frombuffer(block[:2], dtype="<f2")[0].astype(np.float32)
    return np.frombuffer(block[2:34], dtype=np.int8).astype(np.float32) * d

Take a Q4_0 block whose scale is 0.01 and whose first byte is 0x4B. The low nibble is 0xB, 11, so element 0 is (11 - 8) times 0.01, which is 0.03; the high nibble is 4, so element 16 is (4 - 8) times 0.01, which is -0.04. Seek to the tensor's absolute position, read one block, decode it and compare against the same 32 weights from the original checkpoint, remembering the dimension reversal. Errors of about half a quantization step are expected; errors of whole steps or wrong signs mean you have the layout wrong, or the converter does. The k-quants add per-sub-block scales and minimums packed into 6-bit fields; their unpacking is explained in the deep dive linked above.

Metadata that breaks inference when it is wrong

Weights are only half of a model. The runtime builds the computation from metadata, and wrong values produce a model that loads cleanly and then misbehaves.

Key (llama example)What it controlsSymptom when wrong
llama.block_countNumber of transformer layersMissing-tensor error or ignored layers
llama.attention.head_count_kvGrouped-query attention groupsShape errors or garbage output
llama.rope.freq_baseRotary position frequencyFine short prompts, degrading long ones
llama.context_lengthTrained context sizeRuntimes default to an unsupported length
tokenizer.ggml.model and tokensTokenizer type and vocabularyGibberish or broken whitespace
tokenizer.ggml.eos_token_idWhen generation stopsModel never stops, or stops at once
tokenizer.chat_templatePrompt format for chatModel ignores roles, leaks markup, rambles

The chat template is a Jinja template stored as a string. It is the most common reason a model seems worse in GGUF than upstream, so diff it against the upstream tokenizer configuration. Treat it as untrusted input too: it should be rendered in a sandboxed template engine, not every binding has always done so, and a file from an unknown publisher deserves no more trust than other downloaded code.

Producing and changing files

The usual path is a Hugging Face checkpoint converted with llama.cpp's convert_hf_to_gguf.py to an F16 or BF16 GGUF, then llama-quantize to the target type, optionally with an importance matrix from llama-imatrix to weight the error toward activations that matter. Keep the full-precision GGUF: re-quantizing from it is reproducible, while quantizing an already quantized file compounds error. To fix metadata such as a chat template or end-of-sequence id, rewrite the file with the gguf Python package rather than patching bytes, because changing a string length shifts everything after it and the data section must stay aligned. Very large models can be sharded into several numbered files with llama.cpp's split tool, and loaders open them from the first shard; validate each shard. Measure quality after every step with the methods in the quantization evaluation guide.

Failure modes and trade-offs

FailureCauseDetection
Load fails late or mid-runTruncated downloadFile size against last tensor end
Unknown tensor typeRuntime older than the fileType ids against the runtime's table
Unknown architectureConverter newer than runtimegeneral.architecture against supported list
Plausible but poor answersWrong chat template or rope settingsDiff metadata against upstream config
Wrong values after custom conversionDimension order or nibble orderHand-dequantize one block and compare

Against safetensors, GGUF trades generality for deployment convenience: one file carries tokenizer and template, quantized blocks are native, and the layout is built for memory mapping, but it is tied to ggml's types and architecture naming, so other frameworks need converters. For llama.cpp-family runtimes on CPUs, Apple silicon and consumer GPUs it is the natural choice; for training and server frameworks built on PyTorch, keep safetensors as the source of truth.

What to do next

  1. Run the reader on a model you use and print its metadata keys, tensor names, types and dimensions.
  2. Add the validator to your model download or registry step and reject truncated or misaligned files.
  3. Diff the chat template, end-of-sequence id and rope settings against the upstream model configuration.
  4. Hand-dequantize one Q8_0 or Q4_0 block and compare it with the original weights to prove your layout understanding.
  5. Keep the F16 or BF16 GGUF alongside quantized variants and record the llama.cpp version used for each conversion.
  6. Evaluate each quantized variant on your own tasks before deploying it.
Key takeaway: A GGUF file is a typed header, metadata, tensor infos and an aligned data section that runtimes memory-map in place. Offsets are relative to the data section, dimensions are listed fastest first, and quantized tensor size follows from fixed block sizes, which is enough to write a reader and a validator that catch truncation and misalignment. Most quality problems with a well-formed file come from metadata such as the chat template, tokenizer ids or rope settings, so diff those against upstream every time.