GGUF is the file format that carries a model into llama.cpp and the tools built on it. One file holds everything the runtime needs: the weights, usually quantized, the architecture hyperparameters, the tokenizer vocabulary and merges, and the chat template. You can copy that file to a laptop, a phone or a server and run it without the original Python repository, config files or tokenizer folder.

This page is about the file itself: what the bytes are, why the layout looks the way it does, how quantized weights are packed, how a file is produced, and what goes wrong. How llama.cpp schedules work, offloads layers and manages memory at run time is covered separately in running GGUF models. By the end you will be able to parse a GGUF header with a few dozen lines of Python, predict a file's size from its parameter count and quantization type, and debug the common conversion failures.

Advertisement

Why a single, self-describing file

The llama.cpp project originally stored models in GGML files, then a variant called GGJT that added memory-mapping-friendly alignment. Both had a fixed list of hyperparameters in a fixed order, so every new architecture or tokenizer feature broke old loaders or needed a new format revision. GGUF, introduced in August 2023, replaced the fixed fields with typed key-value metadata. A loader looks up the keys it understands and ignores the rest, and new architectures add keys under their own prefix instead of changing the layout.

Compare a typical Hugging Face checkpoint: weights in safetensors shards, hyperparameters in config.json, tokenizer files and a chat template in tokenizer_config.json. That is fine for a Python training stack and awkward for a C++ runtime on a phone. GGUF bundles the same information into one binary that a C program can read with simple pointer arithmetic, and it allows weights to be stored in ggml's quantized block formats, which safetensors does not define.

The layout, byte by byte

A GGUF file has four parts in order. The header is four magic bytes spelling GGUF, a uint32 version (3 is current; version 3 added big-endian support), then two uint64 counts: the number of tensors and the number of metadata entries. Version 1 used 32-bit counts, which is one reason old readers fail on new files.

Next come the metadata entries. Each is a key, a uint32 value type and the value. Strings are a uint64 byte length followed by UTF-8 bytes with no terminator. Arrays are an element type, a uint64 count and the elements, so the whole tokenizer vocabulary is one array of strings. Then the tensor info table: for each tensor, its name, the number of dimensions (at most four), the size of each dimension, the ggml type, and a uint64 offset. Finally, after padding to the alignment, the tensor data.

One GGUF file, front to backHeadermagic GGUF, version uint32, tensor_count uint64, metadata_kv_count uint64Metadata key-value pairskey string, value type uint32, value: general.*, <arch>.*, tokenizer.*Tensor info tableper tensor: name, n_dims, dims[], ggml type, offset into data sectionPadding to general.alignment (default 32 bytes)Tensor datablocks of quantized or float weights, each tensor starting at an aligned offsettoken_embd.weightblk.0.attn_q.weight ... blk.N.ffn_down.weightoutput.weightEverything above the data section is small and parsed once; the data section is memory-mapped and read in place.Offsets are relative to the start of the data section, so a loader can map it without copying.
The four regions of a GGUF file. Metadata and tensor descriptions come first and are small; the large data section starts on an aligned boundary and every tensor offset is relative to it.

The offset design is what makes loading fast. Every tensor starts at a multiple of general.alignment, 32 bytes unless the file says otherwise, measured from the start of the data section. A runtime can memory-map the file and hand the kernel-provided pages straight to compute kernels, so a multi-gigabyte model opens almost instantly, pages in on demand, and is shared between processes that map the same file.

Advertisement

Metadata: the keys that matter

Value types are numbered 0 to 12: unsigned and signed 8, 16, 32 and 64-bit integers, 32 and 64-bit floats, bool, string and array. Keys are dotted names in three main families.

KeyMeaningWhy you care
general.architectureArchitecture name, for example llamaSelects the loader and the prefix for every other model key
general.quantization_versionVersion of the quantized block formatsRequired; mismatches mean the runtime cannot decode the blocks
general.file_typeOverall quantization type of the fileWhat tools display as Q4_K_M and similar
<arch>.context_lengthTraining context lengthUpper bound for the context you should request
<arch>.block_countNumber of transformer layersUsed for layer offload decisions
<arch>.attention.head_count_kvKey-value headsSets KV cache size per token
tokenizer.ggml.modelTokenizer familyWrong value means garbage tokenization
tokenizer.ggml.tokensVocabulary arrayMust match the embedding rows
tokenizer.chat_templateJinja chat templateMissing or wrong template is the top cause of bad chat output

Tensor names follow a convention, such as token_embd.weight, blk.0.attn_q.weight and output.weight, which lets a runtime find tensors without architecture-specific mapping tables. Optional general keys record name, author, licence, base model and source URL; filling them in is cheap and makes model provenance traceable.

Reading a GGUF file yourself

The format is simple enough to parse with the standard library. This reader handles every value type, stops before the data section and prints what a runtime would see. It assumes little-endian files, which is what the conversion tools produce by default.

import struct

SCALAR = {0: "<B", 1: "<b", 2: "<H", 3: "<h", 4: "<I", 5: "<i", 6: "<f",
          7: "<?", 10: "<Q", 11: "<q", 12: "<d"}        # 8 = string, 9 = array
GGML_TYPE = {0: "F32", 1: "F16", 2: "Q4_0", 8: "Q8_0", 10: "Q2_K", 11: "Q3_K",
             12: "Q4_K", 13: "Q5_K", 14: "Q6_K", 30: "BF16"}   # partial list

class GGUFFile:
    def __init__(self, path):
        with open(path, "rb") as self.f:
            if self.f.read(4) != b"GGUF":
                raise ValueError("not a GGUF file")
            self.version = self.u32()                 # 3 today; assumes little-endian
            n_tensors, n_kv = self.u64(), self.u64()
            self.meta = {}
            for _ in range(n_kv):
                key = self.string()
                self.meta[key] = self.value(self.u32())
            self.tensors = []
            for _ in range(n_tensors):
                name = self.string()
                dims = [self.u64() for _ in range(self.u32())]
                ttype, offset = self.u32(), self.u64()
                self.tensors.append((name, dims, GGML_TYPE.get(ttype, ttype), offset))
            align = self.meta.get("general.alignment", 32)
            self.data_start = -(-self.f.tell() // align) * align   # round up

    def read(self, fmt):
        return struct.unpack(fmt, self.f.read(struct.calcsize(fmt)))[0]
    def u32(self): return self.read("<I")
    def u64(self): return self.read("<Q")
    def string(self):
        return self.f.read(self.u64()).decode("utf-8", errors="replace")
    def value(self, vtype):
        if vtype == 8:
            return self.string()
        if vtype == 9:
            item_type, count = self.u32(), self.u64()
            return [self.value(item_type) for _ in range(count)]
        return self.read(SCALAR[vtype])

g = GGUFFile("model-Q4_K_M.gguf")
arch = g.meta["general.architecture"]
print(arch, g.meta.get(f"{arch}.context_length"), g.meta.get(f"{arch}.block_count"))
print("vocab:", len(g.meta.get("tokenizer.ggml.tokens", [])))
print("has chat template:", "tokenizer.chat_template" in g.meta)
for name, dims, ttype, off in g.tensors[:5]:
    print(f"{name:32} {str(dims):20} {ttype:6} @ {g.data_start + off}")

Parsing a file with a 150,000-token vocabulary takes a moment because the token array is read element by element; the weights are never touched. For real work use the gguf package from the llama.cpp repository, which provides a reader and writer and the dump, set-metadata and new-metadata scripts. Writing your own reader once is still the fastest way to understand why a file fails to load.

Quantized tensor data

Each tensor's type decides how its bytes are packed. Floating-point types are plain arrays. Quantized types store weights in blocks, each with its own scale, so that a few outliers do not ruin the precision of a whole tensor. The legacy types use blocks of 32 weights. Q4_0 stores one 16-bit scale and 32 four-bit values in 18 bytes, which is 4.5 bits per weight. Q8_0 stores one scale and 32 bytes, 8.5 bits per weight.

The k-quants use super-blocks of 256 weights split into sub-blocks, with the sub-block scales themselves quantized. Block sizes follow from the struct definitions in ggml's source:

TypeWeights per blockBytes per blockBits per weight
Q8_032348.5
Q4_032184.5
Q6_K2562106.5625
Q5_K2561765.5
Q4_K2561444.5
Q3_K2561103.4375
Q2_K256842.625

Names such as Q4_K_M are recipes, not tensor types. They assign different types to different tensors, for example keeping some attention and feed-forward tensors at a higher precision than the rest. That is why the file's average bits per weight is usually a little above the nominal type's figure. The i-quants, named IQ, push below three bits using lookup grids and depend strongly on an importance matrix. The quantization page covers how to choose between these for accuracy.

Worked example: predicting file size and memory

Take an 8-billion-parameter model quantized with a mix that averages about 4.85 bits per weight. Weight bytes are 8.0e9 times 4.85 divided by 8, roughly 4.85 GB, plus a few megabytes of metadata, mostly the vocabulary. The same model in F16 is 16 GB, and Q8_0 is about 8.5 GB. That arithmetic predicts real files to within a few percent, so use it before downloading anything.

Memory at run time is the mapped weights plus the KV cache plus compute buffers. With 32 layers, 8 key-value heads of dimension 128 and an F16 cache, each token costs 2 times 32 times 8 times 128 times 2 bytes, 128 KB, so an 8,192-token context needs about 1 GB of cache on top of the weights. The file's metadata gives you every number in that formula, which is a good use for the parser above; the KV cache page explains the cache side in depth.

Building a GGUF: convert, quantize, split

Producing a file is a short pipeline using scripts and binaries from the llama.cpp repository. Convert once to a high-precision GGUF, then derive every quantized variant from that, never from another quantized file.

# 1. Hugging Face checkpoint (safetensors + config + tokenizer) -> one f16 GGUF
python convert_hf_to_gguf.py ./Qwen-small-instruct --outtype f16 --outfile model-f16.gguf

# 2. Optional: importance matrix from representative text, used to protect the
#    weights that matter most when quantizing to low bit widths
llama-imatrix -m model-f16.gguf -f calibration.txt -o imatrix.gguf

# 3. Quantize. The last argument picks the per-tensor type mix.
llama-quantize --imatrix imatrix.gguf model-f16.gguf model-Q4_K_M.gguf Q4_K_M

# 4. Optional: split for hosts with per-file size limits, and merge back later
llama-gguf-split --split --split-max-size 4G model-Q4_K_M.gguf model-Q4_K_M
llama-gguf-split --merge model-Q4_K_M-00001-of-00002.gguf model-Q4_K_M.gguf

# 5. Inspect what you produced (pip install gguf)
gguf-dump model-Q4_K_M.gguf | head -60

The calibration text for the importance matrix should look like the model's real inputs: chat transcripts for a chat model, code for a code model. The split tool names shards with five-digit numbers of the form 00001-of-00003, and loaders that support split files open the first shard and find the rest, so keep all shards in one directory.

Failure modes

  • Unknown architecture. An older llama.cpp build cannot load a file whose general.architecture it does not know. Upgrade the runtime rather than editing the key.
  • Unrecognised pre-tokenizer during conversion. The converter identifies how a BPE tokenizer splits text before merging. A new tokenizer it does not recognise stops conversion; forcing it through produces a file that tokenizes subtly wrong.
  • Wrong or missing chat template. The model loads and answers, but badly, because its turns are formatted differently from training. Check tokenizer.chat_template against the original repository; chat templates explains what to look for.
  • Re-quantizing a quantized file. Quantizing from Q8_0 to Q4_K compounds error. The quantize tool refuses by default for this reason; go back to the F16 or BF16 file.
  • Truncated download or missing shard. Shows up as an offset beyond the end of the file. Compare sizes and checksums with the source.
  • Context set beyond training length. The file states its training context; asking for more without a proper RoPE scaling setup degrades quality.

GGUF compared with other formats

FormatBest forLimits
GGUFllama.cpp and ggml-based runtimes on CPU, Apple silicon and consumer GPUs; single-file distributionFew frameworks train in it; tied to ggml's quant types
safetensorsTraining and serving in PyTorch stacksWeights only; config and tokenizer live elsewhere
ONNXCross-vendor graph runtimes and mobile acceleratorsExports large language models with more effort
MLX weightsApple's MLX frameworkApple silicon only

What to do next

  1. Run the parser above on a GGUF you already use and check architecture, context length, KV heads, vocabulary size and chat template.
  2. Predict the file size from parameter count and type, then compare with the real file to calibrate your arithmetic.
  3. Convert one small Hugging Face model to F16 GGUF, quantize it to Q8_0 and Q4_K_M, and compare outputs on a fixed prompt set.
  4. Build an importance matrix from in-domain text and repeat the comparison at a lower bit width.
  5. Fill in provenance metadata: name, licence, base model and source URL.
  6. Keep the F16 GGUF as the master copy and derive every quantized file from it.
  7. Read the runtime page next to learn how these bytes become tokens per second.
Key takeaway: GGUF is a simple, well-designed container: a small header, typed metadata that describes the architecture, tokenizer and chat template, a table of tensor descriptions, and an aligned data section that can be memory-mapped and used in place. Quantized tensors are blocks with their own scales, so file size follows directly from parameter count and bits per weight. Convert once to high precision, quantize from that master, check the metadata, especially the chat template, and most GGUF problems never reach users.