The Hugging Face libraries are usually described one at a time: the Hub stores models, Transformers defines them, Datasets feeds them, PEFT and TRL train them, Accelerate spreads them over devices. On a GPU box what matters is the handoffs. A checkpoint is downloaded, memory-mapped, placed layer by layer on devices, run through attention kernels chosen at load time, trained through small adapters, and handed to a serving engine that is not a Hugging Face library at all. Most production surprises with this stack, such as a slow model, a run you cannot reproduce or an adapter that misbehaves in serving, happen at one of those handoffs.

This article follows one checkpoint along that path and says what each step does with bytes and GPU memory. Two recent changes shape it. Transformers v5, whose first release candidate appeared on 1 December 2025, made PyTorch its only backend, kept a single tokenizer backend instead of fast and slow variants, and reworked weight loading so quantization is handled as part of loading. And Text Generation Inference, Hugging Face's own server, has been in maintenance mode since December 2025; its documentation points to vLLM and SGLang, which can run models through the Transformers definitions.

The path in one picture

One checkpoint, followed from the Hub into GPU memory and back out to a serving engineHub repocommit sha, not mainLocal cacheHF_HOME, one per nodesnapshot_downloadsafetensors shardsJSON header + raw bytesfrom_pretrainedmeta device, device_mapmmapGPU 0layers 0..kGPU 1layers k+1..nCPU offloadcopied in every forwarddatasets (Arrow)tokenised once, mmapPEFT + Accelerateadapter weights onlyDataLoaderAdapter repobase sha recordedpushvLLM or SGLangbase + LoRA
Each arrow is a handoff where a revision, a dtype or a device can change without an error.

Pin the bytes before anything else

Everything downstream depends on knowing exactly which bytes you loaded. A Hub repository is a git repository, and main moves when the author pushes. Pin the full commit hash, download only the files the loader needs, and do it once per node into a shared cache rather than once per training process:

from huggingface_hub import snapshot_download

REPO = "your-org/your-14b-model"
SHA = "full-40-character-commit-hash"

path = snapshot_download(
    REPO, revision=SHA,
    allow_patterns=["*.safetensors", "*.json", "tokenizer*"],   # skip .bin duplicates and extras
)
# Training and serving processes then run with HF_HUB_OFFLINE=1 and load from `path`.

The cache layout, revision handling and offline mode are covered in the Hugging Face Hub in depth. The GPU-relevant point is simpler: eight ranks each downloading 28 GB at job start is a slow start and a rate-limit risk, and two nodes resolving main a minute apart can load different weights into the same job.

What is inside a safetensors file

A .safetensors file is deliberately simple. The first 8 bytes are a little-endian unsigned integer giving the header length. The header is JSON mapping each tensor name to its dtype, shape and a pair of byte offsets, plus an optional __metadata__ entry. After the header come the raw tensor bytes. A sharded checkpoint adds model.safetensors.index.json, which maps each tensor name to its shard file. You can read the header yourself:

import json, struct

def read_header(path):
    with open(path, "rb") as f:
        n = struct.unpack("<Q", f.read(8))[0]      # header length
        header = json.loads(f.read(n))
    meta = header.pop("__metadata__", {})
    return meta, header    # name -> {"dtype", "shape", "data_offsets": [start, end]}

meta, tensors = read_header("model-00001-of-00006.safetensors")
total = sum(t["data_offsets"][1] - t["data_offsets"][0] for t in tensors.values())
print(len(tensors), "tensors,", total / 2**30, "GiB in this shard")

This format has two consequences for loading. The file can be memory-mapped, so a loader reads only the tensors it needs, straight from the page cache, without parsing the rest; a process that needs only some of the layers touches only those bytes. And loading runs no code, unlike the older pickle-based .bin files, which can execute arbitrary Python when unpickled. Prefer safetensors files, and treat any repository that only ships pickles, or asks for trust_remote_code=True, as code you are about to run: pin its revision and read it.

How from_pretrained places weights on devices

When from_pretrained is given a device_map, it first builds the model on PyTorch's meta device, where tensors have shapes but no storage, so building a 14-billion-parameter skeleton allocates nothing. It then decides where each module goes and materialises each tensor directly on its target device from the memory-mapped shards. With device_map="auto", Accelerate fills GPUs in order up to the max_memory you allow, then CPU memory, then disk if you give it an offload folder.

import torch
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    path,
    dtype=torch.bfloat16,                      # older releases call this torch_dtype
    device_map="auto",
    max_memory={0: "12GiB", 1: "12GiB", "cpu": "48GiB"},
    attn_implementation="sdpa",
)
print(model.hf_device_map)                     # module name -> 0, 1, "cpu" or "disk"

Two properties of this placement are easy to miss. First, it is a sequential split, not parallelism: layers on GPU 0 run, then layers on GPU 1 run, so at batch size 1 only one GPU is busy at any moment. It lets a model fit; it does not make it faster. Second, modules placed on cpu are not computed on the CPU. Accelerate attaches hooks that copy their weights to the GPU for each forward pass, so every generated token pays for that copy over PCIe.

Worked example: a 14B model on two 16 GB GPUs

Suppose you want to run a 14B model in bfloat16 on a box with two 16 GB GPUs. The weights are about 14 x 10^9 x 2 bytes, roughly 28 GB or 26 GiB. You cap each GPU at 12 GiB to leave room for activations, the KV cache and the CUDA context, so 24 GiB of weights fit and about 2 GiB of layers land on cpu.

Loading succeeds and the first answer is correct, which is why this goes unnoticed. But each decoding step now copies about 2 GiB from host memory. At an effective 20 to 25 GB/s over PCIe 4.0 x16, that is on the order of 100 ms per token for the copy alone, likely more than the rest of the step. Printing hf_device_map and counting entries that are not GPU indices shows the problem in one line.

There are three honest fixes. Load the weights in 4-bit NF4 with bitsandbytes: about half a byte per parameter plus small per-block scales, roughly 8 GB, which fits on one GPU with room for the KV cache, at some cost in quality and per-token speed. Use a checkpoint already quantised with a method such as AWQ or GPTQ, which Transformers reads from the quantization config in config.json. Or move serving to an engine with real tensor parallelism and check the memory budget there. Raising max_memory until the CPU entries vanish is not a fix; it moves the failure to an out-of-memory error at the first long prompt.

from transformers import BitsAndBytesConfig

bnb = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,     # weights are dequantised to this for each matmul
    bnb_4bit_use_double_quant=True,
)
model = AutoModelForCausalLM.from_pretrained(path, quantization_config=bnb, device_map={"": 0})

Choosing attention kernels

The attn_implementation argument picks the attention code path when the model is built. "sdpa" uses PyTorch's scaled_dot_product_attention, which dispatches to a fused kernel when the dtype, head size and mask allow. "flash_attention_2" needs the separately compiled flash-attn package and half-precision inputs. "eager" is the plain implementation, slow but useful when you need attention weights or are debugging numerics.

Compiling flash-attn against the right CUDA and PyTorch versions is a common source of lost afternoons. The Kernel Hub addresses that: kernels are a repository type on the Hub, built for many PyTorch and CUDA combinations, and the kernels package (which needs PyTorch 2.5 or newer) downloads the build matching your environment. Transformers accepts a Hub kernel repository directly as the attention implementation, and other kernels can be loaded by hand:

model = AutoModelForCausalLM.from_pretrained(
    path, dtype=torch.bfloat16, device_map={"": 0},
    attn_implementation="kernels-community/flash-attn2",
)

from kernels import get_kernel
activation = get_kernel("kernels-community/activation", version=1)
x = torch.randn((10, 10), dtype=torch.float16, device="cuda")
y = torch.empty_like(x)
activation.gelu_fast(y, x)

Whichever path you choose, log it at startup and benchmark against sdpa on your sequence lengths; a kernel that wins on long prompts can lose on short ones.

Feeding the GPU with datasets

On the data side, datasets stores tables as Apache Arrow files and memory-maps them, so a dataset larger than RAM still opens instantly and DataLoader workers share the same pages instead of copying. map results are cached on disk under a fingerprint of the input and the function, so tokenisation runs once rather than every epoch. With several training processes, let one tokenise and the others reuse its cache:

from accelerate import Accelerator
from datasets import load_dataset

accelerator = Accelerator()
ds = load_dataset("parquet", data_files={"train": "data/train-*.parquet"}, split="train")
with accelerator.main_process_first():           # rank 0 builds the cache, others load it
    ds = ds.map(tokenize, batched=True, num_proc=8, remove_columns=ds.column_names)
ds.set_format("torch")
loader = torch.utils.data.DataLoader(ds, batch_size=8, shuffle=True, num_workers=4,
                                     pin_memory=True, collate_fn=collate)

If GPU utilisation dips at regular intervals, the input pipeline is the first suspect: pre-tokenise, raise num_workers, and keep padding per batch rather than to a global maximum.

Adapters out of training

For fine-tuning, PEFT wraps chosen linear layers with low-rank adapters, so gradients and optimizer state exist only for a few million parameters while the base weights stay frozen, often in 4-bit. TRL's trainers and Accelerate handle the loop and the process group; the TRL deep dive covers trainer choice and memory. What leaves the run is small: an adapter_config.json and an adapter_model.safetensors, typically tens of megabytes.

That file names its base model but should not be trusted to pin it. An adapter trained against one revision of a base and served on another produces plausible, subtly worse output, with no error. Record the base commit yourself when you push:

from huggingface_hub import HfApi

model.save_pretrained("out/adapter")
HfApi().upload_folder(
    repo_id="your-org/support-lora", folder_path="out/adapter",
    commit_message=f"LoRA r=16 on {REPO}@{SHA}",
)

Put the same hash in the adapter's model card. PEFT in depth explains loading, merging and multi-adapter batches.

The serving handoff

For serving, the handoff now usually leaves the Hugging Face libraries. vLLM's model_impl setting defaults to auto: use vLLM's own implementation if one exists, otherwise fall back to the Transformers model definition; --model-impl transformers forces the latter. Base and adapter go in together, both pinned:

vllm serve your-org/your-14b-model --revision full-40-character-commit-hash \
  --enable-lora --lora-modules support=/models/support-lora \
  --max-model-len 8192

Serve with the tokenizer and chat template from the same base revision you trained on. A different template changes the token sequence the adapter sees and degrades output with no error. Transformers also includes transformers serve, an OpenAI-compatible server that is handy for checking a model before moving it to a production engine. LLM serving stacks compares the engines.

Failure modes

FailureSymptomFix
Loading mainResults change between runs or nodesPin the commit hash everywhere
Every rank downloadsSlow start, rate limits, full disksOne download per node, then HF_HUB_OFFLINE=1
Silent CPU offloadCorrect but very slow tokensCheck hf_device_map; quantise or shard properly
Pickle weights or remote codeArbitrary code runs at loadPrefer safetensors; pin and read remote code
Attention backend missingLoad error or slower fallbackLog the implementation; benchmark against sdpa
Tokenisation inside the loopPeriodic GPU idle gapsCache map output; main process first
Adapter on another base revisionPlausible but worse answersRecord the base hash with the adapter
Template drift into servingFormat errors, rambling repliesServe tokenizer and template from the training revision

Trade-offs

Staying entirely in Transformers is the shortest path from experiment to answer, but device_map splitting and bitsandbytes are tools for fitting, not throughput. Serving engines are faster and batch better, at the cost of a second runtime whose model support and numerics you must check. 4-bit loading makes large models fit on small GPUs but costs some quality and per-token speed. Hub kernels remove build pain but add a runtime download, so pin the kernel version as you pin the model. For where these libraries sit relative to PyTorch and CUDA, see the Python LLM stack overview.

What to do next

  1. Replace every main in your training and serving configs with a full commit hash.
  2. Download once per node with allow_patterns and run jobs with HF_HUB_OFFLINE=1.
  3. Read one shard header with the script above and check the dtype and total size match your memory plan.
  4. Print hf_device_map after every load and fail the job if anything lands on CPU or disk unintentionally.
  5. Benchmark sdpa, flash attention and a Hub kernel on your real sequence lengths, and log the winner.
  6. Push adapters with the base commit hash in the commit message and model card.
  7. Serve base and adapter pinned, with the training revision's tokenizer and chat template.
Key takeaway: Treat the Hugging Face stack as a chain of handoffs. Pin the commit hash, load safetensors from a per-node cache, read the device map after every load because CPU offload is silent and slow, choose and log the attention backend, tokenise once, record the base revision with every adapter, and serve base, adapter and chat template from the same revision.