The Hugging Face libraries are usually described one at a time: the Hub stores models, Transformers defines them, Datasets feeds them, PEFT and TRL train them, Accelerate spreads them over devices. On a GPU box what matters is the handoffs. A checkpoint is downloaded, memory-mapped, placed layer by layer on devices, run through attention kernels chosen at load time, trained through small adapters, and handed to a serving engine that is not a Hugging Face library at all. Most production surprises with this stack, such as a slow model, a run you cannot reproduce or an adapter that misbehaves in serving, happen at one of those handoffs.
This article follows one checkpoint along that path and says what each step does with bytes and GPU memory. Two recent changes shape it. Transformers v5, whose first release candidate appeared on 1 December 2025, made PyTorch its only backend, kept a single tokenizer backend instead of fast and slow variants, and reworked weight loading so quantization is handled as part of loading. And Text Generation Inference, Hugging Face's own server, has been in maintenance mode since December 2025; its documentation points to vLLM and SGLang, which can run models through the Transformers definitions.
The path in one picture
Pin the bytes before anything else
Everything downstream depends on knowing exactly which bytes you loaded. A Hub repository is a git repository, and main moves when the author pushes. Pin the full commit hash, download only the files the loader needs, and do it once per node into a shared cache rather than once per training process:
from huggingface_hub import snapshot_download
REPO = "your-org/your-14b-model"
SHA = "full-40-character-commit-hash"
path = snapshot_download(
REPO, revision=SHA,
allow_patterns=["*.safetensors", "*.json", "tokenizer*"], # skip .bin duplicates and extras
)
# Training and serving processes then run with HF_HUB_OFFLINE=1 and load from `path`.The cache layout, revision handling and offline mode are covered in the Hugging Face Hub in depth. The GPU-relevant point is simpler: eight ranks each downloading 28 GB at job start is a slow start and a rate-limit risk, and two nodes resolving main a minute apart can load different weights into the same job.
What is inside a safetensors file
A .safetensors file is deliberately simple. The first 8 bytes are a little-endian unsigned integer giving the header length. The header is JSON mapping each tensor name to its dtype, shape and a pair of byte offsets, plus an optional __metadata__ entry. After the header come the raw tensor bytes. A sharded checkpoint adds model.safetensors.index.json, which maps each tensor name to its shard file. You can read the header yourself:
import json, struct
def read_header(path):
with open(path, "rb") as f:
n = struct.unpack("<Q", f.read(8))[0] # header length
header = json.loads(f.read(n))
meta = header.pop("__metadata__", {})
return meta, header # name -> {"dtype", "shape", "data_offsets": [start, end]}
meta, tensors = read_header("model-00001-of-00006.safetensors")
total = sum(t["data_offsets"][1] - t["data_offsets"][0] for t in tensors.values())
print(len(tensors), "tensors,", total / 2**30, "GiB in this shard")This format has two consequences for loading. The file can be memory-mapped, so a loader reads only the tensors it needs, straight from the page cache, without parsing the rest; a process that needs only some of the layers touches only those bytes. And loading runs no code, unlike the older pickle-based .bin files, which can execute arbitrary Python when unpickled. Prefer safetensors files, and treat any repository that only ships pickles, or asks for trust_remote_code=True, as code you are about to run: pin its revision and read it.
How from_pretrained places weights on devices
When from_pretrained is given a device_map, it first builds the model on PyTorch's meta device, where tensors have shapes but no storage, so building a 14-billion-parameter skeleton allocates nothing. It then decides where each module goes and materialises each tensor directly on its target device from the memory-mapped shards. With device_map="auto", Accelerate fills GPUs in order up to the max_memory you allow, then CPU memory, then disk if you give it an offload folder.
import torch
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained(
path,
dtype=torch.bfloat16, # older releases call this torch_dtype
device_map="auto",
max_memory={0: "12GiB", 1: "12GiB", "cpu": "48GiB"},
attn_implementation="sdpa",
)
print(model.hf_device_map) # module name -> 0, 1, "cpu" or "disk"Two properties of this placement are easy to miss. First, it is a sequential split, not parallelism: layers on GPU 0 run, then layers on GPU 1 run, so at batch size 1 only one GPU is busy at any moment. It lets a model fit; it does not make it faster. Second, modules placed on cpu are not computed on the CPU. Accelerate attaches hooks that copy their weights to the GPU for each forward pass, so every generated token pays for that copy over PCIe.
Worked example: a 14B model on two 16 GB GPUs
Suppose you want to run a 14B model in bfloat16 on a box with two 16 GB GPUs. The weights are about 14 x 10^9 x 2 bytes, roughly 28 GB or 26 GiB. You cap each GPU at 12 GiB to leave room for activations, the KV cache and the CUDA context, so 24 GiB of weights fit and about 2 GiB of layers land on cpu.
Loading succeeds and the first answer is correct, which is why this goes unnoticed. But each decoding step now copies about 2 GiB from host memory. At an effective 20 to 25 GB/s over PCIe 4.0 x16, that is on the order of 100 ms per token for the copy alone, likely more than the rest of the step. Printing hf_device_map and counting entries that are not GPU indices shows the problem in one line.
There are three honest fixes. Load the weights in 4-bit NF4 with bitsandbytes: about half a byte per parameter plus small per-block scales, roughly 8 GB, which fits on one GPU with room for the KV cache, at some cost in quality and per-token speed. Use a checkpoint already quantised with a method such as AWQ or GPTQ, which Transformers reads from the quantization config in config.json. Or move serving to an engine with real tensor parallelism and check the memory budget there. Raising max_memory until the CPU entries vanish is not a fix; it moves the failure to an out-of-memory error at the first long prompt.
from transformers import BitsAndBytesConfig
bnb = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16, # weights are dequantised to this for each matmul
bnb_4bit_use_double_quant=True,
)
model = AutoModelForCausalLM.from_pretrained(path, quantization_config=bnb, device_map={"": 0})
Choosing attention kernels
The attn_implementation argument picks the attention code path when the model is built. "sdpa" uses PyTorch's scaled_dot_product_attention, which dispatches to a fused kernel when the dtype, head size and mask allow. "flash_attention_2" needs the separately compiled flash-attn package and half-precision inputs. "eager" is the plain implementation, slow but useful when you need attention weights or are debugging numerics.
Compiling flash-attn against the right CUDA and PyTorch versions is a common source of lost afternoons. The Kernel Hub addresses that: kernels are a repository type on the Hub, built for many PyTorch and CUDA combinations, and the kernels package (which needs PyTorch 2.5 or newer) downloads the build matching your environment. Transformers accepts a Hub kernel repository directly as the attention implementation, and other kernels can be loaded by hand:
model = AutoModelForCausalLM.from_pretrained(
path, dtype=torch.bfloat16, device_map={"": 0},
attn_implementation="kernels-community/flash-attn2",
)
from kernels import get_kernel
activation = get_kernel("kernels-community/activation", version=1)
x = torch.randn((10, 10), dtype=torch.float16, device="cuda")
y = torch.empty_like(x)
activation.gelu_fast(y, x)Whichever path you choose, log it at startup and benchmark against sdpa on your sequence lengths; a kernel that wins on long prompts can lose on short ones.
Feeding the GPU with datasets
On the data side, datasets stores tables as Apache Arrow files and memory-maps them, so a dataset larger than RAM still opens instantly and DataLoader workers share the same pages instead of copying. map results are cached on disk under a fingerprint of the input and the function, so tokenisation runs once rather than every epoch. With several training processes, let one tokenise and the others reuse its cache:
from accelerate import Accelerator
from datasets import load_dataset
accelerator = Accelerator()
ds = load_dataset("parquet", data_files={"train": "data/train-*.parquet"}, split="train")
with accelerator.main_process_first(): # rank 0 builds the cache, others load it
ds = ds.map(tokenize, batched=True, num_proc=8, remove_columns=ds.column_names)
ds.set_format("torch")
loader = torch.utils.data.DataLoader(ds, batch_size=8, shuffle=True, num_workers=4,
pin_memory=True, collate_fn=collate)If GPU utilisation dips at regular intervals, the input pipeline is the first suspect: pre-tokenise, raise num_workers, and keep padding per batch rather than to a global maximum.
Adapters out of training
For fine-tuning, PEFT wraps chosen linear layers with low-rank adapters, so gradients and optimizer state exist only for a few million parameters while the base weights stay frozen, often in 4-bit. TRL's trainers and Accelerate handle the loop and the process group; the TRL deep dive covers trainer choice and memory. What leaves the run is small: an adapter_config.json and an adapter_model.safetensors, typically tens of megabytes.
That file names its base model but should not be trusted to pin it. An adapter trained against one revision of a base and served on another produces plausible, subtly worse output, with no error. Record the base commit yourself when you push:
from huggingface_hub import HfApi
model.save_pretrained("out/adapter")
HfApi().upload_folder(
repo_id="your-org/support-lora", folder_path="out/adapter",
commit_message=f"LoRA r=16 on {REPO}@{SHA}",
)Put the same hash in the adapter's model card. PEFT in depth explains loading, merging and multi-adapter batches.
The serving handoff
For serving, the handoff now usually leaves the Hugging Face libraries. vLLM's model_impl setting defaults to auto: use vLLM's own implementation if one exists, otherwise fall back to the Transformers model definition; --model-impl transformers forces the latter. Base and adapter go in together, both pinned:
vllm serve your-org/your-14b-model --revision full-40-character-commit-hash \
--enable-lora --lora-modules support=/models/support-lora \
--max-model-len 8192Serve with the tokenizer and chat template from the same base revision you trained on. A different template changes the token sequence the adapter sees and degrades output with no error. Transformers also includes transformers serve, an OpenAI-compatible server that is handy for checking a model before moving it to a production engine. LLM serving stacks compares the engines.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
Loading main | Results change between runs or nodes | Pin the commit hash everywhere |
| Every rank downloads | Slow start, rate limits, full disks | One download per node, then HF_HUB_OFFLINE=1 |
| Silent CPU offload | Correct but very slow tokens | Check hf_device_map; quantise or shard properly |
| Pickle weights or remote code | Arbitrary code runs at load | Prefer safetensors; pin and read remote code |
| Attention backend missing | Load error or slower fallback | Log the implementation; benchmark against sdpa |
| Tokenisation inside the loop | Periodic GPU idle gaps | Cache map output; main process first |
| Adapter on another base revision | Plausible but worse answers | Record the base hash with the adapter |
| Template drift into serving | Format errors, rambling replies | Serve tokenizer and template from the training revision |
Trade-offs
Staying entirely in Transformers is the shortest path from experiment to answer, but device_map splitting and bitsandbytes are tools for fitting, not throughput. Serving engines are faster and batch better, at the cost of a second runtime whose model support and numerics you must check. 4-bit loading makes large models fit on small GPUs but costs some quality and per-token speed. Hub kernels remove build pain but add a runtime download, so pin the kernel version as you pin the model. For where these libraries sit relative to PyTorch and CUDA, see the Python LLM stack overview.
What to do next
- Replace every
mainin your training and serving configs with a full commit hash. - Download once per node with
allow_patternsand run jobs withHF_HUB_OFFLINE=1. - Read one shard header with the script above and check the dtype and total size match your memory plan.
- Print
hf_device_mapafter every load and fail the job if anything lands on CPU or disk unintentionally. - Benchmark
sdpa, flash attention and a Hub kernel on your real sequence lengths, and log the winner. - Push adapters with the base commit hash in the commit message and model card.
- Serve base and adapter pinned, with the training revision's tokenizer and chat template.