Hugging Face Transformers is best understood as a library of reference model definitions with a loader, tokenizers and a generation and training stack around them. When a lab releases a new architecture, the modeling file in Transformers is often the first readable implementation, and inference engines such as vLLM and SGLang can run models through it or load the same weights. That role shapes the architecture: one class per model family, explicit code rather than deep abstraction, and configuration files on the Hub that select which class to build.
Version 5, announced in December 2025 and released in January 2026, sharpened this. It dropped TensorFlow and Flax to become PyTorch-only, made dtype="auto" the default when loading, merged the separate "slow" and "fast" tokenizer classes into one class per model with a selectable backend, moved quantization into a reworked weight loader, and ships transformers serve, an OpenAI-compatible server. This article walks the layers in the order a from_pretrained call meets them, with v5 code, and shows where memory goes when you load and run a model.
From Hub repo to Python class
A model on the Hub is a git repository. The files that matter to the loader are config.json (architecture hyperparameters and a model_type string), the tokenizer files and chat template, an optional generation_config.json with default sampling settings, and the weights in safetensors format. Large models split weights into shards and add model.safetensors.index.json, whose weight_map says which shard holds each tensor.
Here is the configuration of a real small model, Qwen/Qwen2.5-0.5B-Instruct, trimmed to the fields the loader and your memory budget care about:
{
"architectures": ["Qwen2ForCausalLM"],
"model_type": "qwen2",
"hidden_size": 896,
"num_hidden_layers": 24,
"num_attention_heads": 14,
"num_key_value_heads": 2,
"vocab_size": 151936,
"max_position_embeddings": 32768,
"tie_word_embeddings": true,
"torch_dtype": "bfloat16"
}AutoConfig reads model_type and maps qwen2 to Qwen2Config; AutoModelForCausalLM then maps that config to Qwen2ForCausalLM. The Auto classes are just these lookup tables, one per task head, which is why the same script runs any supported architecture. Files are downloaded once into the cache under HF_HOME (by default ~/.cache/huggingface) and resolved by commit, so pass revision= with a commit hash in production: a branch name like main can move under you. Set HF_HUB_OFFLINE=1 on machines that must never reach the network, and use token= (v5 removed use_auth_token) for gated models.
From Hub repo to Python class
A model on the Hub is a git repository. The files that matter to the loader are config.json (architecture hyperparameters and a model_type string), the tokenizer files and chat template, an optional generation_config.json with default sampling settings, and the weights in safetensors format. Large models split weights into shards and add model.safetensors.index.json, whose weight_map says which shard holds each tensor.
Here is the configuration of a real small model, Qwen/Qwen2.5-0.5B-Instruct, trimmed to the fields the loader and your memory budget care about:
{
"architectures": ["Qwen2ForCausalLM"],
"model_type": "qwen2",
"hidden_size": 896,
"num_hidden_layers": 24,
"num_attention_heads": 14,
"num_key_value_heads": 2,
"vocab_size": 151936,
"max_position_embeddings": 32768,
"tie_word_embeddings": true,
"torch_dtype": "bfloat16"
}AutoConfig reads model_type and maps qwen2 to Qwen2Config; AutoModelForCausalLM then maps that config to Qwen2ForCausalLM. The Auto classes are just these lookup tables, one per task head, which is why the same script runs any supported architecture. Files are downloaded once into the cache under HF_HOME (by default ~/.cache/huggingface) and resolved by commit, so pass revision= with a commit hash in production: a branch name like main can move under you. Set HF_HUB_OFFLINE=1 on machines that must never reach the network, and use token= (v5 removed use_auth_token) for gated models.
PreTrainedModel, modeling files and modular
Every model class inherits from PreTrainedModel, which supplies loading and saving, device and dtype handling, gradient checkpointing, resizing embeddings and the hooks that attention backends and quantizers plug into. The architecture itself lives in one modeling_<name>.py file that you can read top to bottom: embedding, a stack of decoder layers, a final norm and a head.
Maintaining hundreds of near-identical files used to mean copy-and-paste. The modular approach fixes that: a contributor writes a short modular_<name>.py that inherits from an existing model and overrides only what differs, and a converter generates the full, flat modeling file from it. You still read and debug the flat file; the modular file shows you, in a few dozen lines, how a new model differs from its ancestor.
Models whose code is not in the library can ship it in their repo and be loaded with trust_remote_code=True. That executes Python from the repository on your machine. Pin the revision and review the code before enabling it, exactly as you would for any third-party dependency.
Loading: dtype, device_map, quantization and the memory math
Loading is where most first attempts fail, and the arithmetic is simple. Weight memory is parameter count times bytes per parameter: 2 bytes in bfloat16 or float16, 4 in float32, roughly half a byte for 4-bit formats plus some overhead for scales. A 7-billion-parameter model needs about 14 GB in bfloat16, 28 GB in float32 and around 4 to 5 GB in 4-bit. Before v5 the default dtype was float32 unless you asked otherwise, which silently doubled memory; v5 defaults to dtype="auto", which uses the dtype recorded in the checkpoint.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
model_id = "Qwen/Qwen2.5-0.5B-Instruct"
rev = "main" # replace with a commit hash in production
# Full precision of the checkpoint, spread over available devices
model = AutoModelForCausalLM.from_pretrained(model_id, revision=rev, dtype="auto", device_map="auto")
# 4-bit load: quantization goes through quantization_config in v5
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16)
model_4bit = AutoModelForCausalLM.from_pretrained(model_id, revision=rev,
quantization_config=bnb, device_map="auto")device_map="auto" uses Accelerate's big-model support to place layers on GPUs first, then CPU, then disk, so a model that does not fit still loads, slowly. Check model.hf_device_map after loading: layers on CPU or disk explain most "why is generation so slow" questions. The v5 loader also converts checkpoints on the fly, for example merging or splitting tensors whose layout differs from the current modeling code, which is what lets quantization and new checkpoint formats plug in without a separate conversion step.
Tokenizers, processors and chat templates
A tokenizer turns text into IDs and back, and for chat models it also owns the chat template, a Jinja template stored with the tokenizer that formats a list of messages into the exact string the model was trained on. Getting this wrong is the most common cause of a fine-tuned or served model that "works but answers badly", because the model sees role markers it was never trained with.
tok = AutoTokenizer.from_pretrained(model_id, revision=rev)
messages = [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "Explain KV caching in one sentence."},
]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True,
return_tensors="pt", return_dict=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=60, do_sample=False)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))In v5 each model family has one tokenizer class, backed by the Rust tokenizers library by default, with SentencePiece and other backends available where a model needs them. Calling the tokenizer directly replaces encode_plus, and apply_chat_template returns a dictionary-like encoding with input_ids and attention_mask, which is why the example unpacks it into generate. Multimodal models use a Processor that bundles the tokenizer with an image, audio or video processor. Details are in the tokenizers article.
Attention backends and the KV cache
Attention is the one component whose best implementation depends on hardware and sequence length, so it is pluggable. The modeling file keeps a plain eager implementation for readability, and the AttentionInterface registry provides the others: PyTorch's scaled dot-product attention (sdpa, the usual default), FlashAttention variants and FlexAttention. You choose with attn_implementation="sdpa" (or "flash_attention_2" when that package is installed and your GPU supports it) at load time. Use eager when you need attention weights returned or are debugging numerical differences.
The KV cache is the other memory consumer. Per token it stores keys and values for every layer: 2 x layers x KV heads x head dimension x bytes. For Qwen2.5-0.5B that is 2 x 24 x 2 x 64 x 2 bytes = 12 KiB per token, so one 32,768-token sequence holds 384 MiB, against roughly 1 GB of bfloat16 weights. Grouped query attention, two KV heads instead of fourteen, is what keeps that number small; see the KV cache article for the general case.
Generation and serving
generate() runs the decoding loop: a prefill pass over the prompt that fills the cache, then one forward pass per new token. Its behaviour comes from a GenerationConfig, loaded from the repo and overridable per call: max_new_tokens, do_sample, temperature, top_p, top_k, repetition_penalty, stop strings and beam search. Logits processors let you constrain output token by token, and assisted generation uses a smaller draft model to propose tokens. Read the generate article for the options.
For serving many users, v5 ships continuous batching with a paged KV cache inside the library and the transformers serve command, which exposes an OpenAI-compatible API. It is the quickest way to put a model behind an HTTP endpoint for evaluation or a small internal tool. For high-throughput production traffic, dedicated engines such as vLLM or SGLang usually still win on throughput and scheduling features; the practical pattern is to develop and evaluate in Transformers and serve the same checkpoint, pinned to the same revision, in an engine.
Worked example: LoRA fine-tuning on the same stack
The training side has three layers. Trainer owns the loop: batching, gradient accumulation, mixed precision, evaluation, checkpointing and logging (in v5 report_to defaults to "none", so enable your tracker explicitly). Accelerate sits underneath and turns the same script into DDP, FSDP or DeepSpeed through accelerate launch and a config file. PEFT adds parameter-efficient methods such as LoRA on top of any model. A compact LoRA run:
from datasets import load_dataset
from peft import LoraConfig, get_peft_model
from transformers import DataCollatorForLanguageModeling, Trainer, TrainingArguments
if tok.pad_token is None:
tok.pad_token = tok.eos_token
ds = load_dataset("json", data_files="train.jsonl")["train"] # rows: {"messages": [...]}
ds = ds.map(lambda r: tok(tok.apply_chat_template(r["messages"], tokenize=False),
truncation=True, max_length=1024), remove_columns=ds.column_names)
model = get_peft_model(model, LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
task_type="CAUSAL_LM"))
model.print_trainable_parameters()
args = TrainingArguments(output_dir="out", per_device_train_batch_size=4,
gradient_accumulation_steps=4, learning_rate=2e-4,
num_train_epochs=1, bf16=True, logging_steps=10, save_steps=200)
Trainer(model=model, args=args, train_dataset=ds,
data_collator=DataCollatorForLanguageModeling(tok, mlm=False)).train()
model.save_pretrained("lora-adapter") # adapter weights onlySaved adapters are small, megabytes rather than gigabytes, and load on top of the pinned base model. For serving you can merge them into the base weights or load them in an engine that supports LoRA. See Trainer, PEFT and Accelerate for each layer in depth.
Trade-offs: pipeline, Auto classes or a serving engine
There are three ways into the library, and they trade convenience for control. pipeline() hides tokenization, batching and post-processing behind one call; it is ideal for a demo, a classification job or a quick evaluation, but it chooses defaults you cannot always see, such as truncation length and padding. The Auto classes with generate() or a forward pass give you every tensor, which is what you want for research, custom decoding and fine-tuning. Handing the pinned checkpoint to a serving engine gives the best throughput and latency under concurrent load, at the cost of a second runtime whose numerics and defaults you must check against the reference.
The library also trades speed for readability on purpose. The modeling files avoid fused kernels in the main code path so they stay readable and portable; performance comes from the attention interface, compilation with torch.compile, quantization and the serving features. If a model behaves differently in two engines, the Transformers implementation in eager mode is the one to compare against, because it is the closest to the authors' reference.
Upgrading code to v5
- Replace
torch_dtype=withdtype=, and check any code that assumed float32 weights. - Replace
load_in_4bit=Trueandload_in_8bit=Truearguments withquantization_config=BitsAndBytesConfig(...). - Replace
use_auth_tokenwithtoken; v5 requires a 1.xhuggingface_huband Python 3.10 or newer. - Remove TensorFlow and Flax model classes; convert or retire that code.
- Treat
apply_chat_templateoutput as an encoding, not a bare tensor. - Run your evaluation suite before and after: tokenizer and dtype changes can shift outputs without errors.
Failure modes
- Out of memory at load. Wrong dtype, no
device_map, or a 4-bit expectation without a quantization config. Do the parameter arithmetic first. - Silent CPU offload.
device_map="auto"put layers on CPU and throughput collapsed. Inspecthf_device_map. - Moving revisions. A model or tokenizer changed on
mainand outputs drifted. Pin commit hashes for model, tokenizer and dataset. - Template mismatch. Training data formatted by hand while inference uses the chat template, or the reverse. Use the same template in both.
- Padding side. Batched generation with right padding on a decoder-only model produces garbage for shorter prompts. Set
tok.padding_side = "left"for batched generation. - Remote code.
trust_remote_code=Trueon an unpinned repo runs whatever is pushed next.
What to do next
- Pin a model and tokenizer revision and record them with every experiment.
- Compute weight and KV cache memory for your longest context before choosing hardware.
- Load with
dtype="auto"and an explicitattn_implementation; checkhf_device_map. - Format every prompt with
apply_chat_template, in training and inference alike. - Fine-tune with PEFT first; move to full fine-tuning only if adapters plateau.
- Evaluate in Transformers, then serve the same pinned checkpoint in your production engine.
- Before upgrading to v5, grep for
torch_dtype,load_in_anduse_auth_token.