Hugging Face Transformers is best understood as a library of reference model definitions with a loader, tokenizers and a generation and training stack around them. When a lab releases a new architecture, the modeling file in Transformers is often the first readable implementation, and inference engines such as vLLM and SGLang can run models through it or load the same weights. That role shapes the architecture: one class per model family, explicit code rather than deep abstraction, and configuration files on the Hub that select which class to build.

Version 5, announced in December 2025 and released in January 2026, sharpened this. It dropped TensorFlow and Flax to become PyTorch-only, made dtype="auto" the default when loading, merged the separate "slow" and "fast" tokenizer classes into one class per model with a selectable backend, moved quantization into a reworked weight loader, and ships transformers serve, an OpenAI-compatible server. This article walks the layers in the order a from_pretrained call meets them, with v5 code, and shows where memory goes when you load and run a model.

Hugging Face Transformers v5: the layers a from_pretrained call passes throughYour codepipeline() | AutoModel + generate() | Trainer / TRL | transformers serveHub repo @ revisionconfig.json, tokenizer, safetensorsAuto classesmodel_type -> concrete classLocal cacheHF_HOME/hub, offline modeTokenizer / Processorone backend in v5, chat templatePreTrainedModelmodeling_*.py, built from modular_*.pyWeight loaderdtype auto, device_map, quantizationAttentionInterfaceeager | sdpa | flash | flexgenerate()GenerationConfig, KV cache, logits processorsTrainer + AccelerateDDP, FSDP, DeepSpeedExport / servevLLM, SGLang, serve, PEFT adaptersThe modeling file is the reference implementation; other engines reuse it or its weights.
A from_pretrained call resolves a pinned Hub revision, picks the concrete class through the Auto mappings, loads weights with the requested dtype, placement and quantization, and plugs in an attention backend before generation, training or export.

From Hub repo to Python class

A model on the Hub is a git repository. The files that matter to the loader are config.json (architecture hyperparameters and a model_type string), the tokenizer files and chat template, an optional generation_config.json with default sampling settings, and the weights in safetensors format. Large models split weights into shards and add model.safetensors.index.json, whose weight_map says which shard holds each tensor.

Here is the configuration of a real small model, Qwen/Qwen2.5-0.5B-Instruct, trimmed to the fields the loader and your memory budget care about:

{
  "architectures": ["Qwen2ForCausalLM"],
  "model_type": "qwen2",
  "hidden_size": 896,
  "num_hidden_layers": 24,
  "num_attention_heads": 14,
  "num_key_value_heads": 2,
  "vocab_size": 151936,
  "max_position_embeddings": 32768,
  "tie_word_embeddings": true,
  "torch_dtype": "bfloat16"
}

AutoConfig reads model_type and maps qwen2 to Qwen2Config; AutoModelForCausalLM then maps that config to Qwen2ForCausalLM. The Auto classes are just these lookup tables, one per task head, which is why the same script runs any supported architecture. Files are downloaded once into the cache under HF_HOME (by default ~/.cache/huggingface) and resolved by commit, so pass revision= with a commit hash in production: a branch name like main can move under you. Set HF_HUB_OFFLINE=1 on machines that must never reach the network, and use token= (v5 removed use_auth_token) for gated models.

From Hub repo to Python class

A model on the Hub is a git repository. The files that matter to the loader are config.json (architecture hyperparameters and a model_type string), the tokenizer files and chat template, an optional generation_config.json with default sampling settings, and the weights in safetensors format. Large models split weights into shards and add model.safetensors.index.json, whose weight_map says which shard holds each tensor.

Here is the configuration of a real small model, Qwen/Qwen2.5-0.5B-Instruct, trimmed to the fields the loader and your memory budget care about:

{
  "architectures": ["Qwen2ForCausalLM"],
  "model_type": "qwen2",
  "hidden_size": 896,
  "num_hidden_layers": 24,
  "num_attention_heads": 14,
  "num_key_value_heads": 2,
  "vocab_size": 151936,
  "max_position_embeddings": 32768,
  "tie_word_embeddings": true,
  "torch_dtype": "bfloat16"
}

AutoConfig reads model_type and maps qwen2 to Qwen2Config; AutoModelForCausalLM then maps that config to Qwen2ForCausalLM. The Auto classes are just these lookup tables, one per task head, which is why the same script runs any supported architecture. Files are downloaded once into the cache under HF_HOME (by default ~/.cache/huggingface) and resolved by commit, so pass revision= with a commit hash in production: a branch name like main can move under you. Set HF_HUB_OFFLINE=1 on machines that must never reach the network, and use token= (v5 removed use_auth_token) for gated models.

PreTrainedModel, modeling files and modular

Every model class inherits from PreTrainedModel, which supplies loading and saving, device and dtype handling, gradient checkpointing, resizing embeddings and the hooks that attention backends and quantizers plug into. The architecture itself lives in one modeling_<name>.py file that you can read top to bottom: embedding, a stack of decoder layers, a final norm and a head.

Maintaining hundreds of near-identical files used to mean copy-and-paste. The modular approach fixes that: a contributor writes a short modular_<name>.py that inherits from an existing model and overrides only what differs, and a converter generates the full, flat modeling file from it. You still read and debug the flat file; the modular file shows you, in a few dozen lines, how a new model differs from its ancestor.

Models whose code is not in the library can ship it in their repo and be loaded with trust_remote_code=True. That executes Python from the repository on your machine. Pin the revision and review the code before enabling it, exactly as you would for any third-party dependency.

Loading: dtype, device_map, quantization and the memory math

Loading is where most first attempts fail, and the arithmetic is simple. Weight memory is parameter count times bytes per parameter: 2 bytes in bfloat16 or float16, 4 in float32, roughly half a byte for 4-bit formats plus some overhead for scales. A 7-billion-parameter model needs about 14 GB in bfloat16, 28 GB in float32 and around 4 to 5 GB in 4-bit. Before v5 the default dtype was float32 unless you asked otherwise, which silently doubled memory; v5 defaults to dtype="auto", which uses the dtype recorded in the checkpoint.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

model_id = "Qwen/Qwen2.5-0.5B-Instruct"
rev = "main"  # replace with a commit hash in production

# Full precision of the checkpoint, spread over available devices
model = AutoModelForCausalLM.from_pretrained(model_id, revision=rev, dtype="auto", device_map="auto")

# 4-bit load: quantization goes through quantization_config in v5
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_quant_type="nf4",
                         bnb_4bit_compute_dtype=torch.bfloat16)
model_4bit = AutoModelForCausalLM.from_pretrained(model_id, revision=rev,
                                                  quantization_config=bnb, device_map="auto")

device_map="auto" uses Accelerate's big-model support to place layers on GPUs first, then CPU, then disk, so a model that does not fit still loads, slowly. Check model.hf_device_map after loading: layers on CPU or disk explain most "why is generation so slow" questions. The v5 loader also converts checkpoints on the fly, for example merging or splitting tensors whose layout differs from the current modeling code, which is what lets quantization and new checkpoint formats plug in without a separate conversion step.

Tokenizers, processors and chat templates

A tokenizer turns text into IDs and back, and for chat models it also owns the chat template, a Jinja template stored with the tokenizer that formats a list of messages into the exact string the model was trained on. Getting this wrong is the most common cause of a fine-tuned or served model that "works but answers badly", because the model sees role markers it was never trained with.

tok = AutoTokenizer.from_pretrained(model_id, revision=rev)
messages = [
    {"role": "system", "content": "You are a concise assistant."},
    {"role": "user", "content": "Explain KV caching in one sentence."},
]
inputs = tok.apply_chat_template(messages, add_generation_prompt=True,
                                 return_tensors="pt", return_dict=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=60, do_sample=False)
print(tok.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))

In v5 each model family has one tokenizer class, backed by the Rust tokenizers library by default, with SentencePiece and other backends available where a model needs them. Calling the tokenizer directly replaces encode_plus, and apply_chat_template returns a dictionary-like encoding with input_ids and attention_mask, which is why the example unpacks it into generate. Multimodal models use a Processor that bundles the tokenizer with an image, audio or video processor. Details are in the tokenizers article.

Attention backends and the KV cache

Attention is the one component whose best implementation depends on hardware and sequence length, so it is pluggable. The modeling file keeps a plain eager implementation for readability, and the AttentionInterface registry provides the others: PyTorch's scaled dot-product attention (sdpa, the usual default), FlashAttention variants and FlexAttention. You choose with attn_implementation="sdpa" (or "flash_attention_2" when that package is installed and your GPU supports it) at load time. Use eager when you need attention weights returned or are debugging numerical differences.

The KV cache is the other memory consumer. Per token it stores keys and values for every layer: 2 x layers x KV heads x head dimension x bytes. For Qwen2.5-0.5B that is 2 x 24 x 2 x 64 x 2 bytes = 12 KiB per token, so one 32,768-token sequence holds 384 MiB, against roughly 1 GB of bfloat16 weights. Grouped query attention, two KV heads instead of fourteen, is what keeps that number small; see the KV cache article for the general case.

Generation and serving

generate() runs the decoding loop: a prefill pass over the prompt that fills the cache, then one forward pass per new token. Its behaviour comes from a GenerationConfig, loaded from the repo and overridable per call: max_new_tokens, do_sample, temperature, top_p, top_k, repetition_penalty, stop strings and beam search. Logits processors let you constrain output token by token, and assisted generation uses a smaller draft model to propose tokens. Read the generate article for the options.

For serving many users, v5 ships continuous batching with a paged KV cache inside the library and the transformers serve command, which exposes an OpenAI-compatible API. It is the quickest way to put a model behind an HTTP endpoint for evaluation or a small internal tool. For high-throughput production traffic, dedicated engines such as vLLM or SGLang usually still win on throughput and scheduling features; the practical pattern is to develop and evaluate in Transformers and serve the same checkpoint, pinned to the same revision, in an engine.

Worked example: LoRA fine-tuning on the same stack

The training side has three layers. Trainer owns the loop: batching, gradient accumulation, mixed precision, evaluation, checkpointing and logging (in v5 report_to defaults to "none", so enable your tracker explicitly). Accelerate sits underneath and turns the same script into DDP, FSDP or DeepSpeed through accelerate launch and a config file. PEFT adds parameter-efficient methods such as LoRA on top of any model. A compact LoRA run:

from datasets import load_dataset
from peft import LoraConfig, get_peft_model
from transformers import DataCollatorForLanguageModeling, Trainer, TrainingArguments

if tok.pad_token is None:
    tok.pad_token = tok.eos_token
ds = load_dataset("json", data_files="train.jsonl")["train"]          # rows: {"messages": [...]}
ds = ds.map(lambda r: tok(tok.apply_chat_template(r["messages"], tokenize=False),
                          truncation=True, max_length=1024), remove_columns=ds.column_names)

model = get_peft_model(model, LoraConfig(r=16, lora_alpha=32, lora_dropout=0.05,
                                         target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
                                         task_type="CAUSAL_LM"))
model.print_trainable_parameters()

args = TrainingArguments(output_dir="out", per_device_train_batch_size=4,
                         gradient_accumulation_steps=4, learning_rate=2e-4,
                         num_train_epochs=1, bf16=True, logging_steps=10, save_steps=200)
Trainer(model=model, args=args, train_dataset=ds,
        data_collator=DataCollatorForLanguageModeling(tok, mlm=False)).train()
model.save_pretrained("lora-adapter")    # adapter weights only

Saved adapters are small, megabytes rather than gigabytes, and load on top of the pinned base model. For serving you can merge them into the base weights or load them in an engine that supports LoRA. See Trainer, PEFT and Accelerate for each layer in depth.

Trade-offs: pipeline, Auto classes or a serving engine

There are three ways into the library, and they trade convenience for control. pipeline() hides tokenization, batching and post-processing behind one call; it is ideal for a demo, a classification job or a quick evaluation, but it chooses defaults you cannot always see, such as truncation length and padding. The Auto classes with generate() or a forward pass give you every tensor, which is what you want for research, custom decoding and fine-tuning. Handing the pinned checkpoint to a serving engine gives the best throughput and latency under concurrent load, at the cost of a second runtime whose numerics and defaults you must check against the reference.

The library also trades speed for readability on purpose. The modeling files avoid fused kernels in the main code path so they stay readable and portable; performance comes from the attention interface, compilation with torch.compile, quantization and the serving features. If a model behaves differently in two engines, the Transformers implementation in eager mode is the one to compare against, because it is the closest to the authors' reference.

Upgrading code to v5

  • Replace torch_dtype= with dtype=, and check any code that assumed float32 weights.
  • Replace load_in_4bit=True and load_in_8bit=True arguments with quantization_config=BitsAndBytesConfig(...).
  • Replace use_auth_token with token; v5 requires a 1.x huggingface_hub and Python 3.10 or newer.
  • Remove TensorFlow and Flax model classes; convert or retire that code.
  • Treat apply_chat_template output as an encoding, not a bare tensor.
  • Run your evaluation suite before and after: tokenizer and dtype changes can shift outputs without errors.

Failure modes

  • Out of memory at load. Wrong dtype, no device_map, or a 4-bit expectation without a quantization config. Do the parameter arithmetic first.
  • Silent CPU offload. device_map="auto" put layers on CPU and throughput collapsed. Inspect hf_device_map.
  • Moving revisions. A model or tokenizer changed on main and outputs drifted. Pin commit hashes for model, tokenizer and dataset.
  • Template mismatch. Training data formatted by hand while inference uses the chat template, or the reverse. Use the same template in both.
  • Padding side. Batched generation with right padding on a decoder-only model produces garbage for shorter prompts. Set tok.padding_side = "left" for batched generation.
  • Remote code. trust_remote_code=True on an unpinned repo runs whatever is pushed next.

What to do next

  1. Pin a model and tokenizer revision and record them with every experiment.
  2. Compute weight and KV cache memory for your longest context before choosing hardware.
  3. Load with dtype="auto" and an explicit attn_implementation; check hf_device_map.
  4. Format every prompt with apply_chat_template, in training and inference alike.
  5. Fine-tune with PEFT first; move to full fine-tuning only if adapters plateau.
  6. Evaluate in Transformers, then serve the same pinned checkpoint in your production engine.
  7. Before upgrading to v5, grep for torch_dtype, load_in_ and use_auth_token.
Key takeaway: Transformers is a set of readable reference model definitions plus the machinery to load, run, train and export them. A config's model_type picks the class, the loader applies dtype, device placement and quantization, an attention backend is plugged in, and generate or Trainer drive the model. Pin revisions, do the memory arithmetic before loading, always use the chat template, and when moving to v5 update dtype, quantization and auth arguments and re-run your evaluations.