Hugging Face PEFT, in depth: in-place injection, adapter files, loading, multi-adapter batches, hot-swapping and safe merging

By Sandeep Belgavi · 2026-10-03 · Category: Transformers
Advertisement
Base modelq_proj: nn.Linearget_peft_modelPeftModel > LoraModel > base model (modified in place)base_layerfrozen nn.Linearlora_A[name], lora_B[name]ModuleDict per adaptery = base_layer(x) + scaling x B(A(dropout(x)))scaling = alpha / r, or alpha / sqrt(r) with use_rsloraactive adapter chosen by set_adapter or adapter_namessave_pretrainedadapter onlyfrom_pretrainedis_trainable flagmerge_and_unloadplain model outhotswap_adaptersame shapes, new weightsFiles: adapter_config.json + adapter_model.safetensors. Keys drop the adapter name.
What the PEFT library does to a model: each targeted layer is replaced in place by a wrapper that keeps the original as base_layer and adds per-adapter low-rank modules; the lifecycle runs from saving the adapter alone through loading, merging and hot-swapping.

Choosing a parameter-efficient method is a modelling decision. Using Hugging Face's peft library well is an engineering one: knowing what it does to your model object, what lands on disk, how an adapter is reloaded, how several adapters share one base, and when merging is safe. Most production PEFT bugs are lifecycle bugs, such as an adapter loaded frozen when training was meant to resume, attached to the wrong base revision, or merged into a quantized model, rather than bad hyperparameters.

This article covers the library's runtime surface, with API names and signatures checked against the PEFT 0.21 documentation. Method choice and parameter counts for each technique are in PEFT, in depth, and the mathematics of the low-rank update in LoRA, the math. Examples use LoRA, the method most PEFT features support first, and Transformers v5 conventions such as dtype=.

What get_peft_model does to your model

get_peft_model(model, config) does not copy your model. It walks the module tree, finds modules whose names match target_modules, and replaces each one in place with a tuner layer that keeps the original module as base_layer and adds the adapter's parameters in a ModuleDict keyed by adapter name. Then it freezes everything that is not an adapter parameter and wraps the result in a PeftModel. The object you passed in has been modified, so do not keep using it as if it were the clean base.

import torch
from transformers import AutoModelForCausalLM
from peft import LoraConfig, get_peft_model

base = AutoModelForCausalLM.from_pretrained(MODEL_ID, dtype=torch.bfloat16, device_map="auto")
config = LoraConfig(
    r=16, lora_alpha=32, lora_dropout=0.05,
    target_modules="all-linear",      # every linear layer except the output head
    task_type="CAUSAL_LM",
)
model = get_peft_model(base, config)  # base is modified in place
model.print_trainable_parameters()

print(model.base_model.model.model.layers[0].self_attn.q_proj)
# lora.Linear(
#   (base_layer): Linear(in_features=4096, out_features=4096, bias=False)
#   (lora_dropout): ModuleDict((default): Dropout(p=0.05))
#   (lora_A): ModuleDict((default): Linear(in_features=4096, out_features=16, bias=False))
#   (lora_B): ModuleDict((default): Linear(in_features=16, out_features=4096, bias=False))
#   ...)                                    # abridged

The forward pass computes the frozen base output and adds the scaled low-rank path. The scale is lora_alpha / r by default and lora_alpha / sqrt(r) with use_rslora=True. With the default init_lora_weights=True, lora_B starts at zero, so a fresh adapter is an exact no-op; the alternatives ("gaussian", "pissa", "olora", "eva", "corda", "loftq" and others) start from data- or weight-derived values instead. By default autocast_adapter_dtype=True keeps adapter weights in float32 even when the base is bfloat16, which is what you want for stable training.

Two configuration fields cover the parts LoRA cannot express. modules_to_save makes whole modules, such as a new classification head, fully trainable and saves them with the adapter. trainable_token_indices trains only selected rows of an embedding matrix, the right tool when you added a handful of special tokens and do not want to save the entire vocabulary.

Advertisement

Training with a frozen base

For a quantized base, call prepare_model_for_kbit_training(model) before get_peft_model. Its use_gradient_checkpointing argument defaults to True, and it prepares the frozen quantized model so gradients can flow to the adapters. The QLoRA recipe end to end is in QLoRA fine-tuning explained. With a non-quantized base and gradient checkpointing, the frozen input embeddings produce outputs that do not require gradients, and the backward pass can silently skip the adapters; Transformers' model.enable_input_require_grads() is the usual fix.

Any PyTorch loop or the Transformers Trainer works, because a PeftModel is an nn.Module whose only trainable parameters are the adapter's. Pass model.parameters() to the optimizer or filter on requires_grad; frozen parameters get no optimizer state either way. Check print_trainable_parameters() before the first step: if it reports far more than you expected, a modules_to_save entry or an embedding-saving setting is pulling in a large module.

What lands on disk

save_pretrained(path) writes the adapter, not the model: adapter_model.safetensors with only the adapter's tensors, adapter_config.json with the configuration, and a generated README.md model card. The config records peft_type, target_modules, rank and alpha, plus base_model_name_or_path and revision, which name the base the adapter was trained on.

Tensor keys are prefixed with base_model.model. because of the two wrapper layers, and the adapter name is stripped: in memory a parameter is ...q_proj.lora_A.default.weight, on disk it is ...q_proj.lora_A.weight. That is why one file can be loaded under any name. An adapter whose name is not default is saved in a subdirectory with that name, which surprises scripts that expect the files at the top level. selected_adapters limits which adapters are saved.

Sizes make the case for adapters. Take an illustrative model with hidden size 4096, 32 layers and square query, key, value and output projections. LoRA at rank 16 on those four matrices adds 16 x (4096 + 4096) = 131,072 parameters per matrix, 16,777,216 in total: about 34 MB in bfloat16 or 67 MB in float32, against roughly 14 GB for a 7-billion-parameter base in bfloat16. One base can therefore carry dozens of task adapters.

Loading, switching and inspecting adapters

Reloading takes the base model plus the adapter directory:

from peft import PeftModel

base = AutoModelForCausalLM.from_pretrained(MODEL_ID, revision=BASE_REVISION, dtype=torch.bfloat16)
model = PeftModel.from_pretrained(base, "adapters/support-v3")              # frozen, inference
# to resume training instead:
# model = PeftModel.from_pretrained(base, "adapters/support-v3", is_trainable=True)

model.load_adapter("adapters/billing-v1", adapter_name="billing")
model.set_adapter("billing")                 # one active adapter
with model.disable_adapter():                # temporarily the plain base model
    base_out = model.generate(**inputs, max_new_tokens=64)
model.delete_adapter("billing")
for status in model.get_layer_status()[:3]:
    print(status.name, status.active_adapters, status.merged_adapters)

is_trainable defaults to False. Resuming training without it runs happily and updates nothing, a failure that shows up only as a flat loss curve. Nothing checks that the base weights match the ones the adapter was trained on: any checkpoint with the same module names and shapes will load, so a different revision of the same architecture produces silently degraded output. Pin revision on the base and treat the adapter's base_model_name_or_path and revision as a contract you verify in your loading code.

disable_adapter() is a context manager and the cheapest A/B test you have: the same process, weights and inputs, with and without the adapter. get_layer_status() and get_model_status() report, per layer or in aggregate, which adapters are available, active and merged, and the quantization backend, which is the first thing to inspect when outputs look like the base model.

Many adapters, one base

Several adapters can serve different requests in one batch. Pass adapter_names, one entry per sample, to the forward pass or generate; the reserved name "__base__" routes a sample through the base model alone:

adapter_names = ["__base__", "support", "support", "billing"]
out = model.generate(**batch, adapter_names=adapter_names, max_new_tokens=128)

Inside each tuner layer the batch is split by adapter, so cost grows with the number of distinct adapters in the batch. Dedicated servers implement the same idea with custom kernels and paged adapter memory; see multi-LoRA serving for the production design. add_weighted_adapter builds a new adapter from existing ones, with combination types including "linear", "cat", "svd", "ties" and the dare variants; it is useful for blending, for example, an SFT and a preference adapter, and it must be evaluated like any new model.

To replace an adapter's weights without rebuilding the model, peft.utils.hotswap.hotswap_adapter(model, path, adapter_name="default") swaps them in place. With torch.compile, call prepare_model_for_compiled_hotswap(model, target_rank=max_rank) before compiling so adapters with different ranks or scalings do not trigger recompilation. The documented limits: LoRA only, and the new adapter must target the same layers as the old one or a subset of them.

Merging, and quantized bases

merge_and_unload() adds each adapter's scaled low-rank product into the base weights and returns a plain Transformers model with no PEFT wrapper, which removes the adapter's per-layer latency and lets any inference stack load the result. The price is flexibility: you cannot unmerge, switch or disable adapters afterwards, and you now ship full weights. merge_adapter() and unmerge_adapter() merge temporarily while keeping the adapter, and unload() removes adapters without merging.

Quantized bases need care. The documentation states that AQLM-quantized weights cannot absorb an adapter, that torchao merging is correct only for LoRA with int8 weight-only quantization, and that merge is not supported on Intel Neural Compressor or Transformer Engine layers. The robust practice for a 4-bit training run is to keep the adapter separate, or to reload the base in bfloat16, load the adapter onto it, merge, and quantize the merged model afterwards if you need a small artefact. Merging also rounds: the low-rank update is added to bfloat16 weights and rounded again. Compare merged and unmerged outputs on a fixed evaluation set before shipping.

Worked example: shipping a support adapter

A team ships a support-reply adapter on a 7B base. The pipeline, written as checks rather than hopes:

  1. Train with the base pinned to an exact revision, target_modules="all-linear", rank 16, alpha 32. Before step one, assert that the trainable parameter count matches the expected figure and that a fresh adapter's output equals the base output, which holds because lora_B starts at zero.
  2. Save, then reload into a fresh process with PeftModel.from_pretrained and assert that logits on ten fixed prompts match the in-memory model, which catches key and naming mistakes.
  3. Evaluate with disable_adapter() as the control: the adapter must win on the task set and not regress on a general set by more than an agreed margin.
  4. For serving, keep the adapter unmerged on the shared base, using adapter_names to mix tenants in one batch; merge only for an edge build, starting from the bfloat16 base, and re-run the evaluation on the merged model.

In this illustrative run the reload check failed once: a training script had saved under the adapter name support, so the files sat in a support/ subdirectory and the loader read an older adapter from the parent path. The fixed-prompt comparison caught it before deployment.

Failure modes

Trade-offs

Unmerged adapters keep one base in memory for many tasks and allow switching, disabling and per-request routing, at the cost of an extra low-rank matrix multiply in every targeted layer and an inference stack that understands adapters. Merged models run at base-model speed on any runtime but cost full storage per task and freeze the choice. A middle path is to serve unmerged on shared GPU fleets and merge for edge or single-tenant builds; the serving side of that choice is covered in serving LoRA adapters.

Within the configuration, target_modules="all-linear" is the safest default across architectures and usually trains better than attention-only targets, but it multiplies adapter size and per-layer overhead. Higher rank adds capacity and memory linearly. use_dora=True splits each update into magnitude and direction and often helps at low rank, but it adds work to every forward pass, supports only linear and Conv2D layers, and has narrower quantization and merge support, so check the documentation for your backend first. Hot-swapping is faster than delete-and-load and avoids recompilation, but it constrains every future adapter to the layers the first one targeted, so start with the widest target set you expect to need.

What to do next

  1. Pin the base model revision in training and serving code, and verify it against the adapter's config at load time.
  2. Assert the trainable parameter count and the fresh-adapter no-op property before the first training step.
  3. Add a save-reload-compare test on fixed prompts to every training run.
  4. Use disable_adapter() as the control arm in evaluation.
  5. Pass is_trainable=True whenever you resume training from a saved adapter.
  6. Keep adapters unmerged for multi-tenant serving; merge from a bfloat16 base for single-purpose builds and re-evaluate after merging.
  7. Read the PEFT quantization guide before merging into any quantized model.
Key takeaway: PEFT modifies your model in place, wrapping each targeted layer around its frozen original and saving only the adapter tensors and config. Most failures are lifecycle failures: adapters loaded frozen, attached to the wrong base, saved to an unexpected path or merged into quantized weights. Pin the base revision, test save and reload on fixed prompts, evaluate against disable_adapter, keep adapters unmerged for multi-tenant serving, and merge only from a full-precision base.