Unsloth is an open-source library for fine-tuning language models faster and in less GPU memory than a plain Hugging Face training loop. It keeps the familiar interface, loading a model, attaching LoRA adapters, training with TRL's trainers, saving with the usual methods, and replaces pieces of the model's forward and backward pass underneath with hand-written kernels. The project also ships a web interface, Unsloth Studio, and a desktop app; this article is about the code library, which the project calls Unsloth Core.
The project advertises training around twice as fast with roughly 70 percent less VRAM for many models, with no accuracy loss. Those are vendor figures measured against its own baselines, and the gain on your model, sequence length and GPU will differ, so the right attitude is to understand where the savings come from and measure them yourself. That is what this article sets out to do: explain where fine-tuning memory goes, what Unsloth changes, how to run and export a job end to end, and where it bites.
Where fine-tuning memory goes
Fine-tuning memory has four parts: base weights, trainable parameters with their gradients, optimizer state, and activations saved for the backward pass. Work them through for an 8-billion-parameter Llama-style model with LoRA rank 16 on all seven linear projections in each of 32 layers.
| Item | Calculation | Approximate size |
|---|---|---|
| Base weights, 16-bit | 8e9 x 2 bytes | 16 GB |
| Base weights, 4-bit NF4 | about 7e9 linear weights x 0.5 bytes + constants; embeddings and lm_head stay 16-bit (about 2 GB) | about 5.5-6 GB |
| LoRA parameters | per layer r x (in + out) summed over q,k,v,o,gate,up,down = 1.31M; x 32 | about 42M parameters |
| Adapter weights + gradients | 42M x (4 + 4) bytes in fp32 | about 0.34 GB |
| AdamW state | 42M x 8 bytes | about 0.34 GB |
| Activations | grows with batch x sequence length x layers | often the largest term |
Two conclusions drive everything Unsloth does. First, with QLoRA the trainable side is tiny: well under a gigabyte for adapters, gradients and optimizer combined. Full fine-tuning the same model would need roughly 16 bytes per parameter for weights, gradients and Adam state, around 128 GB before activations. Second, once the base is 4-bit, activations dominate, and they grow linearly with sequence length. A 16k-token sequence holds hundreds of megabytes per layer of saved tensors without checkpointing. So the remaining wins are in activation memory and in the speed of the many small operations around each matrix multiply. The background on adapters is in LoRA fine-tuning.
What Unsloth changes underneath
Triton kernels for the glue. A transformer layer spends real time outside the big matrix multiplies: rotary position embeddings, RMSNorm, the SwiGLU activation, and the cross-entropy loss over a vocabulary that can exceed 100,000 tokens. Each is a memory-bound operation that, run as separate PyTorch ops, reads and writes full tensors several times. Unsloth implements these in Triton with fused forward and backward passes, the same principle as general kernel fusion: fewer round trips to high-bandwidth memory, fewer intermediate tensors kept alive. The loss kernel matters disproportionately for large vocabularies, because materialising logits in fp32 for a long sequence is itself gigabytes.
A hand-derived LoRA backward. Autograd treats a LoRA layer as a generic graph and saves whatever its ops need. Unsloth derived the gradients of the LoRA-augmented projections and MLP by hand, which lets it choose the multiplication order, reuse buffers and skip saving tensors that can be recomputed cheaply. The documentation notes that lora_dropout = 0 and bias = "none" are the optimised paths; other values work but fall off the fast path.
Offloaded gradient checkpointing. Standard checkpointing saves only layer inputs and recomputes the rest during backward. Setting use_gradient_checkpointing = "unsloth" additionally moves those saved inputs to CPU memory and brings them back when the backward pass reaches each layer, overlapping the copies with compute. The docs credit it with about 30 percent extra memory savings and much longer trainable context. The cost is PCIe traffic and pinned host memory, so it helps most when GPU memory, not throughput, is the binding constraint.
Pre-quantized checkpoints. Unsloth publishes 4-bit versions of popular models on Hugging Face, including dynamic quantizations that keep some sensitive layers in higher precision, so loading downloads a quarter of the bytes and skips on-the-fly quantization.
An end-to-end run
A complete supervised fine-tuning run looks like this. The structure follows the project's documented arguments; replace the model and dataset names with your own.
from unsloth import FastLanguageModel # import BEFORE transformers / trl
from unsloth.chat_templates import train_on_responses_only
from trl import SFTTrainer, SFTConfig
from datasets import load_dataset
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="unsloth/Meta-Llama-3.1-8B-Instruct-bnb-4bit",
max_seq_length=4096,
load_in_4bit=True, # QLoRA; load_in_8bit / full_finetuning also exist
)
model = FastLanguageModel.get_peft_model(
model,
r=16,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
"gate_proj", "up_proj", "down_proj"],
lora_alpha=16,
lora_dropout=0, # optimised path
bias="none", # optimised path
use_gradient_checkpointing="unsloth",
random_state=3407,
use_rslora=False,
loftq_config=None,
)
ds = load_dataset("json", data_files="support_chats.jsonl", split="train")
ds = ds.map(lambda ex: {"text": tokenizer.apply_chat_template(ex["messages"], tokenize=False)})
trainer = SFTTrainer(
model=model, tokenizer=tokenizer, train_dataset=ds,
args=SFTConfig(dataset_text_field="text", per_device_train_batch_size=2,
gradient_accumulation_steps=8, learning_rate=2e-4,
num_train_epochs=1, logging_steps=10, output_dir="out", seed=3407),
)
trainer = train_on_responses_only(
trainer,
instruction_part="<|start_header_id|>user<|end_header_id|>\n\n",
response_part="<|start_header_id|>assistant<|end_header_id|>\n\n",
)
trainer.train()Three details deserve attention. The import order matters because Unsloth works by patching classes; importing it after the libraries it patches can silently leave you on the slow path. train_on_responses_only masks the loss on user turns so the model learns to answer rather than to imitate users; the two marker strings must match the model's chat template exactly, and the ones shown are for Llama 3. And max_seq_length bounds both truncation and memory planning, so set it from the length distribution of your data rather than the model's maximum.
The same model object works with TRL's preference and reinforcement trainers, which is how the project's GRPO notebooks are built; the algorithm side is covered in GRPO fine-tuning, and data preparation for the supervised case in supervised fine-tuning.
Worked example: calibrating on a 24 GB GPU
Suppose a support team fine-tunes the 8B model above on 20,000 chats with a mean of 900 tokens and a 99th percentile of 3,500, on a single 24 GB GPU. With max_seq_length=4096, a batch of 2 and accumulation of 8, the effective batch is 16 sequences. The base model takes about 6 GB, adapters and optimizer under 1 GB, leaving the rest for activations and the logits. A sensible procedure is a short calibration before the real run:
- Run 30 steps with plain Hugging Face PEFT and
use_gradient_checkpointing=True, recording peak memory withtorch.cuda.max_memory_allocated()and tokens per second. - Run the same 30 steps with Unsloth and the settings above, same seed and data order.
- Compare loss curves step by step: they should overlap closely. A visible gap means a configuration difference, often the chat template or loss masking, not a kernel effect.
- Raise the batch size until peak memory reaches about 85 percent of the card, since throughput usually rises with batch size until memory runs out.
Read the results as a budget, not a scoreboard. If the plain run already fits with headroom, Unsloth's value is mostly speed, and the question is whether a shorter run justifies another dependency. If the plain run only fits at batch 1 with a 2,048 cap, so that the long tail of chats is truncated, the memory saving changes what you can train at all: full-length examples, a larger effective batch without extra accumulation, or a bigger base model on the same card.
Padding is the other lever. Chats of 900 tokens padded to the longest sequence in a batch waste compute on pad tokens. Packing several short examples into one sequence, with attention kept within each example, removes that waste; recent Unsloth releases advertise padding-free and packing support, but confirm in the current docs how it is enabled for your trainer version, and check that packed examples cannot attend to each other. Finally, hold back 500 chats as a fixed evaluation set and score the adapter on them after training, since a matching loss curve says the kernels are correct, not that the model got better at support.
Saving and exporting
A finished run can leave in three forms, and choosing deliberately avoids most deployment surprises.
| Form | Call | Use when |
|---|---|---|
| LoRA adapter only | model.save_pretrained("lora") | Serving many adapters on one base, or keeping artefacts small |
| Merged 16-bit model | model.save_pretrained_merged("merged", tokenizer, save_method="merged_16bit") | Serving with vLLM, SGLang or any stack that wants plain weights |
| GGUF | model.save_pretrained_gguf("gguf", tokenizer, quantization_method="q4_k_m") | llama.cpp, Ollama and other local runtimes |
Merging into 16-bit weights, not into the 4-bit base, avoids compounding quantization error; you can quantize the merged model afterwards with a method chosen for inference. Saving is itself memory-hungry; the docs describe a maximum_memory_usage argument, default 0.75 of GPU memory, to lower if saving runs out of memory. Always reload the exported artefact and run a fixed evaluation prompt set through it, with the same chat template, before shipping: export is where template and tokenizer mismatches surface.
Scaling past one GPU
Unsloth's design centre is a single GPU. Its documentation describes multi-GPU training through Accelerate or DeepSpeed: for data parallelism, set ddp_find_unused_parameters=False in the training config and launch with accelerate launch train.py or torchrun --nproc_per_node N train.py. For a model too large for one card, device_map="balanced" splits layers across GPUs, which fits the model but runs the GPUs one after another rather than in parallel. Multi-GPU support has changed across releases, so check the current docs for your version, and benchmark it: for a model that fits on one card, several independent single-GPU runs for a hyperparameter sweep often use the hardware better than one distributed run.
Failure modes
- Silent slow path. Wrong import order, a non-zero dropout or an unsupported architecture leaves training on stock code. Check the startup banner and compare throughput with the calibration run.
- Template mismatch. The training template, the response markers and the serving template disagree, and the model trains fine but answers oddly in production. Use one template source for all three.
- Masking that removes everything. If the response marker never matches, every label is masked and the loss is zero or undefined. Decode one batch's unmasked labels before the real run.
- Version drift. Unsloth patches fast-moving libraries; an upgrade of transformers, TRL or PyTorch without a matching Unsloth release can break or slow training. Pin all four together.
- Host memory exhaustion. Offloaded checkpointing moves memory pressure to the CPU; long sequences on a machine with little RAM can swap or be killed.
- Comparing losses across accumulation settings. Loss normalisation under gradient accumulation has had bugs in the wider ecosystem; compare runs at equal effective batch and settings.
Trade-offs
Against a plain PEFT loop, Unsloth gives memory and speed for little code change, at the cost of an extra fast-moving dependency that patches others, and support that lags new architectures until kernels are written. Against configuration-driven frameworks built for multi-node training, it is simpler and stronger on one GPU but less mature at cluster scale. The practical choice: use it when a single GPU or a small node is the budget and the model is supported, keep a plain PEFT path that produces the same artefacts as a fallback, and decide with your own calibration numbers rather than headline figures.
What to do next
- Profile your data's token lengths and set
max_seq_lengthfrom the 99th percentile. - Run the 30-step calibration against plain PEFT and record memory, throughput and loss.
- Verify loss masking by decoding the unmasked labels of one batch.
- Pin Unsloth, transformers, TRL and PyTorch versions together in a lock file.
- Export to the form your server needs, reload it, and run a fixed evaluation set before shipping.
- Keep a fallback training path without Unsloth that produces the same artefact format.