Axolotl is an open-source framework that fine-tunes large language models from a single YAML file. It mostly assembles existing libraries rather than reimplementing training, though it adds optional fused LoRA kernels: Hugging Face Transformers, PEFT for LoRA adapters, bitsandbytes for quantized weights, FlashAttention and PyTorch FSDP or DeepSpeed for sharding, and it adds the parts that are tedious to write well: chat templating, label masking, sample packing, dataset caching and sensible defaults. The result is that a run which would take several hundred lines of custom Python is a reviewed config and one command.
That convenience has a cost: every key is a decision about GPU memory, throughput or what the model learns, and a wrong key fails quietly. This page explains what each block of the config becomes at run time, so you can read a config the way you read code. The generic operational runbook for fine-tuning, including checkpoint policy and eval gates, lives in Fine-Tuning Ops on GPU, in depth; this page is about Axolotl itself. Key names were checked against the Axolotl documentation and repository examples on 2026-10-02; Axolotl moves quickly, so pin a version and check its config reference before copying anything.
Config as program
The design idea is that the YAML file is the program. There is no training script to drift from the config, so the file is a complete record of the run: base model, data, template, adapter, batch shape, optimizer, schedule, precision and distribution strategy. The same file is passed to preprocessing, training, inference and LoRA merging, and each stage reads the keys it needs.
This has practical consequences. Configs can be diffed between two runs to explain a difference in loss. They can be reviewed in a pull request before GPU hours are spent. And they can be generated: a sweep is a loop that writes configs with different learning rates or ranks. The trade-off is that anything Axolotl does not expose as a key needs a plugin or a fork, and the defaults you did not write are part of your program too. Print the fully resolved config at the start of every run and keep it with the output.
A complete QLoRA config, read block by block
The config below fine-tunes an 8B instruction model on support conversations with QLoRA. It follows the shape of the repository example for Llama 3 QLoRA, with a chat dataset instead of an Alpaca one.
base_model: meta-llama/Llama-3.1-8B-Instruct # any Hugging Face causal LM you are licensed to use
load_in_4bit: true # bitsandbytes 4-bit base weights
adapter: qlora # train LoRA matrices on top of the frozen quantized base
lora_r: 32
lora_alpha: 16
lora_dropout: 0.05
lora_target_linear: true # every linear layer: q, k, v, o, gate, up, down
datasets:
- path: ./data/support_chats.jsonl
type: chat_template
field_messages: messages
message_property_mappings:
role: role
content: content
roles_to_train: ["assistant"]
train_on_eos: turn
chat_template: tokenizer_default
dataset_prepared_path: ./prepared/support_v3
val_set_size: 0.05
output_dir: ./outputs/support-qlora-v3
sequence_len: 4096
sample_packing: true
eval_sample_packing: true
attn_implementation: flash_attention_2
micro_batch_size: 2
gradient_accumulation_steps: 8
num_epochs: 2
optimizer: adamw_bnb_8bit
lr_scheduler: cosine
learning_rate: 0.0002
warmup_ratio: 0.05
bf16: auto
gradient_checkpointing: true
evals_per_epoch: 4
saves_per_epoch: 2
logging_steps: 10
loss_watchdog_threshold: 5.0
loss_watchdog_patience: 3The model block loads the base weights in 4-bit and attaches LoRA matrices of rank 32 to every linear layer. Only the LoRA matrices train; the quantized base stays frozen. The dataset block tells Axolotl where messages live in each record and which roles produce loss. The sequence block sets the maximum packed length and the attention backend; attn_implementation: flash_attention_2 replaces the older flash_attention: true, which is now documented as deprecated. The trainer block sets the batch shape and optimizer; adamw_bnb_8bit keeps optimizer state in 8 bits. The watchdog keys stop a run whose loss exceeds the threshold for the given number of steps, which catches divergence before it burns a night.
Datasets, templates and what is actually trained
The most consequential block is the dataset block, because it decides which tokens contribute to the loss. With type: chat_template, Axolotl renders each conversation with a Jinja chat template, tokenizes it, and sets the label of every token that should not be learned to -100, the value PyTorch cross-entropy ignores. roles_to_train chooses which turns are learned, normally only the assistant. train_on_eos: turn also trains the end-of-turn token after each assistant turn, which teaches the model to stop; forgetting it is a common cause of models that ramble past their answer.
chat_template: tokenizer_default uses the template shipped with the tokenizer. That is usually right for instruction models, because the model you deploy will be served with the same template. Choosing a different template, such as chatml on a model that was trained with its own format, is legal but means you must serve with that template forever. Older formats such as type: alpaca remain available for instruction and response pairs.
Never trust masking you have not looked at. Run preprocessing alone, then decode one example twice, once in full and once with only the trained tokens:
# After `axolotl preprocess config.yml`: check what the model will actually be trained on.
from datasets import load_from_disk
from transformers import AutoTokenizer
import glob
tok = AutoTokenizer.from_pretrained("meta-llama/Llama-3.1-8B-Instruct")
path = sorted(glob.glob("./prepared/support_v3/*"))[-1] # hashed subdirectory
ds = load_from_disk(path)
row = ds[0]
ids, labels = row["input_ids"], row["labels"]
trained = [t for t, l in zip(ids, labels) if l != -100]
print("tokens:", len(ids), "trained:", len(trained), f"({len(trained) / len(ids):.0%})")
print("--- full text ---")
print(tok.decode(ids)[:1500])
print("--- trained text only ---")
print(tok.decode(trained)[:800])If the trained text contains the system prompt or user turns, the role mapping is wrong. If it is empty, the role names in your data do not match the template. If the trained fraction is very low, most compute is spent on context; that is fine for long-context tasks but worth knowing when you estimate tokens per step.
The prepared-dataset cache
axolotl preprocess writes the templated, tokenized and masked dataset under dataset_prepared_path, in a subdirectory named by a hash of the dependent settings; packing into batches happens later, in training. Training reuses it, which saves minutes to hours on large datasets and makes multi-GPU starts fast because ranks do not tokenize in parallel. Two habits avoid trouble. Give each data version its own prepared path, as in the example, so a cache built from last week's file is never silently reused. And when you change the tokenizer, template or masking settings, delete the cache rather than relying on the hash to notice; it is cheaper to rebuild than to train on stale tokens.
Batch arithmetic and sample packing
Three keys define the batch. micro_batch_size is the number of sequences per GPU per forward pass, gradient_accumulation_steps is how many micro-batches are summed before an optimizer step, and the data-parallel world size is the number of GPUs holding model replicas or shards. The effective batch is their product: 2 times 8 times 1 GPU gives 16 sequences per optimizer step. Doubling GPUs doubles the effective batch unless you halve accumulation, so learning rate and step counts change when hardware changes. Keep the effective batch fixed across hardware and adjust accumulation.
Without packing, each sequence is padded to the longest in the batch, and chat data is mostly short: a dataset with a median of 600 tokens at a 4,096 limit wastes most of each batch on padding. sample_packing: true concatenates several examples into one sequence up to sequence_len. With FlashAttention, Axolotl passes the boundaries between examples to the variable-length attention kernel, so tokens attend only within their own example and the packed row behaves like several independent rows. Without FlashAttention it falls back to 4D attention masks, which the documentation notes packs less efficiently. With packing on, the micro-batch counts packed rows, so tokens per optimizer step are roughly micro-batch times accumulation times GPUs times sequence length times packing efficiency: about 2 times 8 times 4,096 times 0.9, or 59,000 tokens, in the example.
If evaluation fails with packing, set eval_sample_packing: false, as the config reference suggests; it costs eval speed, not correctness. Packing also changes the number of steps per epoch, so set warmup as a ratio rather than a fixed step count.
Memory budget: QLoRA on an 8B model
It helps to know where memory goes before choosing ranks and lengths. For a Llama-style 8B model with hidden size 4,096, MLP size 14,336, 32 layers and grouped key-value heads projecting to 1,024, the parts are as follows.
| Component | Arithmetic | Approximate size |
|---|---|---|
| Base weights, 4-bit | 8B parameters at about 0.5 bytes plus quantization constants | 5 to 6 GB |
| LoRA parameters, r = 32, all linear | 32 x (8,192 + 5,120 + 5,120 + 8,192 + 3 x 18,432) x 32 layers | about 84 M params, 0.17 GB in bf16 |
| LoRA gradients | same count in bf16 | 0.17 GB |
| Optimizer state, 8-bit Adam | two states at about 1 byte each | about 0.17 GB |
| Activations | micro-batch 2 x 4,096 tokens, gradient checkpointing on | several GB, the part you tune |
| CUDA context, allocator, kernels | fixed overhead | 1 to 2 GB |
The adapter and its optimizer are small; activations dominate. That is why the levers that move memory are sequence length, micro-batch size and gradient checkpointing, not LoRA rank. The configuration above fits comfortably on a 24 GB GPU and leaves room on a 48 GB one to double the micro-batch. Full fine-tuning of the same model is a different budget: bf16 weights, gradients and fp32 Adam state need roughly 16 bytes per parameter before activations, so it needs sharding across several GPUs, covered next. Precision choices are explained in Mixed Precision Training, in depth.
Multi-GPU: FSDP, DeepSpeed and context parallelism
Axolotl uses the same axolotl train config.yml command on one GPU or many, and the distribution strategy is a config block. With no sharding keys, extra GPUs run plain data parallelism, which only helps when the whole model fits on each GPU. To shard, add an FSDP block or point at a DeepSpeed ZeRO config:
fsdp_version: 2
fsdp_config:
offload_params: false
cpu_ram_efficient_loading: true
auto_wrap_policy: TRANSFORMER_BASED_WRAP
transformer_layer_cls_to_wrap: LlamaDecoderLayer
state_dict_type: FULL_STATE_DICT
reshard_after_forward: true
# or, instead of the fsdp keys:
# deepspeed: deepspeed_configs/zero2.jsonFSDP, which the Axolotl documentation recommends, shards parameters, gradients and optimizer state across ranks and gathers each wrapped decoder layer just before it runs; the wrap class must match the model's decoder layer name. cpu_ram_efficient_loading loads weights on one rank and broadcasts, which avoids every rank reading the full checkpoint into host memory. DeepSpeed ZeRO stages 1 to 3 trade the same way, with stage 3 sharding parameters too. The mechanics behind both are in PyTorch FSDP, in depth and DeepSpeed, in depth. For single sequences too long to fit on one GPU even with checkpointing, context_parallel_size splits a sequence across GPUs.
Worked run, end to end
axolotl fetch examples # reference configs to start from
axolotl fetch deepspeed_configs # zero1/zero2/zero3 json files, if you use DeepSpeed
axolotl preprocess config.yml # template, tokenize, mask -> dataset_prepared_path
axolotl train config.yml # same command on one GPU or several
axolotl inference config.yml --lora-model-dir="./outputs/support-qlora-v3"
axolotl merge-lora config.yml --lora-model-dir="./outputs/support-qlora-v3"A disciplined run follows the commands in order. Fetch an example close to your model and start from it. Preprocess and inspect masking with the script above. Train with save_first_step or a short max_steps smoke test first if the config is new, to prove checkpoints save and resume. Watch training loss, eval loss and tokens per second; eval loss rising while training loss falls is the usual sign of overfitting on small datasets, and two epochs is often already enough. Then run interactive inference with the adapter, compare against the base model on held-out prompts, and only then merge the adapter into full weights if your serving stack needs merged weights. Keep the adapter and config even after merging; they are the reproducible artefact.
Failure modes
- Template mismatch between training and serving: the model answers well in Axolotl inference and badly in production. Serve with exactly the template used in training.
- Missing pad or end-of-turn tokens: base models often have no pad token. Set
special_tokensas the examples do, and train on end of turn. - Loss near zero from the first steps: labels include the prompt or the data contains duplicates. Inspect the trained text.
- Out of memory only at evaluation: eval batches can be larger or unpacked. Reduce eval batch size or disable eval packing.
- Stale prepared data after changing the source file under the same path. Version the prepared path.
- Mismatched wrap class or sharding keys: FSDP falls back to wrapping too little and memory looks like plain data parallelism. Confirm per-GPU memory drops as GPUs are added.
- Unpinned versions: the same config gives a different result after an upgrade because a default changed. Record the Axolotl, Transformers and PyTorch versions with the run.
What to do next
- Pin an Axolotl version, fetch its examples, and start from the example closest to your model and adapter type.
- Convert your data to a messages format, set
type: chat_templatewithroles_to_trainandtrain_on_eos, and inspect the decoded trained tokens. - Fix the effective batch, compute tokens per step with packing, and set warmup as a ratio.
- Run a short smoke test that saves and resumes a checkpoint before the real run.
- Use FSDP or DeepSpeed keys only when the model does not fit; check that per-GPU memory falls as you add GPUs.
- Evaluate the adapter against the base on held-out prompts with the serving template, then merge if needed.
- Store the resolved config, versions, adapter and eval results together, and read the FlashAttention guide to understand why packing is cheap.