Hugging Face Accelerate is a thin library that lets one ordinary PyTorch training loop run on a CPU, one GPU, eight GPUs, several machines, or under FSDP or DeepSpeed without being rewritten for each. It does not own your loop the way a framework does. You keep writing the forward pass, the loss and the optimizer step yourself, and Accelerate swaps the objects you hand it for versions that know about devices, processes, precision and sharding.

That design is why Accelerate sits underneath the Hugging Face Trainer and many other libraries, and why it is worth understanding on its own. When a run hangs at step one, double-counts evaluation samples, accumulates gradients wrongly or cannot resume, the cause is almost always in the layer this article describes: the launcher, the process state, the wrapping done by prepare(), and the checkpoint machinery. We will build the architecture from first principles, walk through a complete loop, work through the numbers for data sharding and accumulation, and finish with failure modes and a checklist.

Advertisement

What Accelerate is, and what it is not

Distributed training in PyTorch needs the same chores every time: start one process per device, initialise a process group, place and wrap the model, give each process a different slice of the data, run under autocast, scale the loss for fp16, skip gradient synchronisation on accumulation steps, gather metrics and save checkpoints from one process. All of it is easy to get subtly wrong.

Accelerate packages those chores. It is not a trainer: there are no callbacks, no evaluation strategy and no hyperparameter arguments. It is not a parallelism engine either: sharding comes from PyTorch FSDP, DeepSpeed or Megatron-LM, which Accelerate configures through plugins. Its job is to make the same script correct in every one of those environments.

The architecture

There are five layers. The launcher, accelerate launch, reads a YAML config written by accelerate config (or flags such as --num_processes and --mixed_precision) and starts one Python process per device, exporting the environment variables PyTorch distributed expects. The process state, PartialState and AcceleratorState, is a per-process singleton that reads that environment once and records the rank, the local rank, the device and the distributed type. GradientState is a second singleton that tracks whether this step should synchronise gradients and whether the dataloader has reached its end.

The Accelerator object holds the run's choices: mixed precision, gradient accumulation steps, the dataloader configuration, the plugins and the project configuration for checkpoints and logging. The wrapped objects are what prepare() returns. Finally, two side systems sit beside the loop: state checkpointing and big-model loading.

Accelerate: one launcher, N processes, each running your unchanged loop through a wrapping layeraccelerate launchreads config, sets envspawns NOne process (rank r of N), same script everywherePartialStaterank, device, backendGradientStatesync_gradients, endAcceleratorprecision, plugins, configmodelDDP/FSDP/DSoptimizerAccelerateddataloadershardedschedulerAcceleratedprepare()NCCL / Gloocollectives between ranksall-reducesave_state / load_statemodel, optimizer, scheduler, RNG, scalerinit_empty_weights + dispatchbig-model inference, no launcherYour loop stays plain PyTorch; Accelerate decides what the objects you hand it really are.
The launcher starts N processes. In each, the state singletons describe the process, the Accelerator holds the run's choices, and prepare() returns wrapped versions of the model, optimizer, dataloader and scheduler.
Advertisement

What prepare() does to each object

Calling model, optimizer, train_dl, scheduler = accelerator.prepare(model, optimizer, train_dl, scheduler) is the single most important line. Each argument is recognised by type and replaced.

ObjectWhat you get backWhy it matters
ModelMoved to the process's device and wrapped in DDP, FSDP or a DeepSpeed engine, depending on config; forward runs under autocast when mixed precision is onGradient all-reduce or sharding happens without code changes; the original module is reachable with unwrap_model.
OptimizerAn AcceleratedOptimizer wrapping yoursIts step is skipped when the fp16 grad scaler finds infs, and does nothing on accumulation steps that should not update.
DataLoaderA sharded loader: each process reads a different slice, or the main process reads and dispatchesBatches land on the right device and no two ranks train on the same samples.
LR schedulerAn AcceleratedSchedulerIt only advances when the optimizer really stepped, so skipped steps do not shift the schedule.

Two rules follow. First, pass everything to one prepare() call when you can: FSDP flattens or shards parameters, so an optimizer built on the unwrapped model must be re-pointed at the wrapped parameters, and Accelerate does that when they arrive together. Second, create the optimizer after the model exists but let prepare() do device placement; do not call .cuda() yourself.

The canonical loop

Here is a complete loop with mixed precision, accumulation, clipping, evaluation and checkpointing. Everything that is not plain PyTorch is an accelerator call.

from accelerate import Accelerator
from accelerate.utils import DataLoaderConfiguration, ProjectConfiguration

accelerator = Accelerator(
    mixed_precision="bf16",
    gradient_accumulation_steps=8,
    dataloader_config=DataLoaderConfiguration(even_batches=True),
    project_config=ProjectConfiguration(project_dir="runs/exp1",
                                        automatic_checkpoint_naming=True,
                                        total_limit=3),
)
model = build_model()
optimizer = torch.optim.AdamW(model.parameters(), lr=3e-4)
scheduler = get_cosine_schedule(optimizer, warmup=500, total=20_000)
model, optimizer, train_dl, eval_dl, scheduler = accelerator.prepare(
    model, optimizer, train_dl, eval_dl, scheduler)

for epoch in range(num_epochs):
    model.train()
    for batch in train_dl:
        with accelerator.accumulate(model):
            loss = model(**batch).loss
            accelerator.backward(loss)
            if accelerator.sync_gradients:
                accelerator.clip_grad_norm_(model.parameters(), 1.0)
            optimizer.step()
            scheduler.step()
            optimizer.zero_grad()
    model.eval()
    preds, refs = [], []
    for batch in eval_dl:
        with torch.no_grad():
            logits = model(**batch).logits
        p_, r_ = accelerator.gather_for_metrics((logits.argmax(-1), batch["labels"]))
        preds.append(p_); refs.append(r_)
    if accelerator.is_main_process:
        log_accuracy(torch.cat(preds), torch.cat(refs))
    accelerator.save_state()          # runs/exp1/checkpoints/checkpoint_<i>
accelerator.end_training()

Note what is absent: no dist.init_process_group, no DistributedSampler, no GradScaler, no model.no_sync() and no rank checks around the save. accelerator.backward(loss) replaces loss.backward() because it applies loss scaling and the accumulation divisor, and for DeepSpeed it calls the engine's own backward.

Data sharding and correct evaluation

By default each process builds the same dataloader and the prepared version gives rank r every N-th batch, so with 4 processes and a per-process batch of 8 the global batch is 32. Setting split_batches=True in DataLoaderConfiguration instead treats the loader's batch size as global and splits each batch N ways; use it when you want the global batch to stay fixed as you change the GPU count. dispatch_batches makes the main process read every batch and send slices to the others, which helps for iterable datasets that cannot be indexed but puts all reading on one process.

Evaluation is where silent errors hide. Take 1,003 evaluation samples on 4 processes. With even_batches=True, the default, the sampler pads by reusing samples from the start so every process sees the same number of batches; otherwise ranks would disagree on how many collective calls to make and the run would hang. That padding means a plain accelerator.gather() returns 1,004 predictions, one sample counted twice. gather_for_metrics knows how much padding the last batch carried and drops the duplicates, so the metric is computed on exactly 1,003 samples. The difference is tiny here and large for small or class-imbalanced eval sets.

For inference over a list such as prompts, accelerator.split_between_processes(items) hands each process its share.

Gradient accumulation

Accumulation lets a small per-device batch behave like a large one. With a per-device micro-batch of 4, 8 GPUs and gradient_accumulation_steps=8, each optimizer step sees 4 x 8 x 8 = 256 samples. The arithmetic of why summing scaled micro-batch gradients equals the large-batch gradient is covered in gradient accumulation and microbatches; the systems problem is synchronisation.

Under DDP every backward pass triggers an all-reduce of all gradients. On the seven micro-batches that will not update the weights, that traffic is wasted. Inside with accelerator.accumulate(model):, Accelerate counts micro-batches, enters the model's no_sync context on the non-final ones, divides the loss by the accumulation count in backward, and sets accelerator.sync_gradients to True only on the step that synchronises. The wrapped optimizer and scheduler turn their step() into no-ops on the other steps, which is why the loop above can call them unconditionally.

Use sync_gradients for anything that should happen once per real update: gradient clipping, logging the learning rate, counting optimizer steps. When the dataloader ends mid-cycle, GradientState forces a sync on the last batch so leftover gradients are not carried into the next epoch.

Size the learning-rate schedule carefully. With the default split_batches=False, the prepared scheduler advances num_processes times per optimizer step, because Accelerate assumes the schedule was sized for a single-process run over the whole dataset. So set warmup and total steps in single-process terms: the unsharded loader length divided by the accumulation steps, times epochs. Sized in per-process steps instead, an 8-GPU run exhausts its schedule an eighth of the way through. Log the learning rate early to confirm.

Precision and backend plugins

mixed_precision accepts "no", "fp16", "bf16" and "fp8". With fp16 Accelerate creates a gradient scaler and skips optimizer steps whose gradients overflowed; bf16 needs no scaler. FP8 depends on a supported backend library and hardware, so treat it as a separate project. The trade-offs are covered in mixed precision training.

Sharding arrives through plugins. A FullyShardedDataParallelPlugin configures PyTorch FSDP: which version (fsdp_version 1 or 2), the auto-wrap policy, CPU offload and how state dicts are saved. A DeepSpeedPlugin carries a ZeRO stage and a DeepSpeed config. Both can be set in the YAML file so the same script switches backend by changing only the launch config. How each engine shards memory is covered in PyTorch FSDP and DeepSpeed. The practical rule: start with DDP while the model and optimizer state fit on one device, and move to FSDP or ZeRO only when they do not.

Launching on one machine and many

Run accelerate config once per environment, answer the questions, and commit the resulting YAML next to the script. accelerate env prints the versions and config for bug reports, and accelerate test runs a short sanity script across the configured processes.

# single node, 8 GPUs, bf16
accelerate launch --config_file configs/ddp_8gpu.yaml train.py --epochs 3

# two nodes; run on each with its own --machine_rank
accelerate launch --num_machines 2 --machine_rank 0 \
    --main_process_ip 10.0.0.11 --main_process_port 29500 \
    --num_processes 16 --mixed_precision bf16 train.py

The script itself is launched the same way everywhere, which is the point. Two notes: --num_processes is the total across all machines, and every node must reach the main process's IP and port, which is the first thing to check when a multi-node run waits forever at start-up.

Checkpoints and resume

accelerator.save_state(dir) writes the model, optimizer, scheduler, grad scaler and the random number generator state of every process, and any extra object registered with register_for_checkpointing that has state_dict and load_state_dict. With ProjectConfiguration(automatic_checkpoint_naming=True, total_limit=3) it numbers checkpoints and keeps the last three. accelerator.load_state(dir) restores all of it on the same number of processes.

Restoring state is not the same as resuming the data position. After loading, skip the batches the interrupted epoch had already consumed:

accelerator.load_state(ckpt_dir)
start_epoch, done_steps = read_progress(ckpt_dir)       # your own bookkeeping
for epoch in range(start_epoch, num_epochs):
    dl = train_dl
    if epoch == start_epoch and done_steps:
        dl = accelerator.skip_first_batches(train_dl, done_steps)
    for batch in dl:
        ...

For the final model, use accelerator.save_model(model, out_dir) or accelerator.get_state_dict(model), which gather a full state dict from FSDP or ZeRO-3 shards, rather than calling torch.save on the wrapped module. Call accelerator.wait_for_everyone() before any step that reads files another rank just wrote.

Big-model loading without a launcher

Accelerate also solves a different problem: loading a model larger than one device for inference. init_empty_weights() builds the module skeleton with no memory behind its parameters, and load_checkpoint_and_dispatch(model, checkpoint=path, device_map="auto") fills GPUs first, then CPU, then disk, loading each layer's weights straight to its device. Layer groups that must not be split go in no_split_module_classes. In Transformers the same machinery runs when you pass device_map="auto" to from_pretrained.

Understand the cost: this is sequential model parallelism. Only one device computes at a time and offloaded layers are copied in on every forward pass, so it trades speed for being able to run at all. It is started with plain python, not accelerate launch, and it is the wrong tool for training or throughput serving.

Failure modes

  • Hang at the first collective. Ranks disagree on how many collective calls to make, typically because one rank skipped a batch, ran an extra evaluation step or took a different branch. Keep control flow identical across ranks and keep even_batches on unless you handle uneven inputs explicitly.
  • Inflated evaluation metrics. Using gather where gather_for_metrics was needed counts padded samples twice.
  • Accumulation that does nothing. Calling loss.backward() instead of accelerator.backward(loss) skips scaling and the divisor; forgetting the accumulate context makes every micro-batch an update.
  • Every rank writes the checkpoint. Hand-written saves without is_main_process guards race on the same files. Use save_state, which coordinates this.
  • Unloadable FSDP weights. A torch.save(model.state_dict()) on a sharded model stores one shard per rank, or nothing useful; gather with get_state_dict.

Operating it well

Treat the launch config as code: one YAML per hardware profile, versioned with the script. Log the full accelerator.state, library versions and global batch size at step zero; flags on the command line, the YAML file and the constructor can disagree, and this is how you find out. Keep the loop backend-agnostic, then switch DDP to FSDP in the config and run both for a hundred steps to confirm the loss curves match before a long run. Test resume early by killing a short run mid-epoch and confirming the loss after resume continues from where it was, and include accelerate test in the checks for each new machine image.

The trade-off is control versus convenience. Accelerate keeps you close to PyTorch, which suits custom loops and research code. If you want evaluation strategies and early stopping handled for you, use the Trainer on top of it; if you need pipeline or tensor parallelism at scale, the engines underneath become the real interface and Accelerate is the glue that configures them.

What to do next

  1. Take an existing single-GPU loop, add an Accelerator, route model, optimizer, dataloaders and scheduler through one prepare() call, and replace loss.backward() with accelerator.backward(loss).
  2. Run accelerate config, commit the YAML, and launch on one GPU and then on all of them with the same script.
  3. Switch evaluation to gather_for_metrics and confirm the sample count equals your dataset size.
  4. Add accumulate(), move clipping and step logging under sync_gradients, and compute your global batch size explicitly.
  5. Configure ProjectConfiguration, call save_state each epoch, and rehearse a kill-and-resume with skip_first_batches.
  6. Only when memory demands it, add an FSDP or DeepSpeed plugin through the config and compare a hundred steps against the DDP baseline.
Key takeaway: Accelerate is a wrapping layer, not a framework. The launcher starts one process per device, the state singletons describe each process, and prepare() swaps your model, optimizer, dataloader and scheduler for versions that handle placement, synchronisation, precision and sharding. Most bugs come from going around that layer: calling backward directly, gathering without removing padding, accumulating without the context manager or saving sharded weights by hand. Use its calls for those jobs, keep the launch config in version control, and rehearse resume before you need it.