Favicon of Accelerate

Accelerate

A PyTorch library for training and inference on local machines or clusters, with CPU, GPU and TPU support. Open source under Apache 2.0.

Screenshot of Accelerate website

Accelerate is a Python library for developers and researchers who write their own PyTorch training loops and want to use the same code on a local machine or a distributed cluster. It handles the hardware-specific work while leaving the training logic under your control.

Its main appeal is portability. A script used for local debugging can also run across multiple GPUs or machines without a separate training implementation. CPU-only execution and TPU support make it useful beyond GPU workstations. Accelerate supports both training and inference, and it's open source under the Apache 2.0 license.

The library handles device placement and supports mixed precision with FP16, BF16 and FP8. For larger training workloads, it integrates with DeepSpeed and PyTorch Fully Sharded Data Parallel (FSDP), so existing code can use distributed training methods without replacing the training loop. FP8 support works through Transformer Engine or MS-AMP.

Accelerate is a thin layer around PyTorch, built on torch.distributed and torch_xla. That makes it a fit for people who want help with distributed execution but still need to choose how their model trains. You'll still write the loop yourself.

An optional command-line launcher handles execution across supported hardware. Distributed training can also start from a notebook, including Colab or Kaggle notebooks with a TPU backend.

Similar to Accelerate