
Accelerate is a Python library for developers and researchers who write their own PyTorch training loops and want to use the same code on a local machine or a distributed cluster. It handles the hardware-specific work while leaving the training logic under your control.
Its main appeal is portability. A script used for local debugging can also run across multiple GPUs or machines without a separate training implementation. CPU-only execution and TPU support make it useful beyond GPU workstations. Accelerate supports both training and inference, and it's open source under the Apache 2.0 license.
The library handles device placement and supports mixed precision with FP16, BF16 and FP8. For larger training workloads, it integrates with DeepSpeed and PyTorch Fully Sharded Data Parallel (FSDP), so existing code can use distributed training methods without replacing the training loop. FP8 support works through Transformer Engine or MS-AMP.
Accelerate is a thin layer around PyTorch, built on torch.distributed and torch_xla. That makes it a fit for people who want help with distributed execution but still need to choose how their model trains. You'll still write the loop yourself.
An optional command-line launcher handles execution across supported hardware. Distributed training can also start from a notebook, including Colab or Kaggle notebooks with a TPU backend.
Claim this page and we'll verify you by hand. Accelerate gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Accelerate?Promote it
Something wrong or outdated on this page?
41.4KUpdated 3 days agoApache-2.0
#Distributed execution
Colossal-AI is a Python framework for developers and researchers training or serving large AI models on their own GPU hardware. It addresses the memory and computing demands of models that are difficult to fit on a single GPU, with tools for distributing work across a cluster. It's open source under Apache 2.0.
43.2KUpdated 1 day agoApache-2.0
Windows#Distributed execution
28.6KUpdated 23 hours agoMIT
macOS · Linux#Distributed execution#LoRA
21.9KUpdated 1 day agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
14.7KUpdated 21 hours ago
Docker#Batch processing#Distributed execution#LoRA
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
DeepSpeed is an open-source library for developers and researchers training or running large AI models on their own hardware or compute clusters. It works with PyTorch and focuses on memory use, training speed, and distributing work across GPUs. It's licensed under Apache 2.0.
MLX is a machine learning array framework for researchers and developers building models on their own hardware. Its distinctive feature on Apple silicon is shared CPU and GPU memory: both processors can work on the same arrays without copying data between them. It's open source under the MIT license.
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.