
DeepSpeed is an open-source library for developers and researchers training or running large AI models on their own hardware or compute clusters. It works with PyTorch and focuses on memory use, training speed, and distributing work across GPUs. It's licensed under Apache 2.0.
Its main attraction is making large models more practical within available hardware. ZeRO reduces training memory requirements, while ZeRO-Offload and ZeRO-Infinity support moving work and model state beyond GPU memory. Data, model, and pipeline parallelism let teams spread training across multiple devices.
The library also covers workloads with different demands. DeepSpeed-MoE supports mixture-of-experts training and inference, and Ulysses Sequence Parallelism targets long sequences. DeepSpeed also includes inference optimizations for running large models.
DeepSpeed integrates with Hugging Face Transformers and Accelerate, as well as Lightning and MosaicML. That makes it relevant to teams that want to retain their existing training framework while adding distributed execution and memory optimizations. Models trained with it include BLOOM and GPT-NeoX.
Hardware support includes NVIDIA and AMD GPUs, with CUDA and ROCm used for GPU extensions. Windows supports many training and inference features, though asynchronous I/O and GPU Direct Storage aren't supported there.
Claim this page with an email at deepspeed.ai. DeepSpeed gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find DeepSpeed?Promote it
Something wrong or outdated on this page?
9.9KUpdated 1 day agoApache-2.0
#Distributed execution
Accelerate is a Python library for developers and researchers who write their own PyTorch training loops and want to use the same code on a local machine or a distributed cluster. It handles the hardware-specific work while leaving the training logic under your control.
41.4KUpdated 3 days agoApache-2.0
#Distributed execution
Colossal-AI is a Python framework for developers and researchers training or serving large AI models on their own GPU hardware. It addresses the memory and computing demands of models that are difficult to fit on a single GPU, with tools for distributing work across a cluster. It's open source under Apache 2.0.
21.9KUpdated 1 day agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
28.6KUpdated 23 hours agoMIT
macOS · Linux#Distributed execution#LoRA
14.7KUpdated 21 hours ago
Docker#Batch processing#Distributed execution#LoRA
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
MLX is a machine learning array framework for researchers and developers building models on their own hardware. Its distinctive feature on Apple silicon is shared CPU and GPU memory: both processors can work on the same arrays without copying data between them. It's open source under the MIT license.
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.