Favicon of DeepSpeed

DeepSpeed

Open-source PyTorch optimization library for distributed training and inference, with Apache 2.0 licensing and support for NVIDIA and AMD GPUs.

Screenshot of DeepSpeed website

DeepSpeed is an open-source library for developers and researchers training or running large AI models on their own hardware or compute clusters. It works with PyTorch and focuses on memory use, training speed, and distributing work across GPUs. It's licensed under Apache 2.0.

Its main attraction is making large models more practical within available hardware. ZeRO reduces training memory requirements, while ZeRO-Offload and ZeRO-Infinity support moving work and model state beyond GPU memory. Data, model, and pipeline parallelism let teams spread training across multiple devices.

The library also covers workloads with different demands. DeepSpeed-MoE supports mixture-of-experts training and inference, and Ulysses Sequence Parallelism targets long sequences. DeepSpeed also includes inference optimizations for running large models.

DeepSpeed integrates with Hugging Face Transformers and Accelerate, as well as Lightning and MosaicML. That makes it relevant to teams that want to retain their existing training framework while adding distributed execution and memory optimizations. Models trained with it include BLOOM and GPT-NeoX.

Hardware support includes NVIDIA and AMD GPUs, with CUDA and ROCm used for GPU extensions. Windows supports many training and inference features, though asynchronous I/O and GPU Direct Storage aren't supported there.

Similar to DeepSpeed