DeepSpeed ZeRO: GPU memory sharding and CPU offload

Learn how DeepSpeed ZeRO partitions optimizer state, gradients and weights, and why CPU offload trades GPU memory for slower transfers.

Player not loading? Watch on YouTube

DeepSpeed's ZeRO optimizer is the focus of this explanation of memory use in distributed LLM training. The speaker calculates 208GB of training state for a 13-billion-parameter model using mixed-precision Adam, before activations or the dataset. That figure describes the stated recipe rather than every way to train a model.

The video explains ZeRO's stages: partition optimizer state, then gradients, then model weights across GPUs. It cites Microsoft's 4x and 8x memory reductions for the first two stages, without demonstrating those reductions on a two-card system. Stage three gathers the weights needed for each layer and increases communication between GPUs.

CPU offload moves optimizer state into system RAM, while ZeRO-Infinity extends offloading to NVMe storage. The speaker proposes two RTX 3090s for a 13B full fine-tune, but does not demonstrate a completed run or establish the complete memory requirements of that configuration. Transfers over PCIe can slow training.

The integration discussion covers a JSON configuration and a changed launch command for PyTorch, citing a Hugging Face example with a 3B T5 model. The closing comparison considers renting an A100 and using LoRA, which trains a small adapter instead of every model weight.