Player not loading? Watch on YouTube
DeepSpeed's ZeRO optimizer is the focus of this explanation of memory use in distributed LLM training. The speaker calculates 208GB of training state for a 13-billion-parameter model using mixed-precision Adam, before activations or the dataset. That figure describes the stated recipe rather than every way to train a model.
The video explains ZeRO's stages: partition optimizer state, then gradients, then model weights across GPUs. It cites Microsoft's 4x and 8x memory reductions for the first two stages, without demonstrating those reductions on a two-card system. Stage three gathers the weights needed for each layer and increases communication between GPUs.
CPU offload moves optimizer state into system RAM, while ZeRO-Infinity extends offloading to NVMe storage. The speaker proposes two RTX 3090s for a 13B full fine-tune, but does not demonstrate a completed run or establish the complete memory requirements of that configuration. Transfers over PCIe can slow training.
The integration discussion covers a JSON configuration and a changed launch command for PyTorch, citing a Hugging Face example with a 3B T5 model. The closing comparison considers renting an A100 and using LoRA, which trains a small adapter instead of every model weight.