Favicon of PyTorch Lightning

PyTorch Lightning

Open-source PyTorch training framework for pretraining and fine-tuning models on your own CPUs or GPUs, with Apache 2.0 licensing.

Screenshot of PyTorch Lightning website

PyTorch Lightning is a Python framework for researchers and developers who want to pretrain or fine-tune models on their own hardware or in the cloud. It handles repetitive training code while leaving model logic under your control. The framework is open source under Apache 2.0.

Its main value is keeping a PyTorch project usable as training grows from a CPU to multiple GPUs or machines. Lightning automates backpropagation, mixed precision and distributed training without requiring changes to your core model code. Your models remain PyTorch modules, so you can keep working with PyTorch rather than adopt a separate model format.

Training features include checkpointing, early stopping and experiment manager integrations. It also supports exporting models to ONNX and TorchScript for production use. The framework covers tasks such as classification, segmentation and summarization, rather than focusing only on language models.

For teams that need more control over training loops, the companion Lightning Fabric package provides lower-level tools. Fabric supports Apple Silicon and CUDA GPUs, CPUs and TPUs, plus distributed training strategies such as DDP, FSDP and DeepSpeed.

You can run the framework on your own hardware without choosing Lightning's hosted service. Lightning Cloud is a separate offering that runs workloads on cloud GPUs and provides infrastructure management, autoscaling and monitoring. The hosted platform also supports private cloud and VPC deployments, Kubernetes and Slurm.

Similar to PyTorch Lightning