Favicon of torchtitan

torchtitan

Open-source LLM training software for your own GPU servers, with PyTorch distributed training, NVIDIA and AMD support, and a BSD-3-Clause license.

torchtitan is an open-source training platform for researchers and developers building generative AI models on their own GPU machines or server clusters. It uses PyTorch's distributed training tools and keeps the model code relatively simple when spreading work across GPUs. The Python codebase has extension points and replaceable components for experiments with model architectures and training infrastructure.

Supported model families include Llama 3, Qwen3, DeepSeek V3 and V4, GPT-OSS, Kimi, Muse Glimmer, and Flux. It supports NVIDIA GPUs and AMD GPUs through ROCm, with multi-node training on Slurm clusters. The code uses the BSD-3-Clause license.

Training capabilities include composable data, tensor, pipeline, and context parallelism, plus torch.compile and low-precision training with Float8, MXFP8, and NVFP4. Activation checkpointing and BF16 optimizer states help reduce memory use. For fine-tuning, it accepts chat-formatted datasets; pretraining can use C4 or custom datasets.

Distributed checkpoints support saving asynchronously, and torchtune can load them for further fine-tuning. Checkpoint conversion connects Hugging Face and PyTorch distributed formats, while vLLM can run inference with torchtitan models.

CPU and GPU profiling tools help investigate performance and memory use. Training metrics include loss, GPU memory, and throughput, with logging through TensorBoard or Weights & Biases.

The project is under extensive development. Its latest features require recent PyTorch nightly builds, so check the documented version requirements before setting up a training environment.

Similar to torchtitan