torchtitan is an open-source training platform for researchers and developers building generative AI models on their own GPU machines or server clusters. It uses PyTorch's distributed training tools and keeps the model code relatively simple when spreading work across GPUs. The Python codebase has extension points and replaceable components for experiments with model architectures and training infrastructure.
Supported model families include Llama 3, Qwen3, DeepSeek V3 and V4, GPT-OSS, Kimi, Muse Glimmer, and Flux. It supports NVIDIA GPUs and AMD GPUs through ROCm, with multi-node training on Slurm clusters. The code uses the BSD-3-Clause license.
Training capabilities include composable data, tensor, pipeline, and context parallelism, plus torch.compile and low-precision training with Float8, MXFP8, and NVFP4. Activation checkpointing and BF16 optimizer states help reduce memory use. For fine-tuning, it accepts chat-formatted datasets; pretraining can use C4 or custom datasets.
Distributed checkpoints support saving asynchronously, and torchtune can load them for further fine-tuning. Checkpoint conversion connects Hugging Face and PyTorch distributed formats, while vLLM can run inference with torchtitan models.
CPU and GPU profiling tools help investigate performance and memory use. Training metrics include loss, GPU memory, and throughput, with logging through TensorBoard or Weights & Biases.
The project is under extensive development. Its latest features require recent PyTorch nightly builds, so check the documented version requirements before setting up a training environment.
Claim this page and we'll verify you by hand. torchtitan gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find torchtitan?Promote it
Something wrong or outdated on this page?
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
7.4KUpdated 12 hours agoApache-2.0
Linux#Agent Skills#OpenAI-compatible API#Prompt versioning
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
11.8KUpdated 4 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
10.1KUpdated 2 weeks agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
Reef is self-hosted infrastructure for developers who want AI agents to improve through feedback on actual interactions. It connects inference and learning with versioned deployment, so an agent can update its model weights or its prompts, rules, and skills while continuing to serve requests. It's open source under Apache 2.0.
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
Ludwig is an open-source Python framework for developers and researchers who want to train custom AI models on their own hardware. A YAML file describes the model and training pipeline, while Ludwig handles preprocessing, training and evaluation. It uses the Apache 2.0 license. Install the Python package with the optional LLM dependencies for fine-tuning; current source requires Python 3.12 or later.
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
OpenRLHF is a self-hosted Python framework for researchers and teams training language models with human feedback or custom rewards. It runs on your own NVIDIA GPU hardware, with Docker support and distributed training across servers. It's open source under Apache 2.0.