Favicon of torchtune

torchtune

An open-source LLM fine-tuning library for single-GPU and distributed training, with LoRA and QLoRA support. Development is no longer active.

Screenshot of torchtune website

torchtune is a Python library for developers and researchers who want to adapt LLMs on their own GPU hardware using PyTorch. Its editable training recipes suit work that needs control over the training code and model implementations. The project is no longer actively maintained.

It supports full fine-tuning and memory-saving LoRA and QLoRA methods, with supervised training that can run on one device, multiple devices or multiple machines. NVIDIA hardware examples include the RTX 4090, A6000 and A100. Memory needs depend on the model and training method; QLoRA reduces them by quantizing the base model's weights.

Beyond supervised fine-tuning, torchtune includes knowledge distillation to train smaller models from larger ones, plus preference and reinforcement learning methods such as DPO, PPO and GRPO. Quantization-aware training prepares models for reduced-precision use. Evaluation, quantization and text generation recipes cover the stages after training, though each training method has different support for distributed execution.

Model implementations include Llama, Llama Vision, Gemma, Mistral, Microsoft Phi and Qwen. The library uses PyTorch APIs directly and exposes memory and speed optimizations, including activation checkpointing and compilation. It's open source under the BSD 3-Clause license, so developers can inspect and modify the training recipes and model code.

Similar to torchtune