Favicon of TRL

TRL

Open-source Python library for LLM fine-tuning on your own GPUs, with Transformers support, preference training and Apache 2.0 licensing.

Screenshot of TRL website

TRL is a Python library for developers and researchers who want to adapt foundation models on their own hardware. It builds on Hugging Face Transformers and covers supervised fine-tuning, reinforcement learning and training from preference feedback. It's open source under Apache 2.0.

The training methods serve different needs. Supervised Fine-Tuning (SFT) trains models on example responses, while Direct Preference Optimization (DPO) uses paired preferences. Kahneman-Tversky Optimization (KTO) accepts simpler desirable or undesirable feedback. TRL also supports Group Relative Policy Optimization (GRPO), reward model training and knowledge distillation, with support for different model architectures and modalities.

Hardware flexibility is a practical reason to choose it. Hugging Face Accelerate supports training on a single GPU or across multiple machines, including distributed training with DeepSpeed. PEFT integration supports LoRA and QLoRA adapters, with quantization to reduce the memory needed to train large models. Unsloth integration provides optimized kernels for faster training.

For teams already working with Transformers, TRL brings these methods into the same model ecosystem. Developers can train language models or PEFT adapters on custom datasets through its Python trainers. A command-line interface also supports fine-tuning without writing code.

Similar to TRL