
TRL is a Python library for developers and researchers who want to adapt foundation models on their own hardware. It builds on Hugging Face Transformers and covers supervised fine-tuning, reinforcement learning and training from preference feedback. It's open source under Apache 2.0.
The training methods serve different needs. Supervised Fine-Tuning (SFT) trains models on example responses, while Direct Preference Optimization (DPO) uses paired preferences. Kahneman-Tversky Optimization (KTO) accepts simpler desirable or undesirable feedback. TRL also supports Group Relative Policy Optimization (GRPO), reward model training and knowledge distillation, with support for different model architectures and modalities.
Hardware flexibility is a practical reason to choose it. Hugging Face Accelerate supports training on a single GPU or across multiple machines, including distributed training with DeepSpeed. PEFT integration supports LoRA and QLoRA adapters, with quantization to reduce the memory needed to train large models. Unsloth integration provides optimized kernels for faster training.
For teams already working with Transformers, TRL brings these methods into the same model ecosystem. Developers can train language models or PEFT adapters on custom datasets through its Python trainers. A command-line interface also supports fine-tuning without writing code.
Claim this page and we'll verify you by hand. TRL gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find TRL?Promote it
Something wrong or outdated on this page?
21.7KUpdated 1 day agoApache-2.0
macOS#Distributed execution#Hugging Face integration#LoRA
PEFT is an open-source Python library for developers who want to adapt pretrained models on their own hardware with less compute and storage than full fine-tuning requires. It trains a small subset of parameters, often through adapters, while leaving the base model intact. It's licensed under Apache 2.0.
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
11.8KUpdated 4 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
10.1KUpdated 2 weeks agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
9.4KUpdated 2 days agoApache-2.0
Docker#Distributed execution#LoRA#Multimodal input
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
Ludwig is an open-source Python framework for developers and researchers who want to train custom AI models on their own hardware. A YAML file describes the model and training pipeline, while Ludwig handles preprocessing, training and evaluation. It uses the Apache 2.0 license. Install the Python package with the optional LLM dependencies for fine-tuning; current source requires Python 3.12 or later.
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
OpenRLHF is a self-hosted Python framework for researchers and teams training language models with human feedback or custom rewards. It runs on your own NVIDIA GPU hardware, with Docker support and distributed training across servers. It's open source under Apache 2.0.
Oumi builds specialized AI models for teams that want control over their training data, model weights, and deployment. Its Apache 2.0 open-source stack runs on laptops, clusters, and your own servers, while its hosted service automates model development from a plain-English task description. You own the resulting weights, data, and training recipes.