
verl is a Python library for teams training large language models on their own GPU infrastructure. It's the open-source implementation of HybridFlow, aimed at researchers and engineers who need reinforcement learning after initial model training. It uses the Apache 2.0 license.
The library supports PPO, GRPO and other reinforcement learning methods, as well as supervised fine-tuning. Rewards can come from another model or from functions that check answers, such as math solutions or code. It also supports multi-turn interactions with tool calling and multimodal training with vision-language models including Qwen2.5-vl and Kimi-VL.
Training and response generation can use different backends. FSDP, FSDP2 and Megatron-LM handle training; vLLM, SGLang and Hugging Face Transformers generate responses. Compatible model families include Qwen, Llama3.1, Gemma2 and DeepSeek-LLM, with support for Hugging Face Transformers and Modelscope Hub.
Its HybridFlow approach lets teams extend training workflows and place models across different groups of GPUs. The 3D-HybridEngine reduces duplicate memory use and communication overhead when switching between training and generation. Multi-GPU LoRA training helps reduce memory needs, and expert parallelism supports large models across GPU clusters.
Hardware support covers NVIDIA, AMD and Ascend. On AMD ROCm GPUs, verl supports FSDP, FSDP2 and Megatron-LM training with vLLM for generation. Experiments can report to wandb, swanlab, MLflow or TensorBoard.
Claim this page and we'll verify you by hand. verl gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find verl?Promote it
Something wrong or outdated on this page?
12.2KUpdated 3 days agoMIT
macOS · Windows · Linux · Web#Hugging Face integration#Image-to-image#LoRA
AI Toolkit (ostris) is an MIT-licensed training suite for people who want to fine-tune image and video models on their own hardware or a self-hosted server. It targets consumer NVIDIA GPUs and runs on Linux and Windows, including ARM64 Linux systems such as DGX Spark. An experimental installer also supports Apple Silicon Macs. GPU memory needs depend on the model and training task.
12.5KUpdated 2 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
75.2KUpdated 2 days agoApache-2.0
Web#LoRA#Multimodal input#OpenAI-compatible API
11.8KUpdated 4 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
15.8KUpdated 2 days agoApache-2.0
Web#Distributed execution#Hugging Face integration#LoRA
10.1KUpdated 2 weeks agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
Axolotl is an open-source LLM fine-tuning framework for developers, researchers, and teams training models on their own data. It runs on local hardware or cloud infrastructure you control, including Docker and Kubernetes environments. The framework uses Apache 2.0, which permits commercial use.
LLaMA-Factory is an open-source framework for developers and researchers who want to adapt language and multimodal models on their own hardware. It brings training and inference into one toolkit, with support for LLaMA, Qwen3, Qwen3-VL, DeepSeek, Gemma, Mistral and LLaVA. Its license is Apache 2.0.
Ludwig is an open-source Python framework for developers and researchers who want to train custom AI models on their own hardware. A YAML file describes the model and training pipeline, while Ludwig handles preprocessing, training and evaluation. It uses the Apache 2.0 license. Install the Python package with the optional LLM dependencies for fine-tuning; current source requires Python 3.12 or later.
ms-swift is a Python framework for developers and researchers who want to train and deploy language or multimodal models on their own hardware. It brings fine-tuning, evaluation and model serving into one project, with support for Qwen3, DeepSeek-R1, Llama4 and Mistral, plus multimodal models such as Qwen3-VL and InternVL3.5. It's open source under Apache 2.0.
OpenRLHF is a self-hosted Python framework for researchers and teams training language models with human feedback or custom rewards. It runs on your own NVIDIA GPU hardware, with Docker support and distributed training across servers. It's open source under Apache 2.0.