Favicon of verl

verl

Open-source LLM training library for reinforcement learning and fine-tuning on NVIDIA, AMD and Ascend hardware, licensed under Apache 2.0.

Screenshot of verl website

verl is a Python library for teams training large language models on their own GPU infrastructure. It's the open-source implementation of HybridFlow, aimed at researchers and engineers who need reinforcement learning after initial model training. It uses the Apache 2.0 license.

The library supports PPO, GRPO and other reinforcement learning methods, as well as supervised fine-tuning. Rewards can come from another model or from functions that check answers, such as math solutions or code. It also supports multi-turn interactions with tool calling and multimodal training with vision-language models including Qwen2.5-vl and Kimi-VL.

Training and response generation can use different backends. FSDP, FSDP2 and Megatron-LM handle training; vLLM, SGLang and Hugging Face Transformers generate responses. Compatible model families include Qwen, Llama3.1, Gemma2 and DeepSeek-LLM, with support for Hugging Face Transformers and Modelscope Hub.

Its HybridFlow approach lets teams extend training workflows and place models across different groups of GPUs. The 3D-HybridEngine reduces duplicate memory use and communication overhead when switching between training and generation. Multi-GPU LoRA training helps reduce memory needs, and expert parallelism supports large models across GPU clusters.

Hardware support covers NVIDIA, AMD and Ascend. On AMD ROCm GPUs, verl supports FSDP, FSDP2 and Megatron-LM training with vLLM for generation. Experiments can report to wandb, swanlab, MLflow or TensorBoard.

Similar to verl