Favicon of OpenRLHF

OpenRLHF

An open source RLHF framework for training models on your own NVIDIA GPUs, with HuggingFace model support and Ray, vLLM and DeepSpeed backends.

OpenRLHF is a self-hosted Python framework for researchers and teams training language models with human feedback or custom rewards. It runs on your own NVIDIA GPU hardware, with Docker support and distributed training across servers. It's open source under Apache 2.0.

Its main distinction is the combination of Ray for distributing work, vLLM for generating responses, and DeepSpeed for training. Training models and generation engines can share GPUs to reduce idle time, or run across separate resources for larger workloads. It works directly with HuggingFace Transformers models and produces compatible checkpoints.

The training pipeline covers supervised fine-tuning, reward model training and preference learning through DPO, IPO and cDPO. For reinforcement learning, it supports PPO, REINFORCE++, GRPO and RLOO, along with critic-free FlashREINFORCE for asynchronous agent training. Custom rewards let teams train against task-specific scoring criteria.

The same agent framework handles single responses and multi-turn interactions with environment feedback, independently of the chosen RL algorithm. Vision-language training supports models such as Qwen3.5, including tasks that use images or screenshots as feedback. LoRA and QLoRA are available for fine-tuning, alongside mixture-of-experts model support.

For tool-using agents, OpenRLHF can expose vLLM through a local OpenAI-compatible API and collect training traces across conversations. TensorBoard and Weights & Biases integrations provide training logs, while checkpoint recovery lets interrupted runs resume.

Similar to OpenRLHF