Favicon of LLaMA-Factory

LLaMA-Factory

Open-source LLM fine-tuning framework under Apache 2.0, with LoRA, QLoRA, multimodal training and inference through vLLM or SGLang.

Screenshot of LLaMA-Factory website

LLaMA-Factory is an open-source framework for developers and researchers who want to adapt language and multimodal models on their own hardware. It brings training and inference into one toolkit, with support for LLaMA, Qwen3, Qwen3-VL, DeepSeek, Gemma, Mistral and LLaVA. Its license is Apache 2.0.

Training covers continued pre-training, supervised fine-tuning and reward modeling. Preference training methods include DPO, PPO, KTO and ORPO, so the framework can support both teaching a model a task and adjusting how it responds. Multimodal tasks include image understanding, visual grounding, video recognition and audio understanding; text tasks include multi-turn dialogue and tool use.

You can train all model weights, freeze selected weights or use LoRA and quantized QLoRA to reduce memory requirements. Hardware needs depend heavily on the model and method. Estimated memory needs for a 7B model are roughly 6 GB with 4-bit QLoRA, compared with roughly 16 GB for LoRA. It supports CUDA, along with training acceleration tools such as FlashAttention-2, Unsloth and Liger Kernel.

For comparing training runs, it integrates with LlamaBoard, TensorBoard, Wandb and MLflow. Trained models can be used through a Gradio interface, a command-line interface or an OpenAI-style API, with vLLM or SGLang as inference backends.

Similar to LLaMA-Factory