Favicon of XTuner

XTuner

Open-source LLM training engine for GPU and Ascend NPU hardware, with multimodal fine-tuning, reinforcement learning and Apache 2.0 licensing.

XTuner is an open-source LLM training engine for researchers and teams training large mixture-of-experts (MoE) models on their own hardware. It supports GPU and Ascend NPU training, with an emphasis on memory use and distributed training efficiency at scales reaching a trillion parameters.

Its dropless training approach processes tokens without discarding them when experts receive uneven workloads. It reduces the need to distribute experts across machines compared with traditional 3D parallel training. Memory optimizations also support long training sequences, while DeepSpeed Ulysses integration allows sequence lengths to scale further. The engine handles expert load imbalance during long-sequence training.

XTuner supports vision-language pre-training and supervised fine-tuning, plus Group Relative Policy Optimization (GRPO) for reinforcement learning. Supported models include Intern S1, InternVL, Qwen3 Dense and Qwen3 MoE on both GPUs and Ascend NPUs. GPT OSS, DeepSeek V3 and KIMI K2 support GPU training. GPU paths include FP8 and BF16 precision; supported NPU models use BF16.

For teams choosing a training backend, its focus on large MoE workloads is the main distinction. XTuner uses fully sharded data parallel training and integrates with LMDeploy, vLLM and SGLang inference engines. GraphGen can supply synthetic fine-tuning data. The Python project uses the Apache 2.0 license, while models and datasets retain their own licensing requirements.

Similar to XTuner