Favicon of Musubi Tuner

Musubi Tuner

Local LoRA training scripts for image and video models, with Windows and Linux support, NVIDIA GPU memory savings, and multi-GPU training.

Musubi Tuner is a Python toolkit for training LoRA adapters for image and video generation models on your own hardware. It's aimed at people who want to customize these models using their own datasets and are comfortable working with training scripts. It also includes image and video generation scripts for supported architectures.

Its model coverage includes HunyuanVideo, Wan2.1/2.2 and FramePack for video, alongside FLUX.1 Kontext, FLUX.2 dev/klein, Qwen-Image and Z-Image. Single-frame training is available for Wan and FramePack. The project is unofficial and isn't affiliated with the teams behind these models.

Memory-saving features are a central part of the toolkit. CPU offloading moves model blocks out of GPU memory, while reduced-precision weights can lower memory use for supported architectures. Accelerate provides multi-GPU training. You can generate samples during training and track progress or gradient diagnostics to assess a run before it finishes.

The toolkit supports Windows, including Windows on ARM64, and Linux. Its PyTorch training setup uses NVIDIA CUDA GPUs. Recommended hardware is at least 12GB of VRAM for image training or 24GB for video training, plus 64GB of system RAM; actual needs depend on resolution and training settings. A system with 32GB of RAM and swap may also work.

Dataset tools cache image and video representations and text encoder outputs ahead of training. They can skip broken media files rather than stop an entire caching run. Utilities also merge and convert LoRA weights.

The repository is experimental and is not intended for production use. Breaking changes may occur; the maintainers recommend released versions as stable reference points. Model weights have separate licenses.

Similar to Musubi Tuner