Favicon of diffusion-pipe

diffusion-pipe

An open-source diffusion model trainer that splits large models across GPUs. Uses DeepSpeed and supports Windows through WSL 2 under GPL-3.0.

diffusion-pipe is a local diffusion model training tool for people fine-tuning image and video models on their own GPU hardware. Its main distinction is that it can divide a model across several GPUs when it won't fit on one, while also distributing training work across GPUs. The Python project is open source under GPL-3.0 and uses DeepSpeed.

Model support includes SDXL, Flux, Stable Diffusion 3, Qwen-Image and Qwen-Image-Edit, alongside video models such as Wan, HunyuanVideo and LTX-Video. Image and video training share the same framework, and datasets can contain both media types with captions. MiniMax H3 supports text-to-image and text-to-video-with-audio training; LTX 2.3 training excludes audio.

For hardware with limited VRAM, it provides block swapping on supported models, activation checkpointing and an 8-bit optimizer. It can also train LoRAs directly on quantized models and use ComfyUI quantized weights for supported models. Windows users can run it through WSL 2; native Windows compatibility is limited by DeepSpeed. The documented hardware setup uses NVIDIA GPUs and CUDA.

TensorBoard metrics and evaluation on held-out data help users assess training quality beyond the training loss. Checkpoints preserve training state so interrupted runs can resume. Cached latents and text embeddings stay on disk for reuse between runs, reducing the need to keep encoders loaded during training. Text-encoder LoRA training isn't supported. Saved models and adapters use safetensors weights.

Similar to diffusion-pipe