Favicon of DiffSynth-Studio

DiffSynth-Studio

Open-source diffusion engine in Python for local image, video and music generation, with low-VRAM inference and model training under Apache 2.0.

DiffSynth-Studio is a Python diffusion model engine for developers and researchers who want to generate media and train models on their own hardware. It supports large models on consumer GPUs through memory offloading and quantization, with inference and training in the same framework. It's open source under Apache 2.0.

Its model support covers image generation and editing with Qwen-Image and FLUX, video generation with Wan, and music generation with ACE-Step. MiniMax-H3 supports video with audio, including keyframe guidance and reference-based generation. Other supported tasks include character animation, image layer separation and image quality evaluation.

Memory management moves model weights between disk, system memory and GPU memory to reduce VRAM demand. Quantization reduces memory use for both generation and LoRA training. Supported backends include bitsandbytes, torchao and EntroPack, which can compress weights with either exact recovery or controlled loss.

For training, the framework supports full model updates, LoRAs and adapters with additional inputs. CPU offloading makes large-model LoRA training possible on consumer GPUs, while split training separates data preparation from the computations that update model weights.

DiffSynth-ComfyUI brings the engine's inference capabilities into ComfyUI workflows. DiffSynth-WebUI provides a separate interface for privately deployed LoRA training. ModelScope AIGC Zone and Civision offer hosted experiences powered by the engine.

Some older features are no longer maintained and require historical releases. A small development team limits the pace of new features and issue fixes. Model weights have separate licenses.

Similar to DiffSynth-Studio