Favicon of DiffRhythm

DiffRhythm

An open-source AI music generator that runs locally on macOS, Windows and Linux, with text or audio style prompts and Apache 2.0 code and DiT weights.

DiffRhythm is a local AI music generation model for musicians, developers and researchers who want to create full-length songs on their own hardware. It uses latent diffusion to generate songs with vocals and accompaniment, and can also produce instrumental music. The full model supports songs up to 4 minutes and 45 seconds.

You can guide the musical style with a written description or a reference WAV recording. Text prompts can describe a genre, mood, scene or instrumentation, so you don't need an audio example to request a particular sound. Lyrics use the LRC format, which pairs text with timing information. Song editing and continuation let you work with existing material as well as generate a complete piece.

The software runs on macOS, Windows and Linux, and supports Docker deployment. Local deployment runs generation on your own machine; a hosted Hugging Face Space demo provides a separate way to try the model. The base model needs at least 8 GB of VRAM with chunked decoding, and may need more memory without it. That hardware requirement matters if you're choosing between running it yourself and trying the hosted demo.

DiffRhythm's code and DiT model weights are open source under the Apache License 2.0. The license permits use, modification and redistribution with the required copyright notice and disclaimer. The VAE checkpoint and other downloaded components have separate terms; the Apache statement does not establish their licenses.

Similar to DiffRhythm