Favicon of Riffusion

Riffusion

Local AI music generation library built on Stable Diffusion. Run it on your own hardware with CUDA, Apple Silicon or CPU. MIT licensed and no longer maintained.

Riffusion is a Python library for generating music and audio on your own hardware using Stable Diffusion. It's for developers and musicians who want to experiment with text-driven sound generation or build it into an app. The hobby project is no longer actively maintained.

Its distinctive approach uses spectrogram images, visual representations of sound that it converts into audio clips. The generation pipeline combines image conditioning with interpolation between text prompts, so it can blend between different descriptions rather than generate only isolated clips. It also converts existing audio into spectrogram images.

A Streamlit playground provides an interactive interface for exploration. Developers can use the command-line tools for audio and image tasks, or a self-hosted Flask API to connect the model to an application. The related original Riffusion web app can run locally with that server. This entry describes the hobby library and its original checkpoint, separate from later commercial products.

The library is open source under the MIT license. It works with the riffusion-model-v1 checkpoint and accepts local checkpoints or Hugging Face models in diffusers format. The original checkpoint uses CreativeML Open RAIL-M, with use restrictions separate from the MIT code license.

CUDA is the recommended backend for speed. Real-time generation needs a capable NVIDIA GPU, with the RTX 3090 and A10G given as examples. CPU inference works but is slow. Apple Silicon supports inference through MPS, though some audio processing falls back to the CPU.

Similar to Riffusion