Riffusion is a Python library for generating music and audio on your own hardware using Stable Diffusion. It's for developers and musicians who want to experiment with text-driven sound generation or build it into an app. The hobby project is no longer actively maintained.
Its distinctive approach uses spectrogram images, visual representations of sound that it converts into audio clips. The generation pipeline combines image conditioning with interpolation between text prompts, so it can blend between different descriptions rather than generate only isolated clips. It also converts existing audio into spectrogram images.
A Streamlit playground provides an interactive interface for exploration. Developers can use the command-line tools for audio and image tasks, or a self-hosted Flask API to connect the model to an application. The related original Riffusion web app can run locally with that server. This entry describes the hobby library and its original checkpoint, separate from later commercial products.
The library is open source under the MIT license. It works with the riffusion-model-v1 checkpoint and accepts local checkpoints or Hugging Face models in diffusers format. The original checkpoint uses CreativeML Open RAIL-M, with use restrictions separate from the MIT code license.
CUDA is the recommended backend for speed. Real-time generation needs a capable NVIDIA GPU, with the RTX 3090 and A10G given as examples. CPU inference works but is slow. Apple Silicon supports inference through MPS, though some audio processing falls back to the CPU.
Claim this page and we'll verify you by hand. Riffusion gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Riffusion?Promote it
Something wrong or outdated on this page?
1.8KUpdated 2 months agoApache-2.0
macOS#MLX#Streaming inference
Magenta RealTime 2 is a local AI music model and synthesis engine for musicians and developers who want to play or build AI musical instruments on a laptop. It generates streaming audio in real time, with open weights and code under the Apache 2.0 license.
13.2KUpdated 2 days agoApache-2.0
#ControlNet#Image-to-image#Inpainting
23.7KUpdated 2 years agoMIT
#Multimodal input
MusicGen is Meta AI's music generation model within AudioCraft, a PyTorch library for developers and audio researchers who want to generate music in their own computing environment. It creates music from text descriptions and can use a melody to guide the result. AudioCraft includes both inference and training code, so it's suited to people building audio tools or studying music generation.
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
2.3KUpdated 10 months agoApache-2.0
macOS · Windows · Linux · Docker#Hugging Face integration#Multimodal input
1.3KUpdated 1 day ago
macOS · Windows · Linux#GGUF#Hugging Face integration#LoRA
GPTQModel is a Python toolkit for developers compressing LLMs and running them on their own hardware or servers. It brings model calibration, compression, quality checks and inference into one API, so teams can compare quantization methods without adopting a separate tool for each one.
DiffSynth-Studio is a Python diffusion model engine for developers and researchers who want to generate media and train models on their own hardware. It supports large models on consumer GPUs through memory offloading and quantization, with inference and training in the same framework. It's open source under Apache 2.0.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
DiffRhythm is a local AI music generation model for musicians, developers and researchers who want to create full-length songs on their own hardware. It uses latent diffusion to generate songs with vocals and accompaniment, and can also produce instrumental music. The full model supports songs up to 4 minutes and 45 seconds.