DiffRhythm is a local AI music generation model for musicians, developers and researchers who want to create full-length songs on their own hardware. It uses latent diffusion to generate songs with vocals and accompaniment, and can also produce instrumental music. The full model supports songs up to 4 minutes and 45 seconds.
You can guide the musical style with a written description or a reference WAV recording. Text prompts can describe a genre, mood, scene or instrumentation, so you don't need an audio example to request a particular sound. Lyrics use the LRC format, which pairs text with timing information. Song editing and continuation let you work with existing material as well as generate a complete piece.
The software runs on macOS, Windows and Linux, and supports Docker deployment. Local deployment runs generation on your own machine; a hosted Hugging Face Space demo provides a separate way to try the model. The base model needs at least 8 GB of VRAM with chunked decoding, and may need more memory without it. That hardware requirement matters if you're choosing between running it yourself and trying the hosted demo.
DiffRhythm's code and DiT model weights are open source under the Apache License 2.0. The license permits use, modification and redistribution with the required copyright notice and disclaimer. The VAE checkpoint and other downloaded components have separate terms; the Apache statement does not establish their licenses.
Claim this page and we'll verify you by hand. DiffRhythm gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find DiffRhythm?Promote it
Something wrong or outdated on this page?
2.6KUpdated 2 years ago
macOS · Linux · Web#Batch processing#Hugging Face integration
AudioLDM 2 generates sound effects, music and speech on your own hardware. It's a Python tool for people experimenting with synthetic audio, including sound designers and researchers who want to work with pretrained models. A Gradio browser interface and command-line tools provide access to local generation; a hosted Hugging Face demo is also available.
10.6KUpdated 2 days agoApache-2.0
Linux · Web#Hugging Face integration#Multilingual#Multimodal input
4.9KUpdated 7 months agoApache-2.0
macOS#LoRA#Multilingual
1.8KUpdated 2 months agoApache-2.0
macOS#MLX#Streaming inference
23.7KUpdated 2 years agoMIT
#Multimodal input
MusicGen is Meta AI's music generation model within AudioCraft, a PyTorch library for developers and audio researchers who want to generate music in their own computing environment. It creates music from text descriptions and can use a melody to guide the result. AudioCraft includes both inference and training code, so it's suited to people building audio tools or studying music generation.
4.8KUpdated 3 weeks agoApache-2.0
macOS · Windows · Linux · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
YuE is an open-source music generation project for musicians, songwriters and developers who want to turn lyrics and a style prompt into songs with vocals and accompaniment. Its YuE2 models create an editable melody and chord score before generating the recording, so you can review the composition and change musical details before hearing the result.
ACE-Step is an open-source music generation model for musicians, producers and developers who want to create and edit music on their own hardware. It generates songs with vocals or instrumental tracks from text descriptions and supplied lyrics. You can choose the duration and describe the sound with genre tags, longer prompts or a scene description.
Magenta RealTime 2 is a local AI music model and synthesis engine for musicians and developers who want to play or build AI musical instruments on a laptop. It generates streaming audio in real time, with open weights and code under the Apache 2.0 license.
Lollms WebUI is a local, single-user AI interface for people who want text chat and media generation in one place. It runs on Windows, macOS and Linux, with Docker support, and lets writers, developers and other users choose models and task-specific personalities. It's free and open source under Apache 2.0. The project receives minimal maintenance.