
Stable Audio Open is a text-to-audio model you can run on your own hardware to generate sound effects, field recordings and music samples. It's aimed at artists and machine learning practitioners experimenting with audio generation, and it performs better on environmental sounds and effects than on music.
English prompts guide the output, and you can choose the clip length. It generates stereo audio, with controls for negative prompts and repeatable generation through a fixed seed. It can't produce realistic vocals, and results vary across musical styles and cultures. Prompts in other languages work less well than English.
The model works with stable-audio-tools and Hugging Face Diffusers. Local inference can use a CPU or a CUDA GPU; stable-audio-tools also provides a basic Gradio browser interface for trying trained models. Access to the model files requires a Hugging Face account, acceptance of the terms and agreement to share contact information.
The model uses the Stability AI Community License, while the separate stable-audio-tools Python library uses MIT. Commercial use is subject to Stability AI's licensing terms. The library supports training and fine-tuning, including training across multiple GPUs and machines; its training workflow uses a Weights & Biases account to log outputs and demos.
Training audio comes from Freesound and the Free Music Archive under CC0, CC BY and CC Sampling+ licenses. Attribution for those recordings is available.
Claim this page and we'll verify you by hand. Stable Audio Open gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Stable Audio Open?Promote it
Something wrong or outdated on this page?
2.6KUpdated 2 years ago
macOS · Linux · Web#Batch processing#Hugging Face integration
AudioLDM 2 generates sound effects, music and speech on your own hardware. It's a Python tool for people experimenting with synthetic audio, including sound designers and researchers who want to work with pretrained models. A Gradio browser interface and command-line tools provide access to local generation; a hosted Hugging Face demo is also available.
39.3KUpdated 2 years agoMIT
#Hugging Face integration#Multilingual
Bark is Suno's local text-to-audio model for developers and researchers who want to generate speech alongside other sounds. It can produce laughter, crying, music and background noise within its output. Its generative approach suits audio experiments, though it can depart from the supplied script and doesn't guarantee clean, studio-quality speech.
10.4KUpdated 3 years agoMIT
macOS · Windows · Linux · Docker#Batch processing#Quantization
2.3KUpdated 10 months agoApache-2.0
macOS · Windows · Linux · Docker#Hugging Face integration#Multimodal input
10.6KUpdated 2 days agoApache-2.0
Linux · Web#Hugging Face integration#Multilingual#Multimodal input
4.9KUpdated 7 months agoApache-2.0
macOS#LoRA#Multilingual
Demucs separates a finished song into vocals, drums, bass and the remaining accompaniment on your own computer. It's for musicians who need individual stems or a vocal-free backing track, and developers building audio tools. The project is archived and no longer maintained. Its Python code is open source under the MIT license.
DiffRhythm is a local AI music generation model for musicians, developers and researchers who want to create full-length songs on their own hardware. It uses latent diffusion to generate songs with vocals and accompaniment, and can also produce instrumental music. The full model supports songs up to 4 minutes and 45 seconds.
YuE is an open-source music generation project for musicians, songwriters and developers who want to turn lyrics and a style prompt into songs with vocals and accompaniment. Its YuE2 models create an editable melody and chord score before generating the recording, so you can review the composition and change musical details before hearing the result.
ACE-Step is an open-source music generation model for musicians, producers and developers who want to create and edit music on their own hardware. It generates songs with vocals or instrumental tracks from text descriptions and supplied lyrics. You can choose the duration and describe the sound with genre tags, longer prompts or a scene description.