Favicon of Stable Audio Open

Stable Audio Open

A local text-to-audio model for sound effects and music experiments, with CPU or CUDA support and access under the Stability AI Community License.

Screenshot of Stable Audio Open website

Stable Audio Open is a text-to-audio model you can run on your own hardware to generate sound effects, field recordings and music samples. It's aimed at artists and machine learning practitioners experimenting with audio generation, and it performs better on environmental sounds and effects than on music.

English prompts guide the output, and you can choose the clip length. It generates stereo audio, with controls for negative prompts and repeatable generation through a fixed seed. It can't produce realistic vocals, and results vary across musical styles and cultures. Prompts in other languages work less well than English.

The model works with stable-audio-tools and Hugging Face Diffusers. Local inference can use a CPU or a CUDA GPU; stable-audio-tools also provides a basic Gradio browser interface for trying trained models. Access to the model files requires a Hugging Face account, acceptance of the terms and agreement to share contact information.

The model uses the Stability AI Community License, while the separate stable-audio-tools Python library uses MIT. Commercial use is subject to Stability AI's licensing terms. The library supports training and fine-tuning, including training across multiple GPUs and machines; its training workflow uses a Weights & Biases account to log outputs and demos.

Training audio comes from Freesound and the Free Music Archive under CC0, CC BY and CC Sampling+ licenses. Attribution for those recordings is available.

Similar to Stable Audio Open