Favicon of Stable Audio Tools

Stable Audio Tools

Open-source Python toolkit for local AI audio generation, fine-tuning and training, with a Gradio interface and support for Stable Audio Open.

Stable Audio Tools is an MIT-licensed Python toolkit for developers and audio researchers who want to generate audio on their own hardware or train models on their own recordings. It combines model inference with training and fine-tuning, so you can work with pretrained models or build a model around a specific audio dataset.

The toolkit supports Stable Audio Open, including stabilityai/stable-audio-open-1.0, and can load models from Hugging Face or local checkpoints. A basic Gradio browser interface lets you test trained models. Access to Stable Audio Open on Hugging Face requires accepting the model's terms; the toolkit's MIT license is separate from those terms.

Its model support includes conditional and unconditional diffusion, audio inpainting, autoencoders and language models. You can train from scratch, resume an existing training run or fine-tune pretrained weights. It also supports testing fine-tuned decoders and using pretrained autoencoders within latent diffusion models.

Training uses PyTorch Lightning and can span multiple GPUs or machines. Gradient accumulation helps increase the effective training batch size on smaller GPUs. Datasets can come from local audio folders or WebDataset files in Amazon S3. Training requires a Weights & Biases account for logging outputs and demos, so that part of the workflow uses an external service. The toolkit can export checkpoints containing only the model for inference and further fine-tuning.

Similar to Stable Audio Tools