Stable Audio Tools is an MIT-licensed Python toolkit for developers and audio researchers who want to generate audio on their own hardware or train models on their own recordings. It combines model inference with training and fine-tuning, so you can work with pretrained models or build a model around a specific audio dataset.
The toolkit supports Stable Audio Open, including stabilityai/stable-audio-open-1.0, and can load models from Hugging Face or local checkpoints. A basic Gradio browser interface lets you test trained models. Access to Stable Audio Open on Hugging Face requires accepting the model's terms; the toolkit's MIT license is separate from those terms.
Its model support includes conditional and unconditional diffusion, audio inpainting, autoencoders and language models. You can train from scratch, resume an existing training run or fine-tune pretrained weights. It also supports testing fine-tuned decoders and using pretrained autoencoders within latent diffusion models.
Training uses PyTorch Lightning and can span multiple GPUs or machines. Gradient accumulation helps increase the effective training batch size on smaller GPUs. Datasets can come from local audio folders or WebDataset files in Amazon S3. Training requires a Weights & Biases account for logging outputs and demos, so that part of the workflow uses an external service. The toolkit can export checkpoints containing only the model for inference and further fine-tuning.
Claim this page and we'll verify you by hand. Stable Audio Tools gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Stable Audio Tools?Promote it
Something wrong or outdated on this page?
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
21.9KUpdated 1 day agoMIT
macOS · Windows · Linux · iOS · Android · Web#Distributed execution#ONNX
13.7KUpdated 3 weeks agoApache-2.0
#Hugging Face integration#LoRA#Quantization
14.2KUpdated 5 days ago
#Hugging Face integration#Multimodal input
OpenCLIP is a Python and PyTorch library for developers and researchers who want to match images with text on their own hardware. It implements OpenAI's CLIP approach: images and descriptions become numerical representations that the model can compare. This supports image search and zero-shot classification, where text labels define the categories without a separate classifier trained for each task.
3.5KUpdated 6 days agoApache-2.0
#Hugging Face integration#ONNX#Quantization
10.6KUpdated 3 months agoMIT
#Hugging Face integration#Speaker diarization#Voice activity detection
ONNX Runtime is an open source inference and training engine for developers building AI into apps and services. It runs ONNX models across desktop systems, mobile devices, web browsers and servers. It's a fit when you need the same model format to work in several places, including on a user's device.
LitGPT is a Python toolkit for developers and researchers who want to train, adapt and serve language models on their own hardware or servers. Its model implementations are written directly, with little abstraction between you and the code, so you can inspect model behavior and modify it for research or custom applications. It's open source under Apache 2.0.
Optimum is a collection of Python packages for developers who want to train or run Hugging Face models more efficiently on specific hardware. It extends Transformers, Diffusers, TIMM and Sentence Transformers, with integrations for local machines, mobile and edge devices, and cloud accelerators. It's open source under Apache 2.0.
pyannote.audio is a Python toolkit that separates an audio recording into timed segments labeled by speaker. It's for developers and researchers who need to track who spoke when, with pretrained models that run on their own hardware. The toolkit is open source under the MIT license.