python-audio-separator separates recordings into vocals, instrumentals and individual instruments on your own hardware. It's an open-source Python package under the MIT license, aimed at karaoke creators and developers who want audio separation in scripts or their own applications.
Its model selection draws on Ultimate Vocal Remover (UVR), with support for MDX-Net, VR Arch, Demucs and MDXC, including BS-RoFormer and MelBand RoFormer. Depending on the model, it can extract drums, bass, guitar or piano, or reduce noise, echo and reverb. It accepts common audio formats including WAV, MP3, FLAC and M4A, and supports pretrained models in PTH and ONNX formats.
The command-line interface handles batch processing, while the Python API lets developers incorporate separation into other software. You can compare models by the stems they produce and separation scores where available. Ensembles combine outputs from multiple models, with presets for cleaner instrumentals, fuller instrument preservation or balanced vocals.
Processing can run locally on a CPU, an NVIDIA GPU with CUDA, or Apple Silicon through MPS and CoreML. Docker images support CPU and CUDA workloads on amd64 and arm64 systems. Models download automatically on first use, so obtaining them requires an internet connection.
For long recordings, chunked processing reduces peak memory use and joins the separated audio afterward. Chunk boundaries can occasionally introduce minor artifacts.
Claim this page and we'll verify you by hand. python-audio-separator gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find python-audio-separator?Promote it
Something wrong or outdated on this page?
10.4KUpdated 3 years agoMIT
macOS · Windows · Linux · Docker#Batch processing#Quantization
Demucs separates a finished song into vocals, drums, bass and the remaining accompaniment on your own computer. It's for musicians who need individual stems or a vocal-free backing track, and developers building audio tools. The project is archived and no longer maintained. Its Python code is open source under the MIT license.
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
26.4KUpdated 2 years agoMIT
macOS · Windows · Linux
38.6KUpdated 2 months agoMIT
Windows · Linux · Web#Hugging Face integration#ONNX#Voice conversion
28.5KUpdated 3 months agoMIT
Windows · Docker
Spleeter is Deezer's music source separation library for developers and audio researchers who want to split recordings into separate vocal and instrumental tracks on their own hardware. It includes pretrained models, so you can separate audio without first training a model. The library is open source under the MIT license.
3.3KUpdated 3 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
Ultimate Vocal Remover is a desktop app for separating vocals and other stems from audio files on your own computer. It's for people making karaoke backing tracks, isolating a vocal, or working with separate parts of a song. It runs on Windows, macOS and Linux, and its MIT license allows you to use and modify the software.
RVC WebUI is a local AI voice conversion tool for people who want to train a custom voice, change the voice in a recording, or use a live voice changer. It runs on Windows and Linux, including Ubuntu servers, with a browser interface for training and conversion and a separate interface for live use. It's free and open source under the MIT license.