Favicon of python-audio-separator

python-audio-separator

Local AI audio separator for vocals and instruments, with a Python API, batch processing, CPU and GPU support, and an MIT license.

python-audio-separator separates recordings into vocals, instrumentals and individual instruments on your own hardware. It's an open-source Python package under the MIT license, aimed at karaoke creators and developers who want audio separation in scripts or their own applications.

Its model selection draws on Ultimate Vocal Remover (UVR), with support for MDX-Net, VR Arch, Demucs and MDXC, including BS-RoFormer and MelBand RoFormer. Depending on the model, it can extract drums, bass, guitar or piano, or reduce noise, echo and reverb. It accepts common audio formats including WAV, MP3, FLAC and M4A, and supports pretrained models in PTH and ONNX formats.

The command-line interface handles batch processing, while the Python API lets developers incorporate separation into other software. You can compare models by the stems they produce and separation scores where available. Ensembles combine outputs from multiple models, with presets for cleaner instrumentals, fuller instrument preservation or balanced vocals.

Processing can run locally on a CPU, an NVIDIA GPU with CUDA, or Apple Silicon through MPS and CoreML. Docker images support CPU and CUDA workloads on amd64 and arm64 systems. Models download automatically on first use, so obtaining them requires an internet connection.

For long recordings, chunked processing reduces peak memory use and joins the separated audio afterward. Chunk boundaries can occasionally introduce minor artifacts.

Similar to python-audio-separator