
Spleeter is Deezer's music source separation library for developers and audio researchers who want to split recordings into separate vocal and instrumental tracks on their own hardware. It includes pretrained models, so you can separate audio without first training a model. The library is open source under the MIT license.
Its models produce different sets of stems, the separate audio tracks extracted from a mix:
Written in Python with TensorFlow, Spleeter works as a command-line tool or as a library within an audio processing application. It also supports Docker. This makes it a fit for people building their own separation tools or processing recordings through an existing development pipeline.
GPU processing is a notable strength. Deezer reports that four-stem separation can run at 100 times real-time speed on a GPU. Researchers can also train their own separation models if they have a dataset of isolated sound sources.
Spleeter's pretrained models have been used in audio software including iZotope RX, SpectralLayers, Acoustica and VirtualDJ. The Spleeter 4 Max project brings it into Ableton Live. Deezer also offers a separate commercial product, Spleeter Pro, with professional support.
Claim this page with an email at research.deezer.com. Spleeter gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Spleeter?Promote it
Something wrong or outdated on this page?
10.4KUpdated 3 years agoMIT
macOS · Windows · Linux · Docker#Batch processing#Quantization
Demucs separates a finished song into vocals, drums, bass and the remaining accompaniment on your own computer. It's for musicians who need individual stems or a vocal-free backing track, and developers building audio tools. The project is archived and no longer maintained. Its Python code is open source under the MIT license.
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
1.4KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker#Batch processing#ONNX
3.3KUpdated 3 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
38.6KUpdated 2 months agoMIT
Windows · Linux · Web#Hugging Face integration#ONNX#Voice conversion
26.4KUpdated 2 years agoMIT
macOS · Windows · Linux
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
python-audio-separator separates recordings into vocals, instrumentals and individual instruments on your own hardware. It's an open-source Python package under the MIT license, aimed at karaoke creators and developers who want audio separation in scripts or their own applications.
RVC WebUI is a local AI voice conversion tool for people who want to train a custom voice, change the voice in a recording, or use a live voice changer. It runs on Windows and Linux, including Ubuntu servers, with a browser interface for training and conversion and a separate interface for live use. It's free and open source under the MIT license.
Ultimate Vocal Remover is a desktop app for separating vocals and other stems from audio files on your own computer. It's for people making karaoke backing tracks, isolating a vocal, or working with separate parts of a song. It runs on Windows, macOS and Linux, and its MIT license allows you to use and modify the software.