so-vits-svc is an offline AI framework for changing the voice in an existing singing recording while preserving its pitch and intonation. It's aimed at developers and researchers who want to train their own singing voices, including fictional character voices. The project is archived and no longer maintained.
It runs entirely on your own hardware and doesn't collect user information, input audio or models. The Python project is open source under AGPL-3.0. It doesn't include trained voice models, so users need to train their own before converting audio. It requires source audio rather than text and doesn't support text-to-speech. The maintainers describe it as an academic framework rather than software intended for production use.
Its SoftVC and VITS approach works directly with audio features, without converting the performance into text first. NSF HiFiGAN handles audio generation. Supported encoders include ContentVec, HuBERT, Whisper-PPG and WavLM, giving researchers a choice of how the system represents the source voice.
Conversion controls include pitch shifting and static or dynamic voice mixing. Loudness embedding can preserve the input recording's volume changes, while optional shallow diffusion can improve sound quality and reduce electronic-sounding artifacts. It also supports feature retrieval adapted from RVC. Separate applications can use its models: MoeVoiceStudio works with ONNX models, and w-okada/voice-changer provides a client for real-time conversion.
Claim this page and we'll verify you by hand. so-vits-svc gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find so-vits-svc?Promote it
Something wrong or outdated on this page?
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
38.6KUpdated 2 months agoMIT
Windows · Linux · Web#Hugging Face integration#ONNX#Voice conversion
214Updated 3 weeks agoMIT
Linux · Docker · Web#Home Assistant integration#Hugging Face integration#Multilingual
21.1KUpdated 2 days agoApache-2.0
macOS · Web#GGUF#Hugging Face integration#Multilingual
26.6KUpdated 2 months agoMIT
Linux#Multilingual#Voice cloning#Voice conversion
11.8KUpdated 4 days agoApache-2.0
Docker#Distributed execution#Hugging Face integration#LoRA
RVC WebUI is a local AI voice conversion tool for people who want to train a custom voice, change the voice in a recording, or use a live voice changer. It runs on Windows and Linux, including Ubuntu servers, with a browser interface for training and conversion and a separate interface for live use. It's free and open source under the MIT license.
Wyoming Piper connects Piper's local text-to-speech engine to Home Assistant and other clients that use the Wyoming protocol. It's for people building a voice assistant on their own hardware who need speech generation as a self-hosted service. The project is open source under the MIT license.
Candle is a Rust machine learning framework for developers who want to embed local AI in applications or deploy models on their own servers. It produces lightweight binaries that don't need Python in production, making it a candidate for serverless inference where a large runtime can slow startup. Its API uses tensor operations familiar to PyTorch developers.
Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.
Ludwig is an open-source Python framework for developers and researchers who want to train custom AI models on their own hardware. A YAML file describes the model and training pipeline, while Ludwig handles preprocessing, training and evaluation. It uses the Apache 2.0 license. Install the Python package with the optional LLM dependencies for fine-tuning; current source requires Python 3.12 or later.