Favicon of so-vits-svc

so-vits-svc

An open-source singing voice conversion framework that runs fully offline with user-trained models. Licensed under AGPL-3.0; archived and no longer maintained.

so-vits-svc is an offline AI framework for changing the voice in an existing singing recording while preserving its pitch and intonation. It's aimed at developers and researchers who want to train their own singing voices, including fictional character voices. The project is archived and no longer maintained.

It runs entirely on your own hardware and doesn't collect user information, input audio or models. The Python project is open source under AGPL-3.0. It doesn't include trained voice models, so users need to train their own before converting audio. It requires source audio rather than text and doesn't support text-to-speech. The maintainers describe it as an academic framework rather than software intended for production use.

Its SoftVC and VITS approach works directly with audio features, without converting the performance into text first. NSF HiFiGAN handles audio generation. Supported encoders include ContentVec, HuBERT, Whisper-PPG and WavLM, giving researchers a choice of how the system represents the source voice.

Conversion controls include pitch shifting and static or dynamic voice mixing. Loudness embedding can preserve the input recording's volume changes, while optional shallow diffusion can improve sound quality and reduce electronic-sounding artifacts. It also supports feature retrieval adapted from RVC. Separate applications can use its models: MoeVoiceStudio works with ONNX models, and w-okada/voice-changer provides a client for real-time conversion.

Similar to so-vits-svc