Favicon of GPT-SoVITS

GPT-SoVITS

Local text-to-speech and voice cloning software with a browser interface, multilingual speech generation, and an MIT license. Runs on Windows, Linux and macOS.

GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.

Speech generation supports English, Japanese, Korean, Cantonese and Chinese, including output in a language different from the training recordings. The browser interface covers model training and speech generation.

Included preparation tools separate vocals from accompaniment, remove reverberation and split recordings into training clips. ASR backends include Fun-ASR-Nano, SenseVoice, classic FunASR and Faster Whisper. Text labeling and transcription proofreading help prepare the training data.

Windows, Linux and macOS installations are documented, along with Docker. Hardware options include CPU, NVIDIA CUDA, AMD ROCm on Linux and Apple Silicon. The project warns that GPU training on Macs can produce lower voice quality and recommends CPU training there. Download the required models before local use; classic Chinese ASR also documents downloading its speech, VAD and punctuation models for offline use.

Similar to GPT-SoVITS