GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
Speech generation supports English, Japanese, Korean, Cantonese and Chinese, including output in a language different from the training recordings. The browser interface covers model training and speech generation.
Included preparation tools separate vocals from accompaniment, remove reverberation and split recordings into training clips. ASR backends include Fun-ASR-Nano, SenseVoice, classic FunASR and Faster Whisper. Text labeling and transcription proofreading help prepare the training data.
Windows, Linux and macOS installations are documented, along with Docker. Hardware options include CPU, NVIDIA CUDA, AMD ROCm on Linux and Apple Silicon. The project warns that GPU training on Macs can produce lower voice quality and recommends CPU training there. Download the required models before local use; classic Chinese ASR also documents downloading its speech, VAD and punctuation models for offline use.
Claim this page and we'll verify you by hand. GPT-SoVITS gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find GPT-SoVITS?Promote it
Something wrong or outdated on this page?
3.3KUpdated 3 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
38.6KUpdated 2 months agoMIT
Windows · Linux · Web#Hugging Face integration#ONNX#Voice conversion
3.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Web#Batch processing#Voice conversion
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.
RVC WebUI is a local AI voice conversion tool for people who want to train a custom voice, change the voice in a recording, or use a live voice changer. It runs on Windows and Linux, including Ubuntu servers, with a browser interface for training and conversion and a separate interface for live use. It's free and open source under the MIT license.
Applio is a local AI voice conversion suite for musicians making AI covers, streamers changing their voice live, and creators working with speech. It converts recordings or microphone input into another voice using community models or models you train yourself. Its software uses the MIT license.