Seed-VC changes recorded speech or singing to sound like a voice supplied in a short reference clip, without training a separate model for that speaker. It runs locally on Windows, Linux and Apple Silicon Macs, with uses in audio production, live streaming and online meetings. The project is archived and no longer maintained.
A reference recording of 1 to 30 seconds is enough for voice conversion. Separate models handle recorded speech, singing and live audio, so users can choose according to the material they're working with. Singing conversion includes pitch control, while the voice and accent model can also change accent and emotion or anonymize speech into an average voice without a reference recording.
The software includes a local browser interface for file conversion and a desktop interface for live audio. A GPU is strongly recommended for live use, and other GPU workloads can slow conversion. Model weights download from Hugging Face on first use; the audio conversion runs on your own hardware.
For a particular speaker, optional fine-tuning can improve voice similarity using as little as one utterance per speaker, though it may reduce word accuracy. The file conversion and live interfaces accept custom trained models. The singing model uses OpenAI Whisper to extract speech content and NVIDIA BigVGAN to generate audio. Seed-VC is open source under GPL-3.0.
Claim this page and we'll verify you by hand. Seed-VC gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Seed-VC?Promote it
Something wrong or outdated on this page?
38.6KUpdated 2 months agoMIT
Windows · Linux · Web#Hugging Face integration#ONNX#Voice conversion
RVC WebUI is a local AI voice conversion tool for people who want to train a custom voice, change the voice in a recording, or use a live voice changer. It runs on Windows and Linux, including Ubuntu servers, with a browser interface for training and conversion and a separate interface for live use. It's free and open source under the MIT license.
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
3.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Web#Batch processing#Voice conversion
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.
Applio is a local AI voice conversion suite for musicians making AI covers, streamers changing their voice live, and creators working with speech. It converts recordings or microphone input into another voice using community models or models you train yourself. Its software uses the MIT license.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.