RVC WebUI is a local AI voice conversion tool for people who want to train a custom voice, change the voice in a recording, or use a live voice changer. It runs on Windows and Linux, including Ubuntu servers, with a browser interface for training and conversion and a separate interface for live use. It's free and open source under the MIT license.
Its retrieval approach uses examples from the trained voice to reduce traces of the original speaker's vocal tone in the output. You can train with a small dataset; the project recommends clean voice recordings with little background noise. The base models use the VCTK speech dataset, and training can work on relatively modest graphics cards.
Model merging lets you blend trained models to alter a voice's tone. For work with songs, pymss/MSST models separate vocals from accompaniment, while RMVPE extracts vocal pitch for conversion. RVC accepts voice models in .pth format with accompanying .index files.
NVIDIA GPUs use CUDA acceleration. AMD and Intel hardware can use DirectML on Windows; the documented Linux path for that hardware uses the CPU. Training and conversion run on your own machine or server, with model assets available to download from Hugging Face.
The live voice changer supports low-latency audio. Performance depends heavily on hardware and driver support, including ASIO input and output devices.
Claim this page and we'll verify you by hand. RVC WebUI gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find RVC WebUI?Promote it
Something wrong or outdated on this page?
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
3.9KUpdated 1 year agoGPL-3.0
macOS · Windows · Linux · Web#Hugging Face integration#Streaming inference#Voice conversion
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
3.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Web#Batch processing#Voice conversion
3.3KUpdated 3 weeks agoMIT
Windows · Docker · Web#OpenAI-compatible API
TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
Seed-VC changes recorded speech or singing to sound like a voice supplied in a short reference clip, without training a separate model for that speaker. It runs locally on Windows, Linux and Apple Silicon Macs, with uses in audio production, live streaming and online meetings. The project is archived and no longer maintained.
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.
Applio is a local AI voice conversion suite for musicians making AI covers, streamers changing their voice live, and creators working with speech. It converts recordings or microphone input into another voice using community models or models you train yourself. Its software uses the MIT license.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.