Voice Changer (w-okada), also called VCClient, converts your voice as you speak using AI voice models. It's for people who want live voice conversion on their own computer, including those recording gaming commentary while running demanding software. Processing can stay local.
You can also run voice conversion on a separate computer over a network. This moves the processing load away from the machine running your game or other applications. The app supports Windows, macOS on Apple Silicon M1, and Linux, including x86-64 and ARM systems.
Model support depends on the build you choose. Standard builds support Beatrice, while CUDA builds for Windows and ONNX builds for Windows and M1 Macs also support RVC (Retrieval-based-Voice-Conversion). Beatrice controls include speaker merging, automatic pitch shifting, and pitch and formant adjustments. You can export edited Beatrice models.
For developers building voice conversion into another application, a REST API provides access from custom clients. Google Colab is another supported way to run the processing, but that uses cloud hardware rather than keeping conversion on your own machine; the documented Colab setup requires an ngrok account.
Voice models and included character voices have their own usage terms. Some restrict what converted speech can say or how it can be distributed, and Amitaro requires attribution that makes clear the speech is a converted voice.
Claim this page and we'll verify you by hand. Voice Changer (w-okada) gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Voice Changer (w-okada)?Promote it
Something wrong or outdated on this page?
3.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Web#Batch processing#Voice conversion
Applio is a local AI voice conversion suite for musicians making AI covers, streamers changing their voice live, and creators working with speech. It converts recordings or microphone input into another voice using community models or models you train yourself. Its software uses the MIT license.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
857Updated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API
3.9KUpdated 1 year agoGPL-3.0
macOS · Windows · Linux · Web#Hugging Face integration#Streaming inference#Voice conversion
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
1.5KUpdated 4 months agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).
OpenedAI Speech is a self-hosted text-to-speech server for developers who want local speech generation in apps built around OpenAI's speech API. The project is archived and no longer maintained. It's open source under AGPL-3.0, and it generates audio on your own hardware without an OpenAI API key.
Seed-VC changes recorded speech or singing to sound like a voice supplied in a short reference clip, without training a separate model for that speaker. It runs locally on Windows, Linux and Apple Silicon Macs, with uses in audio production, live streaming and online meetings. The project is archived and no longer maintained.
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.
Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.