
OpenVoice is an open-source voice cloning tool that uses a short recording to reproduce a speaker's voice in generated speech. It's for developers and creators who need a recognizable voice across languages, with control over how that voice sounds. The Python project is MIT licensed for commercial use.
Voice identity and delivery can be controlled separately. OpenVoice reproduces the speaker's vocal tone while allowing changes to emotion, accent, rhythm, pauses and intonation. That gives users room to vary the performance without choosing a different reference speaker. Demonstrated styles include happy and sad speech, as well as British, Indian and Australian accents.
Its cross-language cloning doesn't require the reference recording and generated speech to share a language. It can also clone voices across languages absent from its multilingual speaker training data. OpenVoice V2 natively supports English, Spanish, French, Chinese, Japanese and Korean. This makes it relevant to multilingual narration and speech projects that need to carry one speaker's voice into another language.
OpenVoice's approach emphasizes computational efficiency alongside voice and style control. Developers can use its source code in their own projects under the MIT license. MyShell also uses the model to provide instant voice cloning on its hosted platform.
Claim this page with an email at research.myshell.ai. OpenVoice gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find OpenVoice?Promote it
Something wrong or outdated on this page?
26.6KUpdated 2 months agoMIT
Linux#Multilingual#Voice cloning#Voice conversion
Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
32.9KUpdated 2 weeks ago
#Batch processing#Multilingual#Multimodal input
Fish Speech, currently featuring Fish Audio S2 Pro, is a self-hosted text-to-speech system for creators producing narration and developers building voice applications. It combines voice cloning with control over emotion and delivery within a script. Code and model weights use the custom FISH AUDIO RESEARCH LICENSE.
8.4KUpdated 4 months agoApache-2.0
#Batch processing#Hugging Face integration#Multilingual
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
Higgs Audio is a family of text-to-speech models from Boson AI for developers building narration and conversational audio. Higgs TTS 2 can adapt pacing and intonation to the text and generate dialogue with distinct speakers across multiple languages.
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.