CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
Language support includes Chinese, English, Japanese, Korean, German, Spanish, French, Italian and Russian. It also handles Chinese dialects and accents such as Guangdong, Minnan, Sichuan and Shanghai. Spoken instructions can control the output language or dialect, emotion, speed and volume.
Pronunciation is adjustable through Chinese Pinyin and English CMU phonemes, useful when a name or word needs a specific reading. The system also reads numbers, special symbols and formatted text. Its streaming mode accepts text as it arrives and emits audio before the full response is complete, which suits applications that need to start speaking promptly.
The Python project is open source under Apache 2.0. It includes downloadable pretrained models, a browser demo and training scripts for adapting speech generation. Docker deployment supports NVIDIA GPUs, with FastAPI and gRPC interfaces for connecting other applications. vLLM support provides another inference backend, and NVIDIA TensorRT-LLM can accelerate the CosyVoice2 speech model.
Claim this page and we'll verify you by hand. CosyVoice gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find CosyVoice?Promote it
Something wrong or outdated on this page?
7.2KUpdated 2 years agoApache-2.0
macOS · Linux · Docker · Web#Multilingual#Voice cloning
Zonos is an open-source text-to-speech model for people who want to generate speech and clone voices on their own hardware. It can match a speaker from a short reference recording, with controls for delivery and emotion. The code uses the Apache 2.0 license.
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
19.4KUpdated 10 months agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input#Voice cloning
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.