Favicon of CosyVoice

CosyVoice

Open-source text-to-speech software for local voice cloning and streaming speech generation, with Apache 2.0 licensing and NVIDIA GPU deployment.

CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.

Language support includes Chinese, English, Japanese, Korean, German, Spanish, French, Italian and Russian. It also handles Chinese dialects and accents such as Guangdong, Minnan, Sichuan and Shanghai. Spoken instructions can control the output language or dialect, emotion, speed and volume.

Pronunciation is adjustable through Chinese Pinyin and English CMU phonemes, useful when a name or word needs a specific reading. The system also reads numbers, special symbols and formatted text. Its streaming mode accepts text as it arrives and emits audio before the full response is complete, which suits applications that need to start speaking promptly.

The Python project is open source under Apache 2.0. It includes downloadable pretrained models, a browser demo and training scripts for adapting speech generation. Docker deployment supports NVIDIA GPUs, with FastAPI and gRPC interfaces for connecting other applications. vLLM support provides another inference backend, and NVIDIA TensorRT-LLM can accelerate the CosyVoice2 speech model.

Similar to CosyVoice