Favicon of Coqui TTS (idiap fork)

Coqui TTS (idiap fork)

Open-source text-to-speech toolkit for local speech generation, voice cloning and model training on Linux, macOS and Windows, licensed under MPL-2.0.

Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).

Speech generation runs on your own machine or a self-hosted server. It supports Linux, macOS and Windows, with Docker images for server use. Its PyTorch backend can use a CPU, NVIDIA GPUs through CUDA or AMD GPUs through ROCm.

The model selection includes XTTS, VITS, YourTTS, Tortoise and Bark, plus Fairseq models for broad language coverage. Depending on the model, you can generate speech in multiple languages, choose a preset speaker or clone a voice from reference audio. XTTS also supports streaming speech, and the toolkit can save cloned voices for reuse.

Voice conversion changes an existing recording to match a target voice using FreeVC, kNN-VC or OpenVoice. It can also combine a single-speaker speech model with FreeVC to produce speech in a cloned voice.

For custom voices and research, the library supports training and fine-tuning alongside dataset analysis and curation. Python and command-line interfaces cover pretrained and custom models, while training logs integrate with TensorBoard. The lower-level synthesis API provides additional outputs such as timestamps.

Similar to Coqui TTS (idiap fork)