Favicon of XTTS v2

XTTS v2

Local text-to-speech model with voice cloning and streaming audio, available through Coqui TTS on Linux, macOS and Windows.

XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.

Voice cloning is its main draw. You can use a recorded voice for multilingual speech generation, and Coqui TTS can save cloned voices for reuse without requiring the reference audio each time. XTTS also supports streaming audio, so applications can begin playing speech before the full output is ready. Generated speech can be saved as a WAV file.

The Python library provides access for application development, while a command-line interface supports direct speech generation. Coqui TTS also includes a self-hosted server and Docker images. The toolkit works on Linux, macOS and Windows, with CPU execution and PyTorch backends for NVIDIA CUDA or AMD ROCm hardware.

XTTS fine-tuning code is available for projects that need to adapt the model. The model supports multilingual generation from short reference clips, with 17 languages listed in its original model card.

The maintained Coqui TTS repository is a fork of coqui-ai/TTS. Its library code uses Mozilla Public License 2.0. XTTS-v2 weights use the Coqui Public Model License, which allows only noncommercial use of the model and its outputs.

Similar to XTTS v2