XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
Voice cloning is its main draw. You can use a recorded voice for multilingual speech generation, and Coqui TTS can save cloned voices for reuse without requiring the reference audio each time. XTTS also supports streaming audio, so applications can begin playing speech before the full output is ready. Generated speech can be saved as a WAV file.
The Python library provides access for application development, while a command-line interface supports direct speech generation. Coqui TTS also includes a self-hosted server and Docker images. The toolkit works on Linux, macOS and Windows, with CPU execution and PyTorch backends for NVIDIA CUDA or AMD ROCm hardware.
XTTS fine-tuning code is available for projects that need to adapt the model. The model supports multilingual generation from short reference clips, with 17 languages listed in its original model card.
The maintained Coqui TTS repository is a fork of coqui-ai/TTS. Its library code uses Mozilla Public License 2.0. XTTS-v2 weights use the Coqui Public Model License, which allows only noncommercial use of the model and its outputs.
Claim this page and we'll verify you by hand. XTTS v2 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find XTTS v2?Promote it
Something wrong or outdated on this page?
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
14.9KUpdated 2 years agoApache-2.0
macOS · Windows · Docker#Streaming inference#Voice cloning
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
Tortoise TTS is a local text-to-speech system for developers and creators who want speech with varied voices and natural pacing. It uses reference audio clips to guide a custom voice, with an emphasis on expressive rhythm and intonation.