
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
Its style diffusion approach produces variation in pauses, pace and emotion, rather than giving every reading the same delivery. Training also uses WavLM, a pretrained speech model, to assess generated audio and improve naturalness.
Pretrained models cover single-speaker synthesis with LJSpeech and multispeaker synthesis with LibriTTS. The LibriTTS model can adapt to speakers it didn't encounter during training, carrying over emotion and aspects of the reference recording's acoustic environment. Examples also demonstrate long-form narration and emotional style transfer to text that doesn't itself express an emotion.
You can train or fine-tune models on your own speech data, including other languages when paired with a suitable PL-BERT model. A multilingual PL-BERT is available. Training and fine-tuning can use NVIDIA GPUs, with a single-GPU path available for fine-tuning.
The project requires public disclosure that pretrained-voice output is synthesized unless you have permission to use that voice. Reference inference uses a GPL-licensed dependency; the repository code is MIT.
Claim this page and we'll verify you by hand. StyleTTS 2 gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find StyleTTS 2?Promote it
Something wrong or outdated on this page?
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.