StyleTTS 2

An open-source text-to-speech model you can run locally, with MIT-licensed Python code, pretrained English voices and adaptation to unfamiliar speakers.

Screenshot of StyleTTS 2 website

StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.

Its style diffusion approach produces variation in pauses, pace and emotion, rather than giving every reading the same delivery. Training also uses WavLM, a pretrained speech model, to assess generated audio and improve naturalness.

Pretrained models cover single-speaker synthesis with LJSpeech and multispeaker synthesis with LibriTTS. The LibriTTS model can adapt to speakers it didn't encounter during training, carrying over emotion and aspects of the reference recording's acoustic environment. Examples also demonstrate long-form narration and emotional style transfer to text that doesn't itself express an emotion.

You can train or fine-tune models on your own speech data, including other languages when paired with a suitable PL-BERT model. A multilingual PL-BERT is available. Training and fine-tuning can use NVIDIA GPUs, with a single-GPU path available for fine-tuning.

The project requires public disclosure that pretrained-voice output is synthesized unless you have permission to use that voice. Reference inference uses a GPL-licensed dependency; the repository code is MIT.

Similar to StyleTTS 2