Favicon of Parler-TTS

Parler-TTS

An open-source Python text-to-speech library that runs locally on CPU or CUDA GPUs and controls voice style through text descriptions.

Parler-TTS generates speech locally, with text descriptions that control how the voice sounds. It's a Python library for developers building speech into applications and researchers who want to train or adapt a TTS model. The library uses the Apache 2.0 license and can run on CPU or CUDA GPUs, with support for Apple Silicon.

The words to speak and the description of the speaker are separate inputs. A description can specify gender, pitch, speaking speed and delivery, as well as recording qualities such as background noise and reverberation. Punctuation also affects pauses and phrasing. This gives developers control over both the voice and its acoustic setting without reducing every choice to a fixed voice preset.

Named speakers help maintain a consistent voice across separate generations. The supplied models include speakers such as Laura, Jon and Gary, alongside the ability to generate voices from descriptive prompts. Parler-TTS Mini and Parler-TTS Large provide different model sizes; both use training drawn from audiobook recordings.

The project publishes model weights, datasets and preprocessing code alongside its inference and training code. Developers can fine-tune it on their own speech data, including a dataset from a single speaker. It also supports streaming speech generation and performance options through SDPA, Flash Attention 2 and torch.compile.

Similar to Parler-TTS