Parler-TTS generates speech locally, with text descriptions that control how the voice sounds. It's a Python library for developers building speech into applications and researchers who want to train or adapt a TTS model. The library uses the Apache 2.0 license and can run on CPU or CUDA GPUs, with support for Apple Silicon.
The words to speak and the description of the speaker are separate inputs. A description can specify gender, pitch, speaking speed and delivery, as well as recording qualities such as background noise and reverberation. Punctuation also affects pauses and phrasing. This gives developers control over both the voice and its acoustic setting without reducing every choice to a fixed voice preset.
Named speakers help maintain a consistent voice across separate generations. The supplied models include speakers such as Laura, Jon and Gary, alongside the ability to generate voices from descriptive prompts. Parler-TTS Mini and Parler-TTS Large provide different model sizes; both use training drawn from audiobook recordings.
The project publishes model weights, datasets and preprocessing code alongside its inference and training code. Developers can fine-tune it on their own speech data, including a dataset from a single speaker. It also supports streaming speech generation and performance options through SDPA, Flash Attention 2 and torch.compile.
Claim this page and we'll verify you by hand. Parler-TTS gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Parler-TTS?Promote it
Something wrong or outdated on this page?
9.1KUpdated 1 year agoApache-2.0
macOS · Windows#Batch processing#Multilingual#ONNX
Kokoro is a text-to-speech model and inference library for developers who want to generate speech on their own hardware or servers. Its compact Kokoro-82M model suits personal projects and production applications, with Apache 2.0 licensing for both the library and model weights.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
7.7KUpdated 2 years agoMIT
Docker · Web#Multilingual
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
MeloTTS is a Python text-to-speech library for developers who want to generate speech locally, including on machines without a dedicated GPU. It supports real-time inference on a CPU. Its language and accent choices make it relevant for applications that need spoken output across different audiences.