MeloTTS is a Python text-to-speech library for developers who want to generate speech locally, including on machines without a dedicated GPU. It supports real-time inference on a CPU. Its language and accent choices make it relevant for applications that need spoken output across different audiences.
English voices include American, British, Indian and Australian accents, alongside a default English voice. The library also supports Spanish, French, Chinese, Japanese and Korean. Its Chinese speaker can read text that mixes Chinese and English, a useful distinction for content containing English names or phrases within Chinese sentences.
Developers can integrate speech generation through a Python API, while a web interface and command-line interface provide other ways to use the library. MeloTTS also supports training on a custom dataset, so projects can work beyond the supplied voices. Model cards are available on HuggingFace.
The project is free and open source under the MIT license, which permits commercial and non-commercial use. Its implementation builds on TTS, VITS, VITS2 and Bert-VITS2.
Claim this page and we'll verify you by hand. MeloTTS gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MeloTTS?Promote it
Something wrong or outdated on this page?
5.6KUpdated 2 years agoApache-2.0
macOS#Hugging Face integration
Parler-TTS generates speech locally, with text descriptions that control how the voice sounds. It's a Python library for developers building speech into applications and researchers who want to train or adapt a TTS model. The library uses the Apache 2.0 license and can run on CPU or CUDA GPUs, with support for Apple Silicon.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
9.1KUpdated 1 year agoApache-2.0
macOS · Windows#Batch processing#Multilingual#ONNX
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
Kokoro is a text-to-speech model and inference library for developers who want to generate speech on their own hardware or servers. Its compact Kokoro-82M model suits personal projects and production applications, with Apache 2.0 licensing for both the library and model weights.
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.