IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
It supports Chinese, English, Japanese, Spanish and Arabic, including speech in a different language from the reference clip. Emotion can come from an audio sample, a written description or the script itself. You can also set emotion intensities directly and adjust how strongly an emotional reference affects the result.
Speaking speed is adjustable. For pronunciation, it accepts Chinese Pinyin, English CMU phonemes and Japanese Kana, giving users a way to specify how particular words should sound. These controls suit projects where matching a voice alone doesn't provide enough control over the delivery.
The Python software runs locally with a browser interface, and its Python API supports integration into applications. Linux and Windows are supported, with NVIDIA CUDA acceleration available. Lower-precision inference reduces GPU memory use with a small quality tradeoff, and vLLM supports server deployment. Models come from HuggingFace or ModelScope; the first run also downloads some smaller models and example audio.
The model uses the custom bilibili Model Use License Agreement, which includes use restrictions and requires separate written permission for organizations above specified size or revenue thresholds.
Claim this page and we'll verify you by hand. IndexTTS gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find IndexTTS?Promote it
Something wrong or outdated on this page?
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
7.2KUpdated 2 years agoApache-2.0
macOS · Linux · Docker · Web#Multilingual#Voice cloning
19.4KUpdated 10 months agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input#Voice cloning
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
Zonos is an open-source text-to-speech model for people who want to generate speech and clone voices on their own hardware. It can match a speaker from a short reference recording, with controls for delivery and emotion. The code uses the Apache 2.0 license.
Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.