Favicon of IndexTTS

IndexTTS

Local text-to-speech software clones voices from one audio clip, supports five languages, and provides separate controls for emotion and speaking speed.

IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.

It supports Chinese, English, Japanese, Spanish and Arabic, including speech in a different language from the reference clip. Emotion can come from an audio sample, a written description or the script itself. You can also set emotion intensities directly and adjust how strongly an emotional reference affects the result.

Speaking speed is adjustable. For pronunciation, it accepts Chinese Pinyin, English CMU phonemes and Japanese Kana, giving users a way to specify how particular words should sound. These controls suit projects where matching a voice alone doesn't provide enough control over the delivery.

The Python software runs locally with a browser interface, and its Python API supports integration into applications. Linux and Windows are supported, with NVIDIA CUDA acceleration available. Lower-precision inference reduces GPU memory use with a small quality tradeoff, and vLLM supports server deployment. Models come from HuggingFace or ModelScope; the first run also downloads some smaller models and example audio.

The model uses the custom bilibili Model Use License Agreement, which includes use restrictions and requires separate written permission for organizations above specified size or revenue thresholds.

Similar to IndexTTS