IndexTTS V26 tutorial: saved voices and 8-person dialogue

Learn to save voices, assign up to eight dialogue roles and control emotion in IndexTTS V26, with reported performance near real time on 16GB VRAM.

Player not loading? Watch on YouTube

This Chinese-language tutorial demonstrates a third-party V26 package built around IndexTTS, distributed through the presenter’s download link. This IndexTTS V26 tutorial covers local AI speech synthesis with saved voice references and dialogue for up to eight speakers. The presenter shows how to upload reference audio, name and save a voice, then load it from a dropdown. Dialogue scripts use a role name followed by a colon and spoken text. After parsing the roles, users assign each one a voice and emotion before generating the conversation.

The presenter reports a 2-10 times inference speed improvement. Short passages still spend much of their processing time in GPT generation; longer passages reportedly approach real-time synthesis with 16GB of VRAM. The tutorial also explains sampling settings and warns that setting the maximum token count too low can truncate audio.

Emotion can follow the voice reference, come from separate audio or use manual vector controls. In the presenter's examples, fear transfers less clearly than other emotions, prompting an increase in emotion weight. Experimental text-based control uses a fine-tuned Qwen3 0.6B model, according to the speaker, and the interface warns that results remain unstable. This control applies one overall emotion throughout the output.