MOSS-TTS 1.5: dual-GPU Kaggle setup and voice tests

Learn to set up MOSS-TTS 1.5 on Kaggle with two T4 GPUs, use its cloning modes, and assess voice samples across ten languages.

Player not loading? Watch on YouTube

This tutorial covers MOSS-TTS 1.5, an open source text-to-speech model from OpenMOSS. The presenter describes an 8-billion-parameter release with an Apache 2.0 license and support for 31 languages, compared with 20 in version 1. The walkthrough uses a hosted Kaggle notebook rather than demonstrating how to run models locally.

Setup starts with the notebook linked from the presenter's GitHub repository. Viewers select the GPU T4* 2 accelerator, run all cells, and open the public Gradio link. The presenter estimates five to seven minutes for installation and model downloads. He states that local execution requires 24 GB of GPU VRAM and that a single T4 cannot handle this notebook's model.

The interface accepts text, a language tag or automatic selection, and a reference recording. It includes clone, continuation, and continuation plus clone modes. Continuation requires the reference audio's transcript before the new text. Advanced controls include speed, duration, and text and audio temperatures; the release also supports explicit pause tags.

Tests cover celebrity references and expressive voices, followed by speech in ten languages. The presenter praises the transfer of tone and hesitation, but these are his assessments. He uses Google Translate for the multilingual text and asks native speakers to judge pronunciation and delivery, leaving those results unverified.