Favicon of MOSS TTS

MOSS TTS

Open-source text-to-speech models under Apache 2.0 for multilingual narration, voice cloning and dialogue, with a CPU and browser variant.

Screenshot of MOSS TTS website

MOSS TTS is an open-source family of speech and sound generation models from MOSI.AI and OpenMOSS, licensed under Apache 2.0. It's aimed at developers building narration, dubbing and voice applications. MOSS-TTS-Nano provides a CPU or browser option for speech synthesis and voice cloning.

The main MOSS-TTS model focuses on long recordings with a consistent voice. It can clone a voice from reference speech without training a separate model for that speaker, continue existing audio, and mix languages within a passage. Supported languages include English, Chinese, Cantonese, Japanese, Hindi and Arabic.

Pronunciation controls accept Pinyin and phonemes, while duration and pause controls let creators adjust the timing of spoken content. These capabilities suit scripts where names, pronunciation or pacing need more attention than a basic text-to-speech system allows. The Local Transformer variant supports native stereo audio input and output.

Separate models cover different kinds of audio work. MOSS-TTSD generates dialogue with multiple speakers for podcasts and dubbing. MOSS-VoiceGenerator creates voices and speaking styles from text descriptions without reference recordings, and can supply voices for downstream speech synthesis.

MOSS-TTS-Realtime produces streaming speech for voice agents and uses prior text and user audio to maintain context across turns. MOSS-SoundEffect generates environmental sounds, human actions and musical fragments, with control over their duration.

Similar to MOSS TTS