
MOSS TTS is an open-source family of speech and sound generation models from MOSI.AI and OpenMOSS, licensed under Apache 2.0. It's aimed at developers building narration, dubbing and voice applications. MOSS-TTS-Nano provides a CPU or browser option for speech synthesis and voice cloning.
The main MOSS-TTS model focuses on long recordings with a consistent voice. It can clone a voice from reference speech without training a separate model for that speaker, continue existing audio, and mix languages within a passage. Supported languages include English, Chinese, Cantonese, Japanese, Hindi and Arabic.
Pronunciation controls accept Pinyin and phonemes, while duration and pause controls let creators adjust the timing of spoken content. These capabilities suit scripts where names, pronunciation or pacing need more attention than a basic text-to-speech system allows. The Local Transformer variant supports native stereo audio input and output.
Separate models cover different kinds of audio work. MOSS-TTSD generates dialogue with multiple speakers for podcasts and dubbing. MOSS-VoiceGenerator creates voices and speaking styles from text descriptions without reference recordings, and can supply voices for downstream speech synthesis.
MOSS-TTS-Realtime produces streaming speech for voice agents and uses prior text and user audio to maintain context across turns. MOSS-SoundEffect generates environmental sounds, human actions and musical fragments, with control over their duration.
Claim this page with an email at mosi.cn. MOSS TTS gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find MOSS TTS?Promote it
Something wrong or outdated on this page?
26.6KUpdated 2 months agoMIT
Linux#Multilingual#Voice cloning#Voice conversion
Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
24.3KUpdated 2 days ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
19.4KUpdated 10 months agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input#Voice cloning
15.3KUpdated 2 weeks agoMIT
Docker · Web#Multilingual#Voice cloning
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.