Favicon of MeloTTS

MeloTTS

A local text-to-speech library with real-time CPU inference, multilingual voices and custom dataset training. Free under the MIT license.

MeloTTS is a Python text-to-speech library for developers who want to generate speech locally, including on machines without a dedicated GPU. It supports real-time inference on a CPU. Its language and accent choices make it relevant for applications that need spoken output across different audiences.

English voices include American, British, Indian and Australian accents, alongside a default English voice. The library also supports Spanish, French, Chinese, Japanese and Korean. Its Chinese speaker can read text that mixes Chinese and English, a useful distinction for content containing English names or phrases within Chinese sentences.

Developers can integrate speech generation through a Python API, while a web interface and command-line interface provide other ways to use the library. MeloTTS also supports training on a custom dataset, so projects can work beyond the supplied voices. Model cards are available on HuggingFace.

The project is free and open source under the MIT license, which permits commercial and non-commercial use. Its implementation builds on TTS, VITS, VITS2 and Bert-VITS2.

Similar to MeloTTS