Favicon of Spark-TTS

Spark-TTS

Open-source local text-to-speech built on Qwen2.5, with Chinese and English voice cloning, adjustable voices, and an Apache 2.0 license.

Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.

Voice cloning doesn't require training on each speaker. It uses a reference recording to synthesize text in that voice, with support for Chinese and English, cloning across those languages, and speech that switches between them. Its web interface accepts uploaded audio or a recording made directly in the interface.

For voice creation, you can control gender, pitch, and speaking rate without supplying a speaker recording. This gives users a separate way to create virtual voices when they don't need to reproduce a particular person. Both voice cloning and voice creation are available through the web interface; a command-line interface also supports speech generation.

The system runs on Linux and Windows, with macOS support through Metal Performance Shaders and a CPU fallback. NVIDIA GPU server deployment uses Triton Inference Serving and TensorRT-LLM.

Its speech model is built on Qwen2.5. It reconstructs audio from speech codes predicted by the language model, without a separate generation model such as flow matching.

Similar to Spark-TTS