Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
Voice cloning doesn't require training on each speaker. It uses a reference recording to synthesize text in that voice, with support for Chinese and English, cloning across those languages, and speech that switches between them. Its web interface accepts uploaded audio or a recording made directly in the interface.
For voice creation, you can control gender, pitch, and speaking rate without supplying a speaker recording. This gives users a separate way to create virtual voices when they don't need to reproduce a particular person. Both voice cloning and voice creation are available through the web interface; a command-line interface also supports speech generation.
The system runs on Linux and Windows, with macOS support through Metal Performance Shaders and a CPU fallback. NVIDIA GPU server deployment uses Triton Inference Serving and TensorRT-LLM.
Its speech model is built on Qwen2.5. It reconstructs audio from speech codes predicted by the language model, without a separate generation model such as flow matching.
Claim this page and we'll verify you by hand. Spark-TTS gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Spark-TTS?Promote it
Something wrong or outdated on this page?
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
7.2KUpdated 2 years agoApache-2.0
macOS · Linux · Docker · Web#Multilingual#Voice cloning
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
14.9KUpdated 2 years agoApache-2.0
macOS · Windows · Docker#Streaming inference#Voice cloning
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
Zonos is an open-source text-to-speech model for people who want to generate speech and clone voices on their own hardware. It can match a speaker from a short reference recording, with controls for delivery and emotion. The code uses the Apache 2.0 license.
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
Tortoise TTS is a local text-to-speech system for developers and creators who want speech with varied voices and natural pacing. It uses reference audio clips to guide a custom voice, with an emphasis on expressive rhythm and intonation.