Tortoise TTS is a local text-to-speech system for developers and creators who want speech with varied voices and natural pacing. It uses reference audio clips to guide a custom voice, with an emphasis on expressive rhythm and intonation.
You can generate a short phrase in one or several voices, or turn a longer text into a complete audio file. For longer passages, it produces separate sentence clips as well as a combined recording. If a sentence comes out poorly, you can regenerate that clip without replacing the whole passage.
The software includes a Python API for adding speech generation to other applications, plus socket streaming for delivering audio as it's generated. Its speech model combines autoregressive and diffusion decoders. Generation speed depends on the hardware and settings; acceleration options include DeepSpeed, caching and reduced precision, though DeepSpeed doesn't work on Apple Silicon.
Tortoise runs on your own hardware, with support for NVIDIA GPUs, Windows and Docker. It also supports M1 and M2 Macs running macOS 13 or later. Local inference processes text and voice samples on your machine. A separate Hugging Face Spaces demo runs on hosted hardware and requires a GPU rather than a CPU-only Space.
The project is open source under Apache 2.0. Hugging Face hosts its model weights.
Claim this page and we'll verify you by hand. Tortoise TTS gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Tortoise TTS?Promote it
Something wrong or outdated on this page?
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
7.2KUpdated 2 years agoApache-2.0
macOS · Linux · Docker · Web#Multilingual#Voice cloning
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
19.4KUpdated 10 months agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input#Voice cloning
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
Zonos is an open-source text-to-speech model for people who want to generate speech and clone voices on their own hardware. It can match a speaker from a short reference recording, with controls for delivery and emotion. The code uses the Apache 2.0 license.
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.