Favicon of Tortoise TTS

Tortoise TTS

Open-source text-to-speech software generates custom voices locally on NVIDIA GPUs or Apple Silicon, with Docker support and an Apache 2.0 license.

Tortoise TTS is a local text-to-speech system for developers and creators who want speech with varied voices and natural pacing. It uses reference audio clips to guide a custom voice, with an emphasis on expressive rhythm and intonation.

You can generate a short phrase in one or several voices, or turn a longer text into a complete audio file. For longer passages, it produces separate sentence clips as well as a combined recording. If a sentence comes out poorly, you can regenerate that clip without replacing the whole passage.

The software includes a Python API for adding speech generation to other applications, plus socket streaming for delivering audio as it's generated. Its speech model combines autoregressive and diffusion decoders. Generation speed depends on the hardware and settings; acceleration options include DeepSpeed, caching and reduced precision, though DeepSpeed doesn't work on Apple Silicon.

Tortoise runs on your own hardware, with support for NVIDIA GPUs, Windows and Docker. It also supports M1 and M2 Macs running macOS 13 or later. Local inference processes text and voice samples on your machine. A separate Hugging Face Spaces demo runs on hosted hardware and requires a GPU rather than a CPU-only Space.

The project is open source under Apache 2.0. Hugging Face hosts its model weights.

Similar to Tortoise TTS