Favicon of AllTalk TTS

AllTalk TTS

Local text-to-speech software built on Coqui TTS, with XTTSv2 models, voice fine-tuning and integrations for SillyTavern and Text-generation-webui.

AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.

Voice control is a central feature. You can use local or custom XTTSv2 models and fine-tune a model on a chosen voice to improve its reproduction. For character conversations, AllTalk can assign separate voices to dialogue and narration, with controls for how it treats text that doesn't clearly belong to either. Its bulk speech generator and editor can turn hours of text into a single audio file or read text back to you.

AllTalk connects with SillyTavern and KoboldCPP as well as Text-generation-webui. A JSON API lets other applications request speech, so it can serve as the speech component of a self-hosted AI setup. A browser interface provides settings and built-in documentation, and the server can be made accessible to other computers on your network.

Low VRAM mode helps when GPU memory is limited or shared with a local LLM. DeepSpeed acceleration is available for NVIDIA GPUs on Windows and Linux. Speech generation uses local model files; AllTalk downloads missing models from Hugging Face. The application uses AGPL-3.0; model weights have separate terms. It can delete older generated WAV files. The documented Docker build is an older third-party version.

Similar to AllTalk TTS