AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.
Voice control is a central feature. You can use local or custom XTTSv2 models and fine-tune a model on a chosen voice to improve its reproduction. For character conversations, AllTalk can assign separate voices to dialogue and narration, with controls for how it treats text that doesn't clearly belong to either. Its bulk speech generator and editor can turn hours of text into a single audio file or read text back to you.
AllTalk connects with SillyTavern and KoboldCPP as well as Text-generation-webui. A JSON API lets other applications request speech, so it can serve as the speech component of a self-hosted AI setup. A browser interface provides settings and built-in documentation, and the server can be made accessible to other computers on your network.
Low VRAM mode helps when GPU memory is limited or shared with a local LLM. DeepSpeed acceleration is available for NVIDIA GPUs on Windows and Linux. Speech generation uses local model files; AllTalk downloads missing models from Hugging Face. The application uses AGPL-3.0; model weights have separate terms. It can delete older generated WAV files. The documented Docker build is an older third-party version.
Claim this page and we'll verify you by hand. AllTalk TTS gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find AllTalk TTS?Promote it
Something wrong or outdated on this page?
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
3.8KUpdated 2 days agoMIT
macOS · Windows · Linux · Web#Batch processing#Voice conversion
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
Applio is a local AI voice conversion suite for musicians making AI covers, streamers changing their voice live, and creators working with speech. It converts recordings or microphone input into another voice using community models or models you train yourself. Its software uses the MIT license.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.