TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.
Speech models include Bark, Tortoise and StyleTTS2. Extensions add alternatives such as Piper TTS, Kokoro, XTTSv2, CosyVoice and GPT-SoVITS; these aren't part of the default installation. MusicGen and Stable Audio cover music and audio generation, while ACE-Step is available through an extension.
The audio tools go beyond generation. RVC handles voice conversion, Whisper supports speech transcription, and Demucs and Audio Separator separate audio sources. Resemble Enhance and the PyRNNoise extension provide audio enhancement and noise reduction. The app also manages generated audio files and their metadata.
Installing and enabling the OpenAI-compatible speech API extension connects it to Silly Tavern and OpenWebUI. Text Generation WebUI has a separate integration. These connections let chat interfaces use speech generated by your own audio server.
You can run it locally or in Docker, and it has a Windows launcher. Docker supports NVIDIA CUDA GPUs. Models download to the host for local generation; the built-in MiniMax Cloud TTS option uses a cloud service instead.
The code is open source under MIT. Dependencies and model weights have separate licenses, including noncommercial terms for MusicGen and AudioGen weights.
Claim this page and we'll verify you by hand. TTS WebUI gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find TTS WebUI?Promote it
Something wrong or outdated on this page?
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
1.5KUpdated 4 months agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#OpenAI-compatible API
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
9.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#llama.cpp backend#Multimodal input
2.4KUpdated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker · Web#Hugging Face integration
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
Xinference serves language, speech and multimodal models through a shared API on your own computer or servers. It's an open source platform under Apache 2.0 for developers and researchers who want to build applications around models they host. You can also deploy it on cloud infrastructure.
AllTalk TTS generates speech on your own computer. The project recommends v2 for most users; the saved documentation below describes v1, built on Coqui TTS and XTTSv2 models. It's for people adding voices to AI conversations or producing spoken audio from longer texts. It runs as a standalone application or alongside Text-generation-webui, with support for Windows, Linux and macOS.