Favicon of TTS WebUI

TTS WebUI

A self-hosted text-to-speech and audio generation interface with model extensions, Docker support and an OpenAI-compatible speech API. MIT licensed.

TTS WebUI brings local text-to-speech, music generation and audio processing into one browser interface. It's for people creating spoken audio or music, and for developers who want to add speech to a self-hosted chat app. The interface combines Gradio and React, with extensions that let you choose which audio models to use.

Speech models include Bark, Tortoise and StyleTTS2. Extensions add alternatives such as Piper TTS, Kokoro, XTTSv2, CosyVoice and GPT-SoVITS; these aren't part of the default installation. MusicGen and Stable Audio cover music and audio generation, while ACE-Step is available through an extension.

The audio tools go beyond generation. RVC handles voice conversion, Whisper supports speech transcription, and Demucs and Audio Separator separate audio sources. Resemble Enhance and the PyRNNoise extension provide audio enhancement and noise reduction. The app also manages generated audio files and their metadata.

Installing and enabling the OpenAI-compatible speech API extension connects it to Silly Tavern and OpenWebUI. Text Generation WebUI has a separate integration. These connections let chat interfaces use speech generated by your own audio server.

You can run it locally or in Docker, and it has a Windows launcher. Docker supports NVIDIA CUDA GPUs. Models download to the host for local generation; the built-in MiniMax Cloud TTS option uses a cloud service instead.

The code is open source under MIT. Dependencies and model weights have separate licenses, including noncommercial terms for MusicGen and AudioGen weights.

Similar to TTS WebUI