Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.
It supports Original Chatterbox for English speech with emotion control, Chatterbox Multilingual for languages including Arabic, Chinese, French, Japanese and Spanish, and Chatterbox-Turbo for faster generation with less compute and GPU memory. Turbo can include reactions such as laughter, coughing and chuckling in generated speech. You can switch models without restarting the server, so the same interface and API cover different language and speed needs.
Built-in voices let you generate speech without supplying a recording. Voice cloning accepts WAV or MP3 reference audio to produce a similar voice, while fixed generation seeds help keep results consistent across requests. For longer work, the server splits text at sentence boundaries and joins the generated audio into a single file, including book-length narration. It also supports streamed audio output.
The server runs on Windows, Linux and macOS, with Docker deployment available. It can use NVIDIA GPUs through CUDA, AMD GPUs through ROCm, Apple Silicon through MPS, or a CPU. Speech generation runs on the machine hosting the server; the optional Google Colab demo runs in Google's cloud. Windows also has a portable mode that keeps Python and dependencies in a movable, self-contained folder.
Claim this page and we'll verify you by hand. Chatterbox TTS Server gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Chatterbox TTS Server?Promote it
Something wrong or outdated on this page?
857Updated 2 years agoAGPL-3.0
macOS · Windows · Linux · Docker#Multilingual#ONNX#OpenAI-compatible API
OpenedAI Speech is a self-hosted text-to-speech server for developers who want local speech generation in apps built around OpenAI's speech API. The project is archived and no longer maintained. It's open source under AGPL-3.0, and it generates audio on your own hardware without an OpenAI API key.
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
Coqui TTS (idiap fork) is a local text-to-speech library for developers and speech researchers who want pretrained voices or tools to train their own models. It builds on coqui-ai/TTS, continuing the original unmaintained project. The Python toolkit is open source under the Mozilla Public License 2.0 (MPL-2.0).