Favicon of Chatterbox TTS Server

Chatterbox TTS Server

Self-hosted text-to-speech server runs Chatterbox models on CPU or GPU, with voice cloning, audiobook generation and an OpenAI-compatible API.

Chatterbox TTS Server runs Resemble AI's speech models on your own computer or server, with a browser interface and an OpenAI-compatible API. It's for people producing narration and audiobooks, or developers adding speech to voice agents and other apps. The project is open source under the MIT license.

It supports Original Chatterbox for English speech with emotion control, Chatterbox Multilingual for languages including Arabic, Chinese, French, Japanese and Spanish, and Chatterbox-Turbo for faster generation with less compute and GPU memory. Turbo can include reactions such as laughter, coughing and chuckling in generated speech. You can switch models without restarting the server, so the same interface and API cover different language and speed needs.

Built-in voices let you generate speech without supplying a recording. Voice cloning accepts WAV or MP3 reference audio to produce a similar voice, while fixed generation seeds help keep results consistent across requests. For longer work, the server splits text at sentence boundaries and joins the generated audio into a single file, including book-length narration. It also supports streamed audio output.

The server runs on Windows, Linux and macOS, with Docker deployment available. It can use NVIDIA GPUs through CUDA, AMD GPUs through ROCm, Apple Silicon through MPS, or a CPU. Speech generation runs on the machine hosting the server; the optional Google Colab demo runs in Google's cloud. Windows also has a portable mode that keeps Python and dependencies in a movable, self-contained folder.

Similar to Chatterbox TTS Server