Favicon of Chatterbox

Chatterbox

An open-source text-to-speech model family that runs on your own hardware, clones voices from short clips, and supports offline deployment.

Screenshot of Chatterbox website

Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.

The models cover different needs:

  • Original Chatterbox generates English speech with adjustable emotion and can clone a voice from a short reference clip without training a separate model.
  • Multilingual generates speech and clones voices across supported languages. Dedicated models are available for Chinese, Hindi, Spanish and Portuguese variants when language-specific output matters.
  • Turbo focuses on fast English speech and accepts text tags for sounds such as laughter and breaths. Nano shares that approach and can run on a CPU for on-device use.

Chatterbox also includes voice conversion scripts. Text choices such as capitalization can change emphasis, while emotion control lets a voice range from restrained to dramatic. These controls make it useful for voice agents and interactive media as well as recorded narration.

Every generated clip carries Resemble AI's PerTh watermark, which is designed to remain detectable after common audio processing. The models and weights are available under the MIT license, so teams can modify them and deploy them on their own servers.

Similar to Chatterbox