Favicon of Higgs Audio

Higgs Audio

Higgs Audio V2, now Higgs TTS 2, is a downloadable speech model for expressive narration, multilingual dialogue and voice cloning.

Screenshot of Higgs Audio website

Higgs Audio is a family of text-to-speech models from Boson AI for developers building narration and conversational audio. Higgs TTS 2 can adapt pacing and intonation to the text and generate dialogue with distinct speakers across multiple languages.

Voice cloning uses reference audio without training a separate model for each voice. A conversation can include several cloned voices, or the model can choose voices based on written speaker descriptions. Its audio generation also covers melodic humming in a cloned voice and speech with background music.

Higgs TTS 2 works natively with Hugging Face Transformers and supports batch generation. It builds on Llama-3.2-3B with an audio-specific DualFFN adapter. Its shared audio tokenizer handles speech, music and sound events within the same system.

Higgs Audio V2 is now named Higgs TTS 2. The original checkpoint runs locally through Transformers or the documented V2 Python code, with CUDA and CPU paths in the example. The current repository also introduces a separate V3 release; that successor has different serving code and model terms.

The developer code uses Apache 2.0. Higgs TTS 2 weights use the Boson Higgs Audio 2 Community License, which incorporates Llama terms and imposes additional commercial conditions. Check those model terms before deployment; V3's research and noncommercial license does not describe this V2 checkpoint.

Similar to Higgs Audio