
Higgs Audio is a family of text-to-speech models from Boson AI for developers building narration and conversational audio. Higgs TTS 2 can adapt pacing and intonation to the text and generate dialogue with distinct speakers across multiple languages.
Voice cloning uses reference audio without training a separate model for each voice. A conversation can include several cloned voices, or the model can choose voices based on written speaker descriptions. Its audio generation also covers melodic humming in a cloned voice and speech with background music.
Higgs TTS 2 works natively with Hugging Face Transformers and supports batch generation. It builds on Llama-3.2-3B with an audio-specific DualFFN adapter. Its shared audio tokenizer handles speech, music and sound events within the same system.
Higgs Audio V2 is now named Higgs TTS 2. The original checkpoint runs locally through Transformers or the documented V2 Python code, with CUDA and CPU paths in the example. The current repository also introduces a separate V3 release; that successor has different serving code and model terms.
The developer code uses Apache 2.0. Higgs TTS 2 weights use the Boson Higgs Audio 2 Community License, which incorporates Llama terms and imposes additional commercial conditions. Check those model terms before deployment; V3's research and noncommercial license does not describe this V2 checkpoint.
Claim this page and we'll verify you by hand. Higgs Audio gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Higgs Audio?Promote it
Something wrong or outdated on this page?
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
19.4KUpdated 10 months agoApache-2.0
Docker · Web#Hugging Face integration#Multimodal input#Voice cloning
32.9KUpdated 2 weeks ago
#Batch processing#Multilingual#Multimodal input
Fish Speech, currently featuring Fish Audio S2 Pro, is a self-hosted text-to-speech system for creators producing narration and developers building voice applications. It combines voice cloning with control over emotion and delivery within a script. Code and model weights use the custom FISH AUDIO RESEARCH LICENSE.
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.