Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.
Audio prompts let you clone a voice and influence the emotion and tone of generated speech. Dia can also produce laughter, coughs and other nonverbal sounds within a conversation, though these can yield unexpected results. It isn't tied to a fixed voice: generated voices can vary between runs, while an audio prompt or a fixed seed can help keep speakers consistent.
The Python implementation includes a Gradio interface, Docker files and a command-line interface, and the model also works through Hugging Face Transformers. Hugging Face hosts the downloadable weights and a separate browser demo. Local inference runs on your hardware; the hosted demo runs on Hugging Face and accepts uploaded audio for voice cloning.
Hardware is a practical consideration. Dia has been tested on NVIDIA GPUs using PyTorch and CUDA. Benchmarks on an RTX 4090 used roughly 4.4 GB of VRAM at reduced precision or 7.9 GB at full precision. Generation quality also depends on script length: very short passages can sound unnatural, while long passages can produce speech that sounds rushed.
Claim this page and we'll verify you by hand. Dia gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Dia?Promote it
Something wrong or outdated on this page?
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
7.2KUpdated 2 years agoApache-2.0
macOS · Linux · Docker · Web#Multilingual#Voice cloning
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
11KUpdated 1 year agoApache-2.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Voice cloning
14.9KUpdated 2 years agoApache-2.0
macOS · Windows · Docker#Streaming inference#Voice cloning
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
Zonos is an open-source text-to-speech model for people who want to generate speech and clone voices on their own hardware. It can match a speaker from a short reference recording, with controls for delivery and emotion. The code uses the Apache 2.0 license.
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
Spark-TTS is a local text-to-speech system that can copy a voice from reference audio or create a synthetic speaker with adjustable vocal traits. It's for developers and researchers building speech applications, including personalized narration, assistive technology, and language research. The Python and PyTorch code is open source under Apache 2.0.
Tortoise TTS is a local text-to-speech system for developers and creators who want speech with varied voices and natural pacing. It uses reference audio clips to guide a custom voice, with an emphasis on expressive rhythm and intonation.