Favicon of Dia

Dia

Open-source text-to-speech model for local English dialogue generation, with voice cloning, NVIDIA GPU inference and an Apache 2.0 license.

Dia is the original text-to-speech model from Nari Labs that generates a two-speaker conversation from a written script in one pass. It's for researchers and developers who want to generate English dialogue on their own hardware, with control over speaker voices and delivery. The code and model weights are available under Apache 2.0. Dia2 is a separately linked successor.

Audio prompts let you clone a voice and influence the emotion and tone of generated speech. Dia can also produce laughter, coughs and other nonverbal sounds within a conversation, though these can yield unexpected results. It isn't tied to a fixed voice: generated voices can vary between runs, while an audio prompt or a fixed seed can help keep speakers consistent.

The Python implementation includes a Gradio interface, Docker files and a command-line interface, and the model also works through Hugging Face Transformers. Hugging Face hosts the downloadable weights and a separate browser demo. Local inference runs on your hardware; the hosted demo runs on Hugging Face and accepts uploaded audio for voice cloning.

Hardware is a practical consideration. Dia has been tested on NVIDIA GPUs using PyTorch and CUDA. Benchmarks on an RTX 4090 used roughly 4.4 GB of VRAM at reduced precision or 7.9 GB at full precision. Generation quality also depends on script length: very short passages can sound unnatural, while long passages can produce speech that sounds rushed.

Similar to Dia