Favicon of Zonos

Zonos

Open-source text-to-speech model with voice cloning. Runs locally on Linux and macOS under Apache 2.0, with a hosted audio playground also available.

Screenshot of Zonos website

Zonos is an open-source text-to-speech model for people who want to generate speech and clone voices on their own hardware. It can match a speaker from a short reference recording, with controls for delivery and emotion. The code uses the Apache 2.0 license.

It supports English, Japanese, Chinese, French, and German. Alongside voice cloning, it accepts an audio prefix to guide how the speech sounds. That can help reproduce delivery such as whispering, which can be harder to capture from a voice sample alone. You can adjust speaking rate, pitch variation and audio quality, or guide emotions such as happiness, sadness, fear and anger.

Zonos includes a Gradio browser interface for generating speech and supports Docker deployment. Local deployment runs the model on your hardware; Zyphra also offers a separate hosted audio playground. The model comes in transformer and hybrid variants and generates audio at 44 kHz.

Linux and macOS are supported. GPU use requires at least 6 GB of VRAM, and the hybrid variant additionally needs an NVIDIA GPU from the 3000 series or newer. CPU execution is possible with enough free RAM, but it's much slower and may be unsuitable for interactive use. On an RTX 4090, the model can generate roughly two seconds of speech per second of computation.

Similar to Zonos