Favicon of Unmute

Unmute

An open-source voice interface for text LLMs, with local hosting, Ollama and vLLM support, and speech models that require a CUDA GPU.

Screenshot of Unmute website

Unmute adds spoken conversation to text LLMs using Kyutai's speech recognition and speech synthesis models. It's for developers who want a self-hosted voice interface while keeping their choice of language model. The project uses the MIT license, and a hosted browser demo is available at Unmute.sh.

Both speech models prioritize low latency. Unmute transcribes your speech, passes the text to an LLM, and starts speaking the answer while the model is still generating it. This lets a text model participate in voice conversations without needing built-in audio support.

The backend works with OpenAI-compatible LLM servers, including vLLM and Ollama on your own hardware, or external services such as OpenRouter. The default Docker setup runs the speech services and vLLM locally, so conversation processing can stay on your machine. Choosing an external LLM sends the text exchange to that service; the hosted demo processes conversations on remote servers.

Self-hosting requires an x86_64 machine and a CUDA-capable GPU with at least 16 GB of VRAM. Unmute supports Docker and deployment without Docker. Spreading speech recognition, speech synthesis and the LLM across separate GPUs can reduce response latency compared with running them together on one GPU.

The browser interface includes character choices with distinct voices and prompts, plus subtitles for both sides of the conversation. Developers can customize the voices and character instructions.

Similar to Unmute