Dograh, LM Studio and Speaches AI voice agent setup

Set up a Hindi voice agent on Windows with Docker, LM Studio and Speaches AI, including a port 8001 workaround and website widget configuration.

Player not loading? Watch on YouTube

This tutorial connects Dograh with LM Studio and Speaches AI to build a self-hosted voice agent. Dograh handles orchestration, LM Studio runs the local LLM, and Speaches AI supplies speech recognition and speech synthesis. The example uses Faster Whisper large-v3 and Kokoro, then creates a dental appointment agent with Hindi as its main language.

The presenter works on Windows with Docker, using a server with two Intel Xeon processors, 128GB RAM and two NVIDIA GPUs with 16GB each. He describes 8GB RAM as a minimum for Dograh itself; that figure does not establish the requirements for the full model stack. His zero-per-minute claim excludes server setup and telephone costs.

The walkthrough covers installation problems, including a MinIO image reference and a port conflict. Dograh's interface uses port 3010 and its backend uses 8000, so the presenter moves Speaches AI to 8001. He reports that its interface still expects port 8000 and continues through CLI mode. The configuration also uses host.docker.internal to reach the host from Docker.

After configuring model URLs, voice IDs and transcription language, the presenter tests the agent. He then explains how to copy the website widget code and notes that the Dograh instance needs a publicly accessible domain. Telephone calling requires a separate provider configuration; he uses Voicelink.