Player not loading? Watch on YouTube
This tutorial builds a restaurant voice receptionist with Dograh AI, Kokoro 82M for speech output, and Llama 3.2 3B through Ollama for conversation text. The speaker describes Dograh as an open source voice pipeline framework and says Kokoro can run on a CPU without a GPU. That hardware claim concerns the speech model; the tutorial gives no measured requirements for the complete setup.
The walkthrough uses Google Antigravity as a Python IDE. It covers downloading the language model, checking the Ollama model list, installing eSpeak NG, creating a virtual environment, and installing dependencies from requirements.txt. The demonstrated installation uses PowerShell and an MSI installer, with macOS and Linux commands also referenced. The agent script accepts settings for the model, endpoint, voice, and prompt file. Bella is the default voice.
The microphone demo asks about opening hours, a table booking, and the menu. Although the bot says it has reserved a table, the speaker later explains that it has no connection to a reservation system. SQL integration and real phone calls through a service such as Twilio are proposed extensions. The speaker also suggests faster-whisper for speech recognition to make the pipeline fully offline and private. The example shows conversational replies, with one noted transcription miss, rather than a completed booking workflow.