Player not loading? Watch on YouTube
This tutorial walks through a local AI voice agent built with LiveKit and deployed through Docker. The supplied description identifies Ollama's gemma-4-e2b as the language model, Nemotron or Whisper for speech recognition, and Kokoro for speech output. The speaker presents the setup as private and free, with no tokens or subscriptions required.
Setup starts with downloading the repository, installing Docker and Python, then building the services with Docker Compose. The speaker recommends Python 3.12 or later and 16 GB RAM, and specifies at least 20 GB of free storage. An Nvidia GPU is recommended for faster responses, though the speaker says CPU execution also works. Despite the video's ten-minute framing, the initial build can take up to 45 minutes. The browser interface shows component loading status before a conversation can begin.
For Mac, the tutorial switches the speech recognition provider to Whisper in the environment file and redeploys the system. The speaker recommends this because Whisper is better optimized for Apple silicon, while Nemotron targets Nvidia CUDA.
Customization covers editing agent.py to give the assistant an Alfred persona and adding a function tool that ends a conversation. A separate environment setting selects a Kokoro voice. The final demonstration tests the changed voice and conversation-ending tool; the speaker also reports lower latency than in the earlier setup.