Player not loading? Watch on YouTube
This tutorial builds a browser-based voice AI agent with Pipecat, Sarvam AI and Nebius Token Factory. The application runs on a local server, but the demonstrated speech and language services require provider API keys. It is not an offline setup for running inference on your own hardware.
The speaker explains the speech-to-text, LLM and text-to-speech loop before walking through the code. WebRTC carries audio between the browser and the agent. A context aggregator stores user messages and assistant replies, then passes the conversation history to the language model. The pipeline streams synthesized speech back to the browser and records the assistant's response in that history.
Setup covers creating API keys and loading environment variables. The implementation defines transport settings, initializes the services and wraps the pipeline in a task. Connection handlers start a greeting and cancel the task when the client disconnects. The speaker uses WebRTC for the local demonstration and suggests Daily for production.
The tutorial also discusses model choice: the speaker warns that reasoning models can add response latency and recommends smaller models for tasks that do not need heavy reasoning. The final demonstration launches main.py with uv, connects to localhost port 7860 and checks conversation logs. Tool calling appears as a suggested extension, rather than an implemented feature.