Pipecat and Gemini: build a voice agent with tools

Learn to scaffold a Pipecat voice bot, register tools, save Markdown reports, and connect agents. Flight and lodging searches use mock data.

Player not loading? Watch on YouTube

Chad Bailey walks through building a travel-planning voice bot with Pipecat and a Gemini 3-based real-time multimodal model. The tutorial starts with CLI scaffolding and a web transport, then configures the model and its system prompt. Bailey sets thinking to minimal to reduce response latency and reports that his longer test conversations stayed on topic.

The flight and lodging tools illustrate function handlers, schemas, and registration with the model. Both return mock data, so the spoken travel prices and listings are demonstration output rather than verified booking results. Google Search grounding is enabled separately. A report tool asks the model for Markdown describing the trip and saves that text to disk; Bailey also explains how to load it into a later session's context.

The final section introduces Pipecat Agents. A concierge and a language tutor have different system prompts, with each agent maintaining its own pipeline, LLM service, and tools. An agent runner connects them through a shared message bus.

Bailey describes Pipecat as open source and suitable for a self-hosted bot on the user's infrastructure. This example uses Gemini Live API. The supplied metadata names the public model as gemini-3.1-flash-live and warns that the early-access model identifier shown in the video has been turned down.