Favicon of Pipecat

Pipecat

Open-source Python framework for voice AI agents. Run it locally or on your servers, with WebRTC and telephony support under the BSD 2-Clause license.

Screenshot of Pipecat website

Pipecat is a Python framework for developers building conversational AI agents that handle speech, video, text, and images. You can run it on your own machine or servers, wherever Python runs. It's open source under the BSD 2-Clause license.

Its focus is live conversation. Pipecat combines speech recognition, language models, and speech synthesis in a streaming pipeline, so an agent can begin responding before every stage has finished. Interruption handling, turn detection, and conversation context management help it handle the pauses and overlaps of spoken exchanges. Speech, language, and vision services are interchangeable; supported integrations include Soniox, OpenAI, and Cartesia.

Agents can connect through WebRTC, WebSockets, or telephony using SIP and PSTN. Client SDKs and a UI kit support applications built around those connections. For more involved tasks, agents can hand work to specialists, run tasks in parallel, or coordinate across processes and machines. Pipecat Flows supports conversations that need a defined sequence, and evaluation tools, metrics, traces, and OpenTelemetry help teams test and inspect agent behavior.

The self-hosted framework runs on infrastructure you control. Pipecat Cloud is a separate managed offering that hosts, scales, and monitors agents, while Pipecat Enterprise supports deployment in your own AWS, Azure, or GCP environment. Your choice of speech and model services determines which providers participate in processing a conversation.

Similar to Pipecat