
LiveKit Agents is a framework for developers building voice assistants, phone agents, and apps that combine speech with video or text. Agents join LiveKit rooms as participants, so they can interact with people through web and mobile apps or telephone calls. The Apache 2.0 project lets you run the entire stack on your own servers, including the LiveKit media server.
Its main focus is the conversation itself. It handles streaming speech recognition, language model responses, and speech synthesis, with turn detection and interruption handling built into the framework. A transformer model helps detect when someone has finished speaking. You can combine different speech and language model providers or use realtime model APIs.
Python and Node.js SDKs let developers define agent behavior in code. Agents can call tools, use tools from MCP servers, and hand a conversation to another agent. They can also exchange data with the frontend, which supports interactions beyond spoken replies. LiveKit's SIP integration covers incoming and outgoing phone calls.
Self-hosting controls where the agent and media infrastructure run; model processing depends on the providers you connect. LiveKit Cloud provides hosted agent deployment, transcripts and traces, and model access through LiveKit Inference. Its browser-based Agent Builder offers a separate way to prototype and deploy agents without writing code.
For larger deployments, the framework includes job scheduling, automatic load balancing, and Kubernetes compatibility. Built-in testing supports checks of agent behavior with automated judges.
Claim this page with an email at docs.livekit.io. LiveKit Agents gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LiveKit Agents?Promote it
Something wrong or outdated on this page?
29.8KUpdated 23 hours agoMIT
macOS · Windows · Linux#Code execution#Guardrails#Human approval
OpenAI Agents SDK is an open-source Python framework for developers building AI apps that need to use tools, delegate tasks, or work across multiple steps. Its runtime manages agent turns and conversation state while letting developers express workflows in ordinary Python. It uses the MIT license.
16KUpdated 21 hours agoBSD-2-Clause
#LLM tracing#Multi-agent workflows#Multimodal input
20.3KUpdated 1 hour agoMIT
#Human approval#LLM tracing#MCP
21.3KUpdated 9 months agoApache-2.0
Web#Guardrails#LLM tracing#MCP
11.1KUpdated 18 hours ago
Docker · Web#Multimodal input#Voice activity detection
TEN Framework is a self-hosted framework for developers building voice AI agents and multimodal conversational apps. It focuses on low-latency, real-time conversations and supports both RTC and WebSocket connections. You can run its agent examples locally with Docker or deploy them on your own server.
5KUpdated 23 hours agoApache-2.0
#Code execution#Human approval#MCP
AG2 is an open-source Python framework for developers and researchers building systems where AI agents share work. The code uses the Apache 2.0 license.
Pipecat is a Python framework for developers building conversational AI agents that handle speech, video, text, and images. You can run it on your own machine or servers, wherever Python runs. It's open source under the BSD 2-Clause license.
Pydantic AI is a Python SDK for developers building AI agents into their own applications. Its main draw is Pydantic validation across agent tools and results, so an agent can return structured data that application code can check and use. The SDK is MIT licensed.
Rasa is an AI agent platform for product teams building customer-facing text and voice assistants. Teams can deploy agents on their own infrastructure and choose their models and data arrangements. Its CALM engine combines language model understanding with business flows whose code enforces rules, so an assistant can handle conversational wording while following defined processes.