16KUpdated 21 hours agoBSD-2-Clause
#LLM tracing#Multi-agent workflows#Multimodal input
Pipecat is a Python framework for developers building conversational AI agents that handle speech, video, text, and images. You can run it on your own machine or servers, wherever Python runs. It's open source under the BSD 2-Clause license.
20.3KUpdated 1 hour agoMIT
#Human approval#LLM tracing#MCP
Pydantic AI is a Python SDK for developers building AI agents into their own applications. Its main draw is Pydantic validation across agent tools and results, so an agent can return structured data that application code can check and use. The SDK is MIT licensed.
3.8KUpdated 2 years agoMIT
Docker#Streaming inference#Tool calling
Vocode is an open source Python library for developers building voice AI agents, with a self-hosted telephony server and support for live conversations through a computer's microphone and speakers. It connects speech recognition, an LLM, and speech synthesis in one library. The code uses the MIT license.
382Updated 3 years agoMIT
#Home Assistant integration#Multilingual#Voice activity detection
Rhasspy 3 is an early developer-preview local voice assistant toolkit for developers building their own assistants or adding voice control to Home Assistant. The project is archived and no longer maintained. It keeps data on your computer unless you choose to send it elsewhere, and its speech components support languages beyond English.
778Updated 2 days agoMIT
Docker#Human approval#Multilingual#Tool calling
Bolna is a voice AI agent platform for businesses handling customer support, lead qualification, reminders, and recruitment calls. Its focus is Indian languages and accents, including Hindi, Hinglish, Tamil, and Telugu, alongside English. Developers can self-host its MIT-licensed Python orchestration core, while business teams can use a hosted dashboard to build agents without code.
11.2KUpdated 1 month ago
macOS · Windows · Linux · iOS · Android · Web#Multilingual#Streaming inference
Moonshine is an on-device AI toolkit for developers building voice agents and applications that listen and speak. It combines speech to text, intent recognition and text to speech in one library. Voice processing stays on the device, and you don't need an account or API keys.
295Updated 2 days agoApache-2.0
Linux · Docker#OpenAI-compatible API#Wake word detection#Works offline
OpenVoiceOS is a free, open-source voice AI platform for people building their own smart speakers or adding voice control to devices. Its core builds on a fork of MycroftAI/mycroft-core, and most classic Mycroft skills also work with it. The project uses the Apache 2.0 license, which permits personal and commercial use.
397Updated 1 day agoMIT
#Home Assistant integration#Voice activity detection
Wyoming Protocol connects Home Assistant with separate voice services on a trusted network. It's for developers and people assembling a self-hosted voice assistant who want to choose their own speech recognition, speech synthesis and wake word components. The Python project is open source under the MIT license.
14.4KUpdated 19 hours agoApache-2.0
#MCP#Multi-agent workflows#Multimodal input
LiveKit Agents is a framework for developers building voice assistants, phone agents, and apps that combine speech with video or text. Agents join LiveKit rooms as participants, so they can interact with people through web and mobile apps or telephone calls. The Apache 2.0 project lets you run the entire stack on your own servers, including the LiveKit media server.
21.3KUpdated 9 months agoApache-2.0
Web#Guardrails#LLM tracing#MCP
Rasa is an AI agent platform for product teams building customer-facing text and voice assistants. Teams can deploy agents on their own infrastructure and choose their models and data arrangements. Its CALM engine combines language model understanding with business flows whose code enforces rules, so an assistant can handle conversational wording while following defined processes.
11.2KUpdated 5 months agoApache-2.0
macOS · iOS · Web#Hugging Face integration#MLX#Quantization
Moshi is a voice AI model and dialogue framework that can listen while it speaks. It processes speech directly, retaining information such as emotion and non-verbal cues that a text transcription can miss. It's aimed at researchers and developers building spoken AI applications, with local inference and self-hosted server options.
11.1KUpdated 18 hours ago
Docker · Web#Multimodal input#Voice activity detection
TEN Framework is a self-hosted framework for developers building voice AI agents and multimodal conversational apps. It focuses on low-latency, real-time conversations and supports both RTC and WebSocket connections. You can run its agent examples locally with Docker or deploy them on your own server.
29.8KUpdated 23 hours agoMIT
macOS · Windows · Linux#Code execution#Guardrails#Human approval
OpenAI Agents SDK is an open-source Python framework for developers building AI apps that need to use tools, delegate tasks, or work across multiple steps. Its runtime manages agent turns and conversation state while letting developers express workflows in ordinary Python. It uses the MIT license.