
TEN Framework is a self-hosted framework for developers building voice AI agents and multimodal conversational apps. It focuses on low-latency, real-time conversations and supports both RTC and WebSocket connections. You can run its agent examples locally with Docker or deploy them on your own server.
The voice assistant example provides a starting point for a custom agent. Its extensions cover speech recognition, language models and speech synthesis, with memory, voice activity detection (VAD) and turn detection available to extend the conversation. These capabilities matter for spoken interaction, where an agent needs to detect speech and recognize when someone has finished their turn.
TEN includes TMAN Designer, a visual interface for customizing agent components, alongside a browser interface for the examples. Developers can adapt the speech and model components rather than build the whole conversational app from scratch. The framework's scope also includes video and other multimodal interactions.
Local hosting doesn't mean offline operation. The default example requires credentials for Agora, OpenAI and Deepgram ASR, so it depends on external services even when the application runs on your hardware. Its stated minimum is a two-core CPU and 4 GB of RAM.
For server deployment, TEN supports Docker images. You can host the backend on a container-friendly server and place the frontend separately on services such as Vercel or Netlify.
Claim this page with an email at theten.ai. TEN Framework gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find TEN Framework?Promote it
Something wrong or outdated on this page?
11.2KUpdated 5 months agoApache-2.0
macOS · iOS · Web#Hugging Face integration#MLX#Quantization
Moshi is a voice AI model and dialogue framework that can listen while it speaks. It processes speech directly, retaining information such as emotion and non-verbal cues that a text transcription can miss. It's aimed at researchers and developers building spoken AI applications, with local inference and self-hosted server options.
295Updated 2 days agoApache-2.0
Linux · Docker#OpenAI-compatible API#Wake word detection#Works offline
21.3KUpdated 9 months agoApache-2.0
Web#Guardrails#LLM tracing#MCP
14.4KUpdated 19 hours agoApache-2.0
#MCP#Multi-agent workflows#Multimodal input
16KUpdated 21 hours agoBSD-2-Clause
#LLM tracing#Multi-agent workflows#Multimodal input
382Updated 3 years agoMIT
#Home Assistant integration#Multilingual#Voice activity detection
OpenVoiceOS is a free, open-source voice AI platform for people building their own smart speakers or adding voice control to devices. Its core builds on a fork of MycroftAI/mycroft-core, and most classic Mycroft skills also work with it. The project uses the Apache 2.0 license, which permits personal and commercial use.
Rasa is an AI agent platform for product teams building customer-facing text and voice assistants. Teams can deploy agents on their own infrastructure and choose their models and data arrangements. Its CALM engine combines language model understanding with business flows whose code enforces rules, so an assistant can handle conversational wording while following defined processes.
LiveKit Agents is a framework for developers building voice assistants, phone agents, and apps that combine speech with video or text. Agents join LiveKit rooms as participants, so they can interact with people through web and mobile apps or telephone calls. The Apache 2.0 project lets you run the entire stack on your own servers, including the LiveKit media server.
Pipecat is a Python framework for developers building conversational AI agents that handle speech, video, text, and images. You can run it on your own machine or servers, wherever Python runs. It's open source under the BSD 2-Clause license.
Rhasspy 3 is an early developer-preview local voice assistant toolkit for developers building their own assistants or adding voice control to Home Assistant. The project is archived and no longer maintained. It keeps data on your computer unless you choose to send it elsewhere, and its speech components support languages beyond English.