Favicon of TEN Framework

TEN Framework

A self-hosted voice AI framework with RTC and WebSocket support, Docker deployment, and examples that use OpenAI, Deepgram and Agora.

Screenshot of TEN Framework website

TEN Framework is a self-hosted framework for developers building voice AI agents and multimodal conversational apps. It focuses on low-latency, real-time conversations and supports both RTC and WebSocket connections. You can run its agent examples locally with Docker or deploy them on your own server.

The voice assistant example provides a starting point for a custom agent. Its extensions cover speech recognition, language models and speech synthesis, with memory, voice activity detection (VAD) and turn detection available to extend the conversation. These capabilities matter for spoken interaction, where an agent needs to detect speech and recognize when someone has finished their turn.

TEN includes TMAN Designer, a visual interface for customizing agent components, alongside a browser interface for the examples. Developers can adapt the speech and model components rather than build the whole conversational app from scratch. The framework's scope also includes video and other multimodal interactions.

Local hosting doesn't mean offline operation. The default example requires credentials for Agora, OpenAI and Deepgram ASR, so it depends on external services even when the application runs on your hardware. Its stated minimum is a two-core CPU and 4 GB of RAM.

For server deployment, TEN supports Docker images. You can host the backend on a container-friendly server and place the frontend separately on services such as Vercel or Netlify.

Similar to TEN Framework