ten-framework is an open-source framework for developing and deploying real-time multimodal conversational AI agents. It provides extensive examples and deployment options while addressing the need to integrate various models and services with low latency.
Project overview
ten-framework covers a notably wide array of use cases—from real-time voice assistants and transcription to lip-synced avatars and SIP phone call integration—consolidating multimodal conversational AI building blocks into a single framework.
Project type
AI Agent · Video · Audio & Speech
Use cases
Knowledge Q&A · Automation
Deployment
Refer to project documentation
License
License pending
Best for
Developers and AI engineers who need to integrate various models and services into real-time multimodal conversational AI agents.
Key capabilities
A low-latency, high-quality real-time assistant supporting RTC and WebSocket connections, extendable with memory, VAD, and turn detection.
A doodle board that turns spoken or typed prompts into simple hand-drawn sketches with a crayon palette and real-time drawing.
Real-time diarization that detects and labels speakers, demonstrated via an interactive game use case.
Works with multiple avatar vendors, featuring anime characters with MotionSync-powered lip sync and support for realistic avatars.
SIP extension that enables phone calls powered by TEN.
A transcription tool that transcribes audio to text.
Limitations and risks
On Apple Silicon, unchecking Rosetta for x86/amd64 emulation in Docker settings may result in slower build times on ARM.
Requires multiple paid API keys for core functionalities, including OpenAI, Deepgram, and ElevenLabs.
Getting started
Install Docker / Docker Compose and Node.js. Clone the repo and create a .env file. Set up Agora App ID, OpenAI, Deepgram, and ElevenLabs API keys in the .env file. Run docker compose up -d and docker exec to enter the container. Build the agent using task commands. Access the Agent Examples UI at localhost:3000. The build step takes 5-8 minutes. Minimum hardware: CPU >= 2 cores, RAM >= 4 GB. Coding is optional.
Evidence and sources
GitHub project description: Open-source framework for conversational voice AI agents
README: TEN is an open-source framework for real-time multimodal conversational AI.
README: Open-source framework for conversational AI Agents.
README: This low-latency, high-quality real-time assistant supports both RTC and [WebSocket][websocket-example] connections, and you can extend it with [Memory][memory-example], [VAD][voi…
README: A doodle board that turns spoken or typed prompts into simple hand-drawn sketches, complete with a crayon palette and real-time drawing.
AI Search
Find projects, verify facts, compare options, or turn a complex need into an actionable plan
Try a searchA click only fills the search box; you stay in control