Verified project record
moeru-ai/airi
A self-hosted AI VTuber companion application supporting realtime voice chat, VRM/Live2D avatar animation, and interaction with chat platforms. It is built with web technologies, offers optional native GPU acceleration on desktop, and is currently in an early stage of development.
Project overview
The project provides a self-hosted companion application built natively around web technologies like WebGPU, WebAudio, and WebAssembly, combined with the ability to animate VRM and Live2D avatars and interact via voice and text.
- Project type
- AI Agent · Audio & Speech
- Use cases
- Automation
- Deployment
- Refer to project documentation
- License
- MIT
Best for
- General users, developers, and creators seeking a self-hosted AI VTuber companion for realtime voice chat and avatar animation.
- Users who prefer web technology stacks and want to run the application on modern browsers or desktop environments with optional GPU acceleration.
Key capabilities
- Supports multiple voice synthesis providers including ElevenLabs, Microsoft/Azure Speech, OpenAI-compatible TTS, Alibaba Cloud Model Studio, and local Kokoro TTS.
- Supports loading, controlling, and animating VRM and Live2D models with auto blink, look at, and idle eye movement.
- Chats directly in Telegram and Discord.
- Pure in-browser database support using DuckDB WASM or pglite.
- Desktop version capable of using native NVIDIA CUDA and Apple Metal, while browser version uses WebGPU.
- Ability to run inference locally inside the browser using WebGPU.
Limitations and risks
- The project is still in the early stage of development and is seeking talented developers to help make the application a reality.
- The project carries an adoption risk because it is still in the early stage of development.
- External services like OpenAI API and ElevenLabs API are optional. Paid services may be used optionally for LLM and TTS providers.
Getting started
- Setup is documented as easy because desktop binaries and a hosted web version are available without code compilation. The first success path is to download the installer or open the hosted demo.
Alternatives and comparisons
- A framework to create conversational multi-modal voice agents running as realtime programmable participants on servers.
- An open-source framework for real-time multimodal conversational AI with low-latency extensions and ready-to-use agent examples.
- A documented, private on-device open source AI toolkit for real-time voice tasks without requiring external API keys.
Project comparisons
Evidence and sources
- GitHub project description: 💖🧸 Self hosted, you-owned Grok Companion, a container of souls of waifu, cyber livings to bring them into our worlds, wishing to achieve Neuro-sama's altitude. Capable of realtime…
- README: Re-creating Neuro-sama, a soul container of AI waifu / virtual characters to bring them into our world.
- README: Unlike the other AI driven VTuber open source projects, アイリ was built with support of many Web technologies such as [WebGPU](https://www.w3.org/TR/webgpu/), [WebAudio](https://dev…
- README: We are still in the early stage of development where we are seeking out talented developers to join us and help us to make アイリ a reality.
- README: - [x] Multi-provider voice synthesis, including [ElevenLabs](https://elevenlabs.io/), Microsoft/Azure Speech, OpenAI-compatible TTS, Alibaba Cloud Model Studio, and local Kokoro T…
AI Search
Find projects, verify facts, compare options, or turn a complex need into an actionable plan
Try a searchA click only fills the search box; you stay in control
Project Details
0