duixcom/duix-avatar provides offline digital human cloning and video synthesis capabilities for creators and developers with supported NVIDIA GPU hardware. It requires substantial local disk and memory resources and does not support real-time interaction in the open-source version.
Project overview
The project is designed to operate fully offline without an internet connection, offering local digital human cloning and video synthesis for privacy-conscious workflows.
Project type
Video
Use cases
Video Creation
Deployment
Refer to project documentation
License
License pending
Best for
Creators and developers who need to produce digital human videos locally and have an NVIDIA graphics card with at least 32GB of RAM and 100GB+ of disk space.
Key capabilities
Uses AI algorithms to capture facial features and clone voices to build realistic virtual models.
Drives virtual avatars using natural language processing to convert text to speech or using direct voice input; the open-source version does not support real-time interaction.
Synchronizes digital human video images with sound for natural lip-syncing.
Scripts support eight languages: English, Japanese, Korean, Chinese, French, German, Arabic, and Spanish.
Exposes local APIs for model training, audio synthesis, and video synthesis.
Limitations and risks
Requires an NVIDIA GPU, 32GB of RAM, and 100GB+ of free disk space.
Does not support real-time interaction in the open-source version.
Officially supports only Windows 10 and Ubuntu 22.04.
Deployment requires downloading large Docker images, approximately 70GB in size.
Getting started
To get started, users must meet the hardware requirements, install Docker, download the Docker images, execute docker-compose up, and install the client application.
Alternatives and comparisons
Provides an end-to-end automated video dubbing pipeline combining transcription, translation, text-to-speech, voice conversion, and lip synchronization.
Provides a complete workflow for video translation and audio transcription that supports local offline deployment and a wide variety of mainstream online APIs.
Provides an all-in-one automated pipeline to generate single-line subtitles and high-quality dubbing for videos, cutting, translating, aligning, and voicing content.