Verified project record
k2-fsa/OmniVoice
OmniVoice is a text-to-speech model from k2-fsa covering more than 600 languages, with zero-shot voice cloning, attribute-driven voice design, automatic voice selection, and pronunciation controls. It runs locally via a Python package with CLI, library, and web GUI interfaces, and includes batch multi-GPU inference and training examples.
Project overview
The project reports support for more than 600 languages in a zero-shot TTS model and a real-time factor as low as 0.025 (described as 40 times faster than real time), with one unified generation API covering voice cloning, voice design, and automatic voice selection.
- Project type
- Model Development · Model Runtime · Audio & Speech
- Use cases
- Audio & Speech
- Deployment
- Refer to project documentation
- License
- Apache-2.0
Best for
- Developers and AI engineers who need text-to-speech across a wide range of languages (more than 600) with voice cloning, voice design, and automatic voice selection in one API.
- Users comfortable with a Python-based setup who can install PyTorch for their hardware backend and work via CLI, library, or the bundled web GUI.
Key capabilities
- Generates speech from input text in more than 600 supported languages.
- Generates speech in a cloned voice using a short reference recording and its transcript; the recommended reference length is 3–10 seconds.
- Creates a voice from attributes such as gender, age, pitch, whisper style, English accent, and Chinese dialect without requiring reference audio. Voice design is trained only on Chinese and English data.
- Generates speech without a reference recording or voice-design instruction by letting the model choose a voice automatically.
- Accepts inline non-verbal tags and Chinese pinyin or English phoneme overrides to control expressive sounds and pronunciation.
- Distributes JSONL-defined batch text-to-speech inference across multiple GPUs for large-scale generation tasks.
- Includes an examples pipeline covering data preparation, training, evaluation, and fine-tuning.
Limitations and risks
- Voice design is trained only on Chinese and English data.
- Voice design may produce unstable results for some low-resource languages or edge cases.
- Cross-lingual voice cloning carries an accent from the reference audio language.
- Reference audio longer than the recommended 3–10 seconds slows inference and may degrade cloning quality.
- Flash attention is unavailable on Intel XPU, so the model falls back to SDPA.
- Training with packed sequences has only partial Intel XPU support.
- Voice cloning can be misused for unauthorized impersonation, fraud, scams, or other illegal or unethical activities; adopters should implement consent and usage safeguards.
Getting started
- Setup difficulty is medium: install PyTorch for the selected hardware backend, then run pip install omnivoice. Pretrained model weights may need to be downloaded before local use; the documented inference examples load the k2-fsa/OmniVoice pretrained model, and downloading uses Hugging Face connectivity or the documented hf-mirror endpoint.
- Install PyTorch for the selected hardware backend, run pip install omnivoice, then run omnivoice-demo --ip 0.0.0.0 --port 8001 to launch the demo.
Evidence and sources
- GitHub project description: High-Quality Voice Cloning TTS for 600+ Languages
- README: OmniVoice is a state-of-the-art massively multilingual zero-shot text-to-speech (TTS) model supporting over 600 languages. Built on a novel diffusion language model-style architec…
- README: - **600+ Languages Supported**: The broadest language coverage among zero-shot TTS models ([full list](docs/languages.md)). - **Voice Cloning**: State-of-the-art voice cloning qua…
- README: **Step 1**: Install PyTorch
- README: pip install omnivoice
AI Search
Find projects, verify facts, compare options, or turn a complex need into an actionable plan
Try a searchA click only fills the search box; you stay in control
Project Details
0