Verified project record
huggingface/speech-to-speech
A modular voice-agent pipeline that combines VAD, STT, LLM, and TTS components in separate threads with swappable backends selected via CLI flags. It exposes an OpenAI Realtime-compatible WebSocket API and supports local microphone interaction, raw audio streaming, and multi-language configurations.