Verified project record
m-bain/whisperX
WhisperX adds batched inference, word-level timestamps via wav2vec2 alignment, and pyannote-audio speaker diarization to Whisper large-v2, reaching up to 70x realtime transcription. It is designed for developers and researchers who need precise word timings and speaker labels but should be evaluated against its handling of overlapping speech and its reliance on language-specific alignment models.