WhisperX
Whisper with word-level timestamps and speaker diarization
About
WhisperX is a Whisper-based automatic speech recognition pipeline that adds word-level timestamps via forced phoneme alignment and uses voice-activity detection to batch audio for fast inference (around 70x realtime with the large-v2 model). It supports speaker diarization through external models and runs on GPU with CUDA 12.8 or on CPU. Useful for subtitles, transcript editing, and meeting-style audio where word offsets matter.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Audio & Speech
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- BSD-4-Clause
- Added
- Jan 29, 2026
Related Tools
Free text-to-speech generator with multiple voices, accents, and languages. No signup required.
Deep learning toolkit for text-to-speech synthesis
CTranslate2-based Whisper with 4x faster transcription
Universal neural vocoder from NVIDIA that converts mel spectrograms into waveforms up to 44 kHz.
End-to-end Chinese and English spoken dialogue model from Zhipu AI with streaming speech output.
Transformer-based text-to-audio model from Suno
Mentioned in
Beyond Whisper: Parakeet, SenseVoice and ASR in 2026
Whisper is no longer the default: how Parakeet, SenseVoice, Kimi-Audio, Ultravox and Moshi compare on...
Max P
whisper.cpp vs faster-whisper: Speed and Accuracy Compared
Two leading open source paths to running OpenAI Whisper. One is a CPU-friendly C/C++ port, the other rides...
Billy C