Pyannote Audio
Open-source speaker diarization and voice activity detection toolkit.
About
pyannote.audio is a Python toolkit for speaker diarization, voice activity detection, overlapped-speech detection, and speaker embedding extraction, built on PyTorch. It ships pretrained pipelines that can be fine-tuned on user data and exposes a Pipeline API that loads models directly from the Hugging Face Hub. The community-1 pipeline is the current open-source baseline; a paid precision-2 pipeline is offered alongside. MIT licensed.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- MIT
- Minimum VRAM
- 4 GB
- Added
- Apr 3, 2026
Related Tools
Convolution-augmented transformer for speech recognition in ESPnet toolkit.
End-to-end speech processing toolkit covering ASR, TTS, and speech translation.
CLI tool that transcribes audio 10x faster using pipeline optimizations.
Established speech recognition toolkit used in research and production systems.
Self-supervised speech representation model by Meta for ASR.
Multilingual ASR model by NVIDIA supporting 4 languages with translation.