Whisper
OpenAI's powerful speech recognition model
About
Whisper by OpenAI is a general-purpose speech recognition model trained on a large and diverse audio dataset. It is a Transformer sequence-to-sequence model that handles multilingual transcription, speech translation, and spoken language identification as a single multitask system, and it holds up well across accents, background noise, and technical speech. Pretrained model checkpoints and inference code are released under the MIT license.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Audio & Speech
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- MIT
- Added
- Jan 29, 2026
Related Tools
Free text-to-speech generator with multiple voices, accents, and languages. No signup required.
CTranslate2-based Whisper with 4x faster transcription
Universal neural vocoder from NVIDIA that converts mel spectrograms into waveforms up to 44 kHz.
End-to-end Chinese and English spoken dialogue model from Zhipu AI with streaming speech output.
Audio foundation model unifying speech recognition, understanding, and conversation in one 7B model.
Deep learning toolkit for text-to-speech synthesis