Vosk

Offline speech recognition toolkit supporting 20+ languages with small models.

Open SourceSelf HostedOffline Capable
0.0 (0)

About

Vosk makes speech recognition work where the cloud cannot reach. The toolkit from Alpha Cephei performs fully offline transcription in more than 20 languages and dialects, from English and Spanish to Chinese, Russian, and Japanese, using compact models around 50 MB that run on hardware as small as a Raspberry Pi or an Android phone while also scaling up to server clusters. Its streaming API returns partial results with effectively zero latency, and it supports large-vocabulary continuous transcription, runtime-reconfigurable vocabularies for command grammars, and speaker identification. Bindings cover Python, Java, Node.js, C#, C++, Rust, Go, and more, which is why it shows up in everything from smart home devices and IVR systems to subtitle generators and lecture transcription tools. Released under the Apache 2.0 license with freely downloadable models, Vosk is a frequent choice for developers who need private, on-device voice input without per-minute API fees, accepting somewhat lower accuracy than large cloud models in exchange.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Easy (2/5)
License
Apache-2.0
Added
Apr 3, 2026

Related Tools

Convolution-augmented transformer for speech recognition in ESPnet toolkit.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

End-to-end speech processing toolkit covering ASR, TTS, and speech translation.

Open SourceSelf HostedOfflineGPU 8GB+
Expert
0.0 (0)

CLI tool that transcribes audio 10x faster using pipeline optimizations.

Open SourceSelf HostedOfflineGPU 6GB+
Easy
0.0 (0)

Established speech recognition toolkit used in research and production systems.

Open SourceSelf HostedOffline
Expert
0.0 (0)

Self-supervised speech representation model by Meta for ASR.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Multilingual ASR model by NVIDIA supporting 4 languages with translation.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)
Browse all Speech-to-Text / Speech Recognition tools