Canary (NVIDIA NeMo)

Multilingual ASR model by NVIDIA supporting 4 languages with translation.

Open SourceSelf HostedOffline CapableGPU Required (8GB+ VRAM)
0.0 (0)

About

Canary is NVIDIA's family of multilingual speech models built on the NeMo toolkit, combining a FastConformer encoder with a Transformer decoder in an encoder-decoder design. The original Canary-1B has one billion parameters across 24 encoder and 24 decoder layers and covers English, German, French, and Spanish, performing both transcription and speech-to-text translation between English and the other three languages, with or without punctuation and capitalization. Training drew on 85,000 hours of speech, and the model uses concatenated SentencePiece tokenizers, one per language. Later releases extended the family to more languages with lower word error rates. Models load through NeMo, an open source Apache 2.0 speech framework, with tasks and languages specified through prompts or manifest files; weights are published on Hugging Face under their own terms, and the original 1B checkpoint carries a CC-BY-NC-4.0 license. Speech teams reach for Canary when transcription and translation are needed from a single model.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
8 GB
Added
Apr 3, 2026

Related Tools

End-to-end speech processing toolkit covering ASR, TTS, and speech translation.

Open SourceSelf HostedOfflineGPU 8GB+
Expert
0.0 (0)

CLI tool that transcribes audio 10x faster using pipeline optimizations.

Open SourceSelf HostedOfflineGPU 6GB+
Easy
0.0 (0)

Established speech recognition toolkit used in research and production systems.

Open SourceSelf HostedOffline
Expert
0.0 (0)

Self-supervised speech representation model by Meta for ASR.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Open-source speaker diarization and voice activity detection toolkit.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Convolution-augmented transformer for speech recognition in ESPnet toolkit.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)
Browse all Speech-to-Text / Speech Recognition tools