SpeechBrain

All-in-one conversational AI toolkit for speech recognition, enhancement, and more.

Open SourceSelf HostedOffline CapableGPU Required (4GB+ VRAM)
0.0 (0)

About

SpeechBrain is an open-source PyTorch toolkit for conversational AI covering more than 20 speech and text processing tasks, including speech recognition, speaker and language identification, speech separation and enhancement, text-to-speech and vocoding, emotion recognition, spoken language understanding, and voice activity detection, with additional recipes for EEG signal analysis. Training is orchestrated through its Brain class with YAML-based hyperparameter files, and the toolkit supports dynamic batching, mixed precision, and multi-GPU distributed training. Users can start from over 100 pretrained models published on the Hugging Face Hub or reproduce results using more than 200 training recipes spanning 40+ datasets, backed by extensive tutorials and documentation. Development is community driven with roots at Mila and Concordia University, and the code is released under the Apache 2.0 license, which permits commercial use. The audience is largely speech researchers and graduate students, though the pretrained models also see production use as ready-made components.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Advanced (4/5)
License
Apache-2.0
Minimum VRAM
4 GB
Added
Apr 3, 2026

Related Tools

Convolution-augmented transformer for speech recognition in ESPnet toolkit.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

End-to-end speech processing toolkit covering ASR, TTS, and speech translation.

Open SourceSelf HostedOfflineGPU 8GB+
Expert
0.0 (0)

CLI tool that transcribes audio 10x faster using pipeline optimizations.

Open SourceSelf HostedOfflineGPU 6GB+
Easy
0.0 (0)

Established speech recognition toolkit used in research and production systems.

Open SourceSelf HostedOffline
Expert
0.0 (0)

Self-supervised speech representation model by Meta for ASR.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Multilingual ASR model by NVIDIA supporting 4 languages with translation.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)
Browse all Speech-to-Text / Speech Recognition tools