Mozilla TTS

Deep learning TTS library by Mozilla with Tacotron and WaveRNN implementations.

Open SourceSelf HostedOffline CapableGPU Required (4GB+ VRAM)
0.0 (0)

About

Before Coqui, there was Mozilla TTS: a deep learning text-to-speech library that collected research-grade implementations of text-to-spectrogram models (Tacotron, Tacotron2, Glow-TTS, Speedy-Speech) and neural vocoders (MelGAN, Multiband-MelGAN, ParallelWaveGAN, WaveGrad, WaveRNN) behind a common training and inference framework, along with a GE2E-based speaker encoder and dataset quality analysis tools. Pretrained models were released in PyTorch, TensorFlow, and TFLite formats, and the stack saw use in products and research projects across more than 20 languages. Training expects a GPU while inference can run on CPU. The repository is licensed under MPL 2.0, but it is no longer where active development happens: the core team continued the codebase as Coqui TTS after leaving Mozilla. It remains relevant as a well-documented historical codebase for studying classic neural TTS architectures and for projects still pinned to its released models.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
MPL-2.0
Minimum VRAM
4 GB
Added
Apr 3, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools