WhisperSpeech

Text-to-speech system built on top of Whisper encoder representations.

Open SourceSelf HostedOffline CapableGPU Required (6GB+ VRAM)
0.0 (0)

About

WhisperSpeech is an open source text-to-speech system from Collabora built by inverting Whisper: where OpenAI's model maps audio to text, WhisperSpeech runs the pipeline in reverse to generate natural speech. The architecture follows Google's SPEAR-TTS design, using semantic tokens derived from the Whisper encoder, acoustic tokens from Meta's EnCodec codec, and the Vocos vocoder for final audio. Training uses only properly licensed open speech data, starting with the English LibriLight corpus, which keeps the released models safe to build on under the MIT license. The system supports voice cloning from a short reference clip, and the team reports inference running more than ten times faster than real time on a consumer GPU. It installs from PyPI as the whisperspeech package, runs on PyTorch with CUDA, and ships Colab notebooks plus a hosted Hugging Face Space for quick testing. English is the primary language today, with multilingual support an explicit roadmap goal.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Advanced (4/5)
License
MIT
Minimum VRAM
6 GB
Added
Apr 3, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools