Parler-TTS

TTS model that generates speech from text descriptions of the desired voice.

Open SourceSelf HostedOffline CapableGPU Required (6GB+ VRAM)
0.0 (0)

About

Rather than selecting from fixed voices, Parler-TTS generates speech from a plain-language description of the desired speaker, such as a female voice with a warm tone speaking quickly. Maintained by Hugging Face, the project reproduces research by Lyth and King from Stability AI and the University of Edinburgh on natural language guidance of text-to-speech, and it is fully open: training code, datasets, and model weights are all released under Apache 2.0. Two checkpoints trained on 45,000 hours of audiobook audio shipped in August 2024, Mini at 880M parameters and Large at 2.3B, along with 34 named speakers for consistent voice reproduction across generations. Descriptions can specify gender, pitch, speaking rate, and even recording quality, while punctuation shapes prosody. Compatibility with SDPA, Flash Attention 2, and model compilation speeds up inference. Researchers and developers use it as a controllable, reproducible alternative to closed text-to-speech APIs.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
6 GB
Added
Apr 3, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools