StyleTTS 2

Style diffusion and adversarial training for human-level TTS with style transfer.

Open SourceSelf HostedOffline CapableGPU Required (6GB+ VRAM)
0.0 (0)

About

StyleTTS 2 approaches text-to-speech with two ideas that pushed it to human-level quality in evaluations: style diffusion and adversarial training with large speech language models. Rather than needing reference audio at synthesis time, it models speaking style as a latent variable sampled through diffusion, so each utterance can receive natural, varied delivery. During training, a pretrained speech language model such as WavLM serves as a discriminator, paired with differentiable duration modeling for end-to-end learning. The Columbia University research project reports speech that surpasses human recordings on the single-speaker LJSpeech benchmark and matches them on the multi-speaker VCTK set, and it supports zero-shot speaker adaptation when trained on LibriTTS. The code is MIT licensed, with pretrained models carrying a condition that synthesized speech be disclosed as such unless the speaker gave permission. Researchers and developers use it for single and multi-speaker synthesis and style transfer from reference audio.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Advanced (4/5)
License
MIT
Minimum VRAM
6 GB
Added
Apr 3, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools