StyleTTS 2
Style diffusion and adversarial training for human-level TTS with style transfer.
About
StyleTTS 2 approaches text-to-speech with two ideas that pushed it to human-level quality in evaluations: style diffusion and adversarial training with large speech language models. Rather than needing reference audio at synthesis time, it models speaking style as a latent variable sampled through diffusion, so each utterance can receive natural, varied delivery. During training, a pretrained speech language model such as WavLM serves as a discriminator, paired with differentiable duration modeling for end-to-end learning. The Columbia University research project reports speech that surpasses human recordings on the single-speaker LJSpeech benchmark and matches them on the multi-speaker VCTK set, and it supports zero-shot speaker adaptation when trained on LibriTTS. The code is MIT licensed, with pretrained models carrying a condition that synthesized speech be disclosed as such unless the speaker gave permission. Researchers and developers use it for single and multi-speaker synthesis and style transfer from reference audio.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- MIT
- Minimum VRAM
- 6 GB
- Added
- Apr 3, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.