Dia TTS
Open-source dialogue TTS model by Nari Labs supporting multi-speaker conversations.
About
Dia by Nari Labs is a 1.6B parameter text-to-speech model that generates realistic dialogue directly from a transcript, supporting multiple speakers with distinct voices and nonverbal sounds like laughter and coughing. Output can be conditioned on a reference clip for emotion and tone control, and the model is available through Hugging Face Transformers. The initial release supports English. Released under the Apache 2.0 license.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- Apache-2.0
- Minimum VRAM
- 6 GB
- Added
- Apr 3, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.