IndexTTS
Zero-shot TTS model with high naturalness and speaker similarity.
About
IndexTTS is a zero-shot text-to-speech model that synthesizes speech matching a short reference clip's voice. The IndexTTS2 release adds a method for controlling the duration of generated speech in autoregressive synthesis, which helps tasks like video dubbing that need audio-visual synchronization, alongside emotionally expressive output. It supports multiple languages and runs on a GPU. Released as an open-source model with inference code.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- Minimum VRAM
- 6 GB
- Added
- Apr 3, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.