VoiceCraft
Zero-shot speech editing and TTS using neural codec language models.
About
VoiceCraft from UT Austin is a token-infilling neural codec language model that performs both speech editing and zero-shot text-to-speech on in-the-wild audio such as audiobooks, podcasts, and videos. It can edit existing recordings or clone an unseen voice from a few seconds of reference, and ships code and demos. Inference uses a GPU with around 8 GB of VRAM. Distributed as open-source research.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- Minimum VRAM
- 8 GB
- Added
- Apr 3, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.