VibeVoice
Neural TTS model by Microsoft with expressive speech synthesis.
About
VibeVoice by Microsoft Research is a neural text-to-speech model aimed at expressive, natural synthesis with controllable prosody and emotion. The project later added VibeVoice-ASR, a unified speech-to-text model available through the Hugging Face Transformers library. Inference runs on a GPU and demo notebooks are provided. Released under the MIT license with model collections published on Hugging Face.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- MIT
- Minimum VRAM
- 6 GB
- Added
- Apr 3, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.