EmotiVoice
Emotion-controllable TTS engine by NetEase with 2000+ voices.
About
EmotiVoice is an open-source text-to-speech engine from NetEase Youdao whose distinguishing feature is emotion control: generated speech can sound happy, excited, sad, angry, or otherwise expressive based on a prompt. It ships more than 2000 preset voices covering English and Chinese, and supports voice cloning when trained on a user's own recordings. Style factors such as pitch, speed, and energy are controllable alongside emotion, following a transformer-based design inspired by PromptTTS. The project provides several ways to run it: an interactive Streamlit web demo, batch synthesis via scripts, and a FastAPI server exposing an OpenAI-compatible TTS HTTP API, with a Docker image for setup on NVIDIA GPUs. Released under the Apache 2.0 license, it is used by developers and content creators who need expressive narration for audiobooks, voiceovers, dubbing, and accessibility features without relying on a commercial TTS service.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 4 GB
- Added
- Apr 3, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Compact open-source speech synthesizer supporting 100+ languages.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.