Matcha-TTS
Fast TTS with conditional flow matching for efficient speech synthesis.
About
Developed at KTH Royal Institute of Technology and published at ICASSP 2024, Matcha-TTS is a non-autoregressive text-to-speech model that uses optimal-transport conditional flow matching, a technique similar to rectified flows, to cut the number of ODE solver steps needed for synthesis. The result is fast, probabilistic speech generation with a compact memory footprint and natural-sounding output, with controls for speaking rate, sampling temperature, and solver step count to trade speed against quality. The package installs with pip and offers a command line tool, a Gradio web demo, and a Jupyter notebook, plus full training scripts for custom datasets and ONNX export for deployment. Released as open source under the MIT license, it serves speech researchers as a flow-matching baseline and developers who need lightweight local TTS, and its approach has influenced later speech systems that adopt flow matching.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- MIT
- Minimum VRAM
- 4 GB
- Added
- Apr 3, 2026
Related Tools
Lightweight and expressive TTS model with 82M parameters for fast local inference.
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.