Kokoro TTS
Lightweight and expressive TTS model with 82M parameters for fast local inference.
About
At just 82 million parameters, Kokoro is an open weight text-to-speech model that produces speech quality comparable to much larger systems while running faster and at lower cost. The accompanying Python library installs via pip and handles inference, including streaming audio output at a 24 kHz sample rate. Kokoro speaks eight languages, American and British English, Spanish, French, Hindi, Italian, Japanese, Brazilian Portuguese, and Mandarin Chinese, with multiple voices available. It depends on espeak-ng for phonemization and runs on CPU, CUDA GPUs, or Apple Silicon, making real time synthesis feasible on consumer hardware without a dedicated graphics card. Because the weights are Apache licensed, commercial deployment is permitted, and the model shows up in production voice features as well as hobby projects where cloud TTS would be too slow or too expensive. Developers building reading apps, voice assistants, and audio pipelines are typical users.
Reviews (1)
Leave a Review
Decent tool but just doesn't have the best sounding voices
Details
- Category
- Text-to-Speech (TTS)
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- Apache-2.0
- Added
- Apr 3, 2026
Related Tools
Conversational TTS model optimized for dialogue and chat applications.
Multilingual large voice generation model with full-stack inference, training, and deployment.
Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.
Emotion-controllable TTS engine by NetEase with 2000+ voices.
Compact open-source speech synthesizer supporting 100+ languages.
Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.
Mentioned in
Building Real-Time Voice Agents: TEN, Pipecat, and LiveKit
A working guide to real-time voice agent stacks: latency budgets, turn detection, interruption handling,...
Max P
Open-Weight Text to Speech Models in 2026: The XTTS Successors
A working developer's comparison of Kokoro, Zonos, Kyutai TTS, F5-TTS, Piper, Chatterbox and the rest:...
Billy C