OpenVoice

Instant voice cloning TTS by MyShell requiring only a short audio reference.

Open SourceSelf HostedOffline CapableGPU Required (4GB+ VRAM)
0.0 (0)

About

OpenVoice, developed by MyShell with researchers from MIT and Tsinghua University, clones a voice from a short audio clip and then generates speech in that voice with detailed control over style parameters including emotion, accent, rhythm, pauses, and intonation. Tone color is decoupled from those style controls, which also enables zero-shot cross-lingual cloning: the target language does not need to appear in the reference audio or the training data. Version 2, released in April 2024, improved audio quality and added native support for English, Spanish, French, Chinese, Japanese, and Korean. The model builds on established TTS and VITS techniques and has powered voice cloning on the myshell.ai platform since May 2023, handling tens of millions of requests within its first months. Both versions are open source under the MIT license, permitting free commercial use, which makes the project a common starting point for developers adding voice cloning to their applications.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
MIT
Minimum VRAM
4 GB
Added
Apr 3, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools