CosyVoice 2

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOffline CapableGPU Required (8GB+ VRAM)
0.0 (0)

About

The second generation of the CosyVoice line, CosyVoice 2 is a 0.5B parameter speech synthesis model developed by the FunAudioLLM team at Alibaba. Its headline addition over the original 300M model is streaming: the model supports bidirectional streaming synthesis with first-packet latency reported as low as 150 ms while keeping quality close to offline generation, which makes it practical for live voice agents. Like the rest of the family it performs zero-shot voice cloning from a brief audio prompt, cross-lingual synthesis in which a cloned voice speaks another language, and instruction-controlled generation covering emotion, dialect, speaking rate, and volume. The model handles Chinese, English, Japanese, and Korean along with Chinese dialects, and the repository provides training and inference code, a web UI, Docker deployment, and weights on ModelScope and Hugging Face under the Apache 2.0 license. Within the same repository it has since been followed by Fun-CosyVoice 3.0, which broadens language coverage further.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Advanced (4/5)
License
Apache-2.0
Minimum VRAM
8 GB
Added
Apr 3, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Compact open-source speech synthesizer supporting 100+ languages.

Open SourceSelf HostedOffline
Beginner
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools