CosyVoice

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOffline CapableGPU Required
0.0 (0)

About

CosyVoice is a family of large text-to-speech models from the FunAudioLLM team that treats speech synthesis as a language modeling problem. The current Fun-CosyVoice 3.0 release, a 0.5B parameter model, covers nine languages including Chinese, English, Japanese, Korean, German, Spanish, French, Italian, and Russian, plus more than 18 Chinese dialects and accents, and supports zero-shot voice cloning from a short reference sample, including cross-lingual cloning. Streaming synthesis works in both directions, text in and audio out, with latency as low as about 150 ms, and instruction inputs control language, dialect, emotion, speaking rate, and volume. Pronunciation inpainting accepts Chinese Pinyin and English CMU phonemes for precise reads, and reported benchmarks include a 0.81 percent character error rate on Chinese and a 1.68 percent word error rate on English test sets. The open source project ships a web UI, Docker images, and a Python API, with weights on ModelScope and Hugging Face, serving developers building production voice features.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Added
May 7, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Compact open-source speech synthesizer supporting 100+ languages.

Open SourceSelf HostedOffline
Beginner
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools