ChatTTS

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOffline CapableGPU Required (4GB+ VRAM)
0.0 (0)

About

ChatTTS is a generative text-to-speech model built for dialogue scenarios such as LLM assistant voices rather than long-form narration. Trained on roughly 40,000 hours of Chinese and English speech, the autoregressive model produces conversational prosody with natural pauses and can insert paralinguistic elements like laughter through fine-grained control tokens at the word and sentence level. It supports multiple speakers, mixed language input, and streaming audio generation built on a DVAE-based pipeline, and it can synthesize a 30 second clip on a GPU with about 4 GB of VRAM. The code is published under AGPLv3+, while the released weights carry a CC BY-NC 4.0 license restricting them to non-commercial use; the authors state the open model is intended for academic purposes and note that compressed quality and embedded noise were added to deter misuse. Researchers and developers exploring conversational agents in Chinese and English make up its main audience.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
Minimum VRAM
4 GB
Added
Apr 3, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Compact open-source speech synthesizer supporting 100+ languages.

Open SourceSelf HostedOffline
Beginner
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools