MetaVoice

Real-time voice cloning and TTS model with 1.2B parameters by MetaVoice.

Open SourceSelf HostedOffline CapableGPU Required (6GB+ VRAM)
0.0 (0)

About

MetaVoice-1B is a 1.2 billion parameter text-to-speech base model trained on 100,000 hours of speech, released by MetaVoice under the Apache 2.0 license with no usage restrictions. Its focus is voice cloning and expressive delivery: zero-shot cloning works for American and British accents from about 30 seconds of reference audio, while cross-lingual or other-accent cloning is achieved through fine-tuning, reportedly with as little as one minute of data for Indian-accented speakers. The model handles emotional rhythm and tone in English and supports synthesis of arbitrarily long text. Running it requires a GPU with at least 12 GB of VRAM and Python 3.10 or 3.11, and on Ampere or newer NVIDIA architectures the compiled synthesis path reaches faster than real-time generation. The repository includes fine-tuning scripts driven by simple CSV datasets of audio and captions. It suits developers who need a permissively licensed, cloning-capable TTS model they can host themselves.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
6 GB
Added
Apr 3, 2026

Related Tools

Featured

Lightweight and expressive TTS model with 82M parameters for fast local inference.

Open SourceSelf HostedOffline
Easy
4.0 (1)

Conversational TTS model optimized for dialogue and chat applications.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)

Multilingual large voice generation model with full-stack inference, training, and deployment.

Open SourceSelf HostedOfflineGPU
Intermediate
0.0 (0)

Large-scale multilingual TTS model by Alibaba with zero-shot voice cloning.

Open SourceSelf HostedOfflineGPU 8GB+
Advanced
0.0 (0)

Emotion-controllable TTS engine by NetEase with 2000+ voices.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Featured

Transformer-based text-to-audio model by Suno that generates speech, music, and sound effects.

Open SourceSelf HostedOfflineGPU 4GB+
Intermediate
0.0 (0)
Browse all Text-to-Speech (TTS) tools