MuseTalk
Real-time high-quality lip-sync model for audio-driven talking face generation.
About
MuseTalk is an audio-driven lip-sync model from Tencent that generates 30+ fps high-resolution talking-face video from a reference image or video and an input audio track. It operates in the latent space of an FT-MSE-VAE and uses a spatio-temporal sampling approach with perceptual, GAN, and sync losses to balance visual quality and lip accuracy. Version 1.5 ships inference, training code, and weights for self-hosted use.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- AI Animation & Motion
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- Minimum VRAM
- 6 GB
- Added
- Apr 3, 2026
Related Tools
Animates a still human photo with 3D SMPL parametric motion guidance extracted from a driving video.
Free markerless motion capture system that works with ordinary cameras and no special hardware.
Audio-driven Tencent model that animates avatar images into emotion-controllable dialogue videos.
ByteDance's audio-conditioned latent diffusion model for lip-syncing video to new speech.
Audio-driven talking head animation from a single image.
Effective whole-body pose estimation with few-shot keypoint detection.