SadTalker
Audio-driven talking head animation from a single image.
About
SadTalker generates a talking-head video from a single portrait image and an audio clip by predicting 3D motion coefficients that drive natural head movement and lip sync. Decoupling expression and pose from the audio improves realism over earlier single-image methods. The project provides a Colab notebook, a Hugging Face Space, and a Stable Diffusion WebUI extension. Developed at Xi'an Jiaotong University and released under the MIT license.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- AI Animation & Motion
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- MIT
- Minimum VRAM
- 6 GB
- Added
- Apr 3, 2026
Related Tools
Animates a still human photo with 3D SMPL parametric motion guidance extracted from a driving video.
Free markerless motion capture system that works with ordinary cameras and no special hardware.
Audio-driven Tencent model that animates avatar images into emotion-controllable dialogue videos.
ByteDance's audio-conditioned latent diffusion model for lip-syncing video to new speech.
Real-time high-quality lip-sync model for audio-driven talking face generation.
Effective whole-body pose estimation with few-shot keypoint detection.