DWPose
Effective whole-body pose estimation with few-shot keypoint detection.
About
DWPose implements the ICCV 2023 work on whole-body pose estimation with two-stage distillation, producing models that detect body, foot, face, and hand keypoints from images. The family spans tiny through large variants built on the RTMPose architecture, ranging from about 0.5 to 10 GFLOPS and reaching up to 0.665 average precision on the COCO-WholeBody benchmark, so users can trade accuracy against compute. Distillation lets the smaller models keep much of the teacher's accuracy at a fraction of the size. The project builds on the MMPose framework but also publishes ONNX exports that run without MMPose installed, with checkpoints distributed through Hugging Face and Google Drive. Beyond research use, DWPose has become a common conditioning signal for ControlNet in Stable Diffusion pipelines, where its full-body skeletons guide pose-controlled image and animation generation. The code is released by IDEA Research under the Apache 2.0 license.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- AI Animation & Motion
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Easy (2/5)
- License
- Apache-2.0
- Minimum VRAM
- 4 GB
- Added
- Apr 3, 2026
Related Tools
Free markerless motion capture system that works with ordinary cameras and no special hardware.
Audio-driven Tencent model that animates avatar images into emotion-controllable dialogue videos.
ByteDance's audio-conditioned latent diffusion model for lip-syncing video to new speech.
Audio-driven talking head animation from a single image.
Real-time high-quality lip-sync model for audio-driven talking face generation.
Animates a still human photo with 3D SMPL parametric motion guidance extracted from a driving video.