I2VGen-XL
Image-to-video generation model by Alibaba DAMO Academy.
About
I2VGen-XL is an image-to-video synthesis method from Alibaba's Tongyi Lab, published within the VGen video generation codebase. It animates a still image into high-definition video through cascaded diffusion: a base stage establishes semantics and motion coherent with the input image, and a refinement stage sharpens detail and raises resolution, optionally guided by a text caption. Trained on WebVid-10M and LAION-400M, the released models are intended for research and non-commercial use, and the authors note weaker results on anime-style images and black backgrounds due to training data coverage. Inference is available through the VGen repository along with a local Gradio demo, plus hosted versions on ModelScope, Hugging Face, and Replicate. Researchers exploring image conditioning in video diffusion use it both as a baseline and as a component of VGen, which bundles several related Tongyi Lab video generation models in one codebase.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Apache-2.0
- Minimum VRAM
- 12 GB
- Added
- Apr 3, 2026
Related Tools
Open-source video generation model by Tencent with text and image conditioning.
Updated CogVideo model by Zhipu AI with improved video quality.
Infinite-length music-driven video generation with visual conditioning.
Text-to-video generation framework with cascaded latent diffusion.
Open-source video generation model with controllable camera and subject motion.
Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.