Wan 2.1
Open-source video generation model suite by Alibaba with text-to-video and image-to-video.
About
Alibaba's Wan 2.1 established the Wan line as one of the leading open video generation suites. It spans several diffusion transformer models: text-to-video in 1.3B and 14B parameter sizes at 480p and 720p, a 14B image-to-video model, a first-and-last-frame-to-video variant, VACE models for video editing and creation, and text-to-image, with the notable ability to render legible Chinese and English text inside generated footage. Underpinning the suite is Wan-VAE, a 3D causal variational autoencoder that encodes and decodes 1080p video of arbitrary length while preserving temporal information. The 1.3B text-to-video model needs only about 8 GB of VRAM, putting local generation within reach of consumer cards like the RTX 4090. Everything is open source under Apache 2.0, with weights on Hugging Face and ModelScope, inference code and Gradio demos in the repository, and integrations in Diffusers and ComfyUI. Researchers and creators run it for self-hosted video generation without usage restrictions, and Wan 2.2 continues the series.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Apache-2.0
- Minimum VRAM
- 12 GB
- Added
- Apr 3, 2026
Related Tools
Open-source video generation model by Tencent with text and image conditioning.
Image-to-video generation model by Alibaba DAMO Academy.
Updated CogVideo model by Zhipu AI with improved video quality.
Infinite-length music-driven video generation with visual conditioning.
Text-to-video generation framework with cascaded latent diffusion.
Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.