ModelScope Text-to-Video
Text-to-video generation model by Alibaba DAMO Academy on ModelScope.
About
ModelScope Text-to-Video is an early open text-to-video diffusion model from Alibaba DAMO Academy, distributed through the ModelScope model-as-a-service library. It generates short clips from English text prompts and was one of the first openly released models of its kind. The core ModelScope library provides unified interfaces for loading and running the model. Inference benefits from a GPU with 12 GB or more of VRAM.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- Minimum VRAM
- 12 GB
- Added
- Apr 3, 2026
Related Tools
Open-source video generation model by Tencent with text and image conditioning.
Image-to-video generation model by Alibaba DAMO Academy.
Updated CogVideo model by Zhipu AI with improved video quality.
Infinite-length music-driven video generation with visual conditioning.
Text-to-video generation framework with cascaded latent diffusion.
Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.