HunyuanVideo
Open-source video generation model by Tencent with text and image conditioning.
About
Tencent's HunyuanVideo is a video foundation model with over 13 billion parameters, among the largest video generators with open weights. Text and video tokens are processed first in separate streams and then merged in a dual-stream to single-stream transformer, while a decoder-only multimodal LLM replaces the usual CLIP or T5 text encoder and a 3D VAE compresses videos across time, space, and channels. A prompt rewrite module adapts user input into phrasing the model prefers. It generates clips up to 5 seconds and 129 frames at 540p to 720p in multiple aspect ratios, and in professional human evaluation it compared favorably with closed systems including Runway Gen-3. Hardware demands are high, about 45 GB of GPU memory minimum and 80 GB recommended, though FP8 quantized weights, CPU offloading, and multi-GPU parallelism via xDiT reduce the burden. Distributed under the Tencent Hunyuan Community License, it is integrated into Diffusers, ComfyUI, and GGUF quantization workflows.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Tencent Community
- Minimum VRAM
- 24 GB
- Added
- Apr 3, 2026
Related Tools
Image-to-video generation model by Alibaba DAMO Academy.
Updated CogVideo model by Zhipu AI with improved video quality.
Infinite-length music-driven video generation with visual conditioning.
Text-to-video generation framework with cascaded latent diffusion.
Open-source video generation model with controllable camera and subject motion.
Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.