Mochi 1
Open-source video generation model by Genmo with state-of-the-art motion quality.
About
Genmo released Mochi 1 as an open text-to-video model with an emphasis on motion quality and prompt adherence. The generator is a 10 billion parameter Asymmetric Diffusion Transformer (AsymmDiT), which gives the visual stream roughly four times the parameters of the text stream, paired with AsymmVAE, a 362 million parameter video autoencoder that compresses input 128 times through 8x8 spatial and 6x temporal reduction. Prompts are encoded with a single T5-XXL model. Output is 480p video, and hardware demands are substantial: around 60 GB of VRAM for single-GPU inference with an H100 recommended, though ComfyUI optimizations bring requirements under 20 GB. Weights are published on Hugging Face and by torrent under the Apache 2.0 license, and the repository includes a LoRA trainer for fine-tuning on a single H100 or A100. Researchers and technically equipped creators use it as one of the more capable fully open video generation baselines available for local experimentation.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Video Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Advanced (4/5)
- License
- Apache-2.0
- Minimum VRAM
- 24 GB
- Added
- Apr 3, 2026
Related Tools
Open-source video generation model by Tencent with text and image conditioning.
Image-to-video generation model by Alibaba DAMO Academy.
Updated CogVideo model by Zhipu AI with improved video quality.
Infinite-length music-driven video generation with visual conditioning.
Text-to-video generation framework with cascaded latent diffusion.
Open-source text-to-video model by Zhipu AI/Tsinghua with 2B and 5B variants.