Würstchen
Efficient latent diffusion model using highly compressed latent space.
About
Wurstchen is a text-to-image latent diffusion framework that moves the expensive text-conditional stage into a highly compressed latent space. It adds a second compression stage so that Stage A and Stage B reconstruct images while Stage C learns the text-conditional model in a low-dimensional space, reaching a 42x compression factor and reducing training and inference cost while keeping image quality. Created by Pablo Pernias under the MIT license.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Image Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- MIT
- Minimum VRAM
- 4 GB
- Added
- Apr 3, 2026
Related Tools
State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.
Next-generation image generation model by Black Forest Labs.
Bilingual text-to-image diffusion transformer by Tencent with Chinese and English support.
Zero-shot identity-preserving image generation from a single face photo.
Image prompt adapter for pre-trained text-to-image diffusion models.
Neural network architecture for adding spatial control to diffusion models.