PixArt-Sigma
Efficient text-to-image diffusion transformer producing 4K resolution images.
About
Generating images directly at up to 4K resolution, PixArt-Sigma is a text-to-image Diffusion Transformer developed by researchers from Huawei Noah's Ark Lab together with DLUT, HKU, and HKUST. Despite that output range the transformer holds only about 0.6 billion parameters, an efficiency the team attributes to weak-to-strong training, a strategy that evolves the model from its PixArt-alpha predecessor using higher quality data rather than training from scratch. Compared with PixArt-alpha it extends the T5 text token length from 120 to 300 for richer prompts and swaps in the SDXL VAE, supporting resolutions from 256 up through 2K and 4K across varied aspect ratios. Training and inference code, checkpoints at multiple resolutions, a Gradio demo, and Hugging Face Diffusers integration are all released openly. The model draws researchers studying efficient diffusion transformers and practitioners who want high-resolution generation on modest compute budgets.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Image Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- Minimum VRAM
- 8 GB
- Added
- Apr 3, 2026
Related Tools
State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.
Next-generation image generation model by Black Forest Labs.
Bilingual text-to-image diffusion transformer by Tencent with Chinese and English support.
Zero-shot identity-preserving image generation from a single face photo.
Image prompt adapter for pre-trained text-to-image diffusion models.
Neural network architecture for adding spatial control to diffusion models.