Playground v2.5
Open-weight aesthetic-focused image generation model rivaling DALL-E 3.
About
Aesthetic quality is the stated focus of Playground v2.5, an open-weight latent diffusion model for text-to-image generation from Playground. It follows the Stable Diffusion XL architecture with two fixed pre-trained text encoders, OpenCLIP-ViT/G and CLIP-ViT/L, and produces 1024x1024 images across portrait and landscape aspect ratios. A notable technical change is its EDM formulation, supported in Diffusers through dedicated EDM schedulers, which the team credits for crisper detail. On the MJHQ-30K benchmark the roughly 3 billion parameter model reports an FID of 4.48 versus 9.55 for SDXL, and accompanying user studies found its outputs preferred over earlier Playground versions and several closed commercial models. It is distributed on Hugging Face under the Playground v2.5 Community License and loads into Diffusers, AUTOMATIC1111, and ComfyUI, making it a drop-in option for hobbyists and products that want an SDXL-class model tuned for visual appeal.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Category
- Image Generation
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- Minimum VRAM
- 8 GB
- Added
- Apr 3, 2026
Related Tools
State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.
Next-generation image generation model by Black Forest Labs.
Bilingual text-to-image diffusion transformer by Tencent with Chinese and English support.
Zero-shot identity-preserving image generation from a single face photo.
Image prompt adapter for pre-trained text-to-image diffusion models.
Neural network architecture for adding spatial control to diffusion models.