Kandinsky 3

Text-to-image model by Sber AI with latent diffusion architecture.

Open SourceSelf HostedOffline CapableGPU Required (8GB+ VRAM)
0.0 (0)

About

Sber AI's Kandinsky 3 is a large scale text-to-image model based on latent diffusion. The pipeline pairs an 8.6 billion parameter Flan-UL2 text encoder with a 3 billion parameter diffusion U-Net and a MoVQ image encoder and decoder, giving it stronger prompt understanding than earlier Kandinsky releases. Beyond text-to-image generation the repository covers inpainting, outpainting, image fusion, and image variations, while a distilled Kandinsky Flash variant samples in four steps for roughly three times faster inference. Extensions include an HED based ControlNet, an IP-Adapter for image conditioning, and the KandiSuperRes upscaler. Training paid particular attention to Russian cultural content, and the model handles both Russian and English prompts. Code and checkpoints are openly released, with demos on Hugging Face and the FusionBrain web service. Researchers and developers seeking an alternative to Stable Diffusion lineage models are the typical users. Apache 2.0 license.

Reviews (0)

Leave a Review

No reviews yet. Be the first to review!

Details

Price
Free
Platform
Local/Desktop
Difficulty
Intermediate (3/5)
License
Apache-2.0
Minimum VRAM
8 GB
Added
Apr 3, 2026

Related Tools

Featured

State-of-the-art open image generation model by Black Forest Labs with rectified flow transformers.

Open SourceSelf HostedOfflineGPU 12GB+
Intermediate
0.0 (0)
Featured

Next-generation image generation model by Black Forest Labs.

Open SourceSelf HostedOfflineGPU 12GB+
Intermediate
0.0 (0)

Bilingual text-to-image diffusion transformer by Tencent with Chinese and English support.

Open SourceSelf HostedOfflineGPU 12GB+
Intermediate
0.0 (0)

Zero-shot identity-preserving image generation from a single face photo.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)

Image prompt adapter for pre-trained text-to-image diffusion models.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)
Featured

Neural network architecture for adding spatial control to diffusion models.

Open SourceSelf HostedOfflineGPU 8GB+
Intermediate
0.0 (0)
Browse all Image Generation tools