SigLIP
Improved vision-language model by Google using sigmoid loss for contrastive learning.
About
SigLIP by Google is a vision-language model that replaces the softmax contrastive loss of CLIP with a sigmoid loss computed on each image-text pair, which scales better and improves zero-shot performance. It produces strong image and text embeddings for zero-shot classification and retrieval and is available through Hugging Face. It is trained with the big_vision JAX codebase. Released under the Apache 2.0 license.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 4 GB
- Added
- Apr 3, 2026
Related Tools
Contrastive language-image pre-training model by OpenAI for zero-shot visual classification.
Lightweight face recognition and analysis framework wrapping multiple models.
Foundation model for monocular depth estimation by TikTok.
Monocular depth estimation model producing detailed depth maps from single images.
Meta AI research platform for object detection, segmentation, and pose estimation.
Simple and effective multi-object tracking using every detection box.