OWL-ViT
Open-vocabulary object detection model by Google using vision transformers.
About
OWL-ViT, Vision Transformer for Open-World Localization from Google, performs open-vocabulary object detection by accepting free-text queries rather than a fixed label set. It transfers image-text pretraining in the style of CLIP to detection without task-specific training data, so it can localize objects described by arbitrary text. It is distributed within Google's Scenic research codebase for attention-based vision models. Released under the Apache 2.0 license.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- Apache-2.0
- Minimum VRAM
- 6 GB
- Added
- Apr 3, 2026
Related Tools
Contrastive language-image pre-training model by OpenAI for zero-shot visual classification.
Lightweight face recognition and analysis framework wrapping multiple models.
Foundation model for monocular depth estimation by TikTok.
Monocular depth estimation model producing detailed depth maps from single images.
Meta AI research platform for object detection, segmentation, and pose estimation.
Simple and effective multi-object tracking using every detection box.