YOLO-World
Open-vocabulary real-time object detection using YOLO with text prompts.
About
YOLO-World from Tencent AI Lab adds open-vocabulary detection to the YOLO architecture by fusing vision and language features, so it can detect objects described by arbitrary text prompts without retraining on those classes. It keeps real-time inference speed, which separates it from heavier open-vocabulary detectors. Pretrained weights and demos are available on Hugging Face and Roboflow. Released under the GPL-3.0 license.
Reviews (0)
Leave a Review
No reviews yet. Be the first to review!
Details
- Price
- Free
- Platform
- Local/Desktop
- Difficulty
- Intermediate (3/5)
- License
- GPL-3.0
- Minimum VRAM
- 4 GB
- Added
- Apr 3, 2026
Related Tools
Contrastive language-image pre-training model by OpenAI for zero-shot visual classification.
Lightweight face recognition and analysis framework wrapping multiple models.
Foundation model for monocular depth estimation by TikTok.
Monocular depth estimation model producing detailed depth maps from single images.
Meta AI research platform for object detection, segmentation, and pose estimation.
Simple and effective multi-object tracking using every detection box.