Papers › Simple Open-Vocabulary Object Detection with Vision Transformers
Simple Open-Vocabulary Object Detection with Vision Transformers
Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, Xiao Wang, Xiaohua Zhai, Thomas Kipf, Neil Houlsby
Combining simple architectures with large-scale pre-training has led to massive improvements in image classification. For object detection, pre-training and scaling approaches are less well established, especially in the long-tailed and open-vocabulary setting, where training data is relatively scarce. In this paper, we propose a strong recipe for transferring image-text models to open-vocabulary object detection. We use a standard Vision Transformer architecture with minimal modifications, contrastive image-text pre-training, and end-to-end detection fine-tuning. Our analysis of the scaling properties of this setup shows that increasing image-level pre-training and model size yield consistent improvements on the downstream detection task. We provide the adaptation strategies and regularizations needed to attain very strong performance on zero-shot text-conditioned and one-shot image-conditioned object detection. Code and models are available on GitHub.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Described Object Detection | Description Detection Dataset | OWL-ViT-base | Intra-scenario ABS mAP | 8.8 | #7 of 8 | Archive leaderboard | report |
| Described Object Detection | Description Detection Dataset | OWL-ViT-base | Intra-scenario FULL mAP | 8.6 | #7 of 8 | Archive leaderboard | report |
| Described Object Detection | Description Detection Dataset | OWL-ViT-base | Intra-scenario PRES mAP | 8.5 | #7 of 8 | Archive leaderboard | report |
| One-Shot Object Detection | COCO (Common Objects in Context) | OWL-ViT (R50+H/32) | AP 0.5 | 41.8 | #1 of 4 | Archive leaderboard | report |
| Open Vocabulary Object Detection | LVIS v1.0 | OWL-ViT (CLIP-L/14) | AP novel-LVIS base training | 25.6 | #15 of 28 | Archive leaderboard | report |
| Open Vocabulary Object Detection | LVIS v1.0 | OWL-ViT (CLIP-L/14) | AP novel-Unrestricted open-vocabulary training | 31.2 | #15 of 28 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections