Papers › Simple Open-Vocabulary Object Detection with Vision Transformers

Simple Open-Vocabulary Object Detection with Vision Transformers

12 May 2022arXiv:2205.06230archive 2025-07-28

Matthias Minderer, Alexey Gritsenko, Austin Stone, Maxim Neumann, Dirk Weissenborn, Alexey Dosovitskiy, Aravindh Mahendran, Anurag Arnab, Mostafa Dehghani, Zhuoran Shen, Xiao Wang, Xiaohua Zhai, Thomas Kipf, Neil Houlsby

Combining simple architectures with large-scale pre-training has led to massive improvements in image classification. For object detection, pre-training and scaling approaches are less well established, especially in the long-tailed and open-vocabulary setting, where training data is relatively scarce. In this paper, we propose a strong recipe for transferring image-text models to open-vocabulary object detection. We use a standard Vision Transformer architecture with minimal modifications, contrastive image-text pre-training, and end-to-end detection fine-tuning. Our analysis of the scaling properties of this setup shows that increasing image-level pre-training and model size yield consistent improvements on the downstream detection task. We provide the adaptation strategies and regularizations needed to attain very strong performance on zero-shot text-conditioned and one-shot image-conditioned object detection. Code and models are available on GitHub.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Described Object DetectionImage ClassificationObjectObject DetectionOne-Shot Object DetectionOpen Vocabulary Object DetectionOpen-vocabulary object detectionimage-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Described Object Detection Description Detection Dataset OWL-ViT-base Intra-scenario ABS mAP 8.8 #7 of 8 Archive leaderboard report
Described Object Detection Description Detection Dataset OWL-ViT-base Intra-scenario FULL mAP 8.6 #7 of 8 Archive leaderboard report
Described Object Detection Description Detection Dataset OWL-ViT-base Intra-scenario PRES mAP 8.5 #7 of 8 Archive leaderboard report
One-Shot Object Detection COCO (Common Objects in Context) OWL-ViT (R50+H/32) AP 0.5 41.8 #1 of 4 Archive leaderboard report
Open Vocabulary Object Detection LVIS v1.0 OWL-ViT (CLIP-L/14) AP novel-LVIS base training 25.6 #15 of 28 Archive leaderboard report
Open Vocabulary Object Detection LVIS v1.0 OWL-ViT (CLIP-L/14) AP novel-Unrestricted open-vocabulary training 31.2 #15 of 28 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformerVision Transformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections