Papers › Localized Vision-Language Matching for Open-vocabulary Object Detection

Localized Vision-Language Matching for Open-vocabulary Object Detection

12 May 2022arXiv:2205.06160archive 2025-07-28

Maria A. Bravo, Sudhanshu Mittal, Thomas Brox

In this work, we propose an open-vocabulary object detection method that, based on image-caption pairs, learns to detect novel object classes along with a given set of known classes. It is a two-stage training approach that first uses a location-guided image-caption matching technique to learn class labels for both novel and known classes in a weakly-supervised manner and second specializes the model for the object detection task using known class annotations. We show that a simple language model fits better than a large contextualized language model for detecting novel objects. Moreover, we introduce a consistency-regularization technique to better exploit image-caption pair information. Our method compares favorably to existing open-vocabulary detection approaches while being data-efficient. Source code is available at https://github.com/lmb-freiburg/locov .

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

lmb-freiburg/locov officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModelingLanguage ModellingObjectObject DetectionOpen Vocabulary Attribute DetectionOpen Vocabulary Object DetectionOpen World Object DetectionOpen-vocabulary object detectionobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Open Vocabulary Attribute Detection OVAD benchmark LocOv (ResNet50) mean average precision 14.9 #4 of 5 Archive leaderboard report
Open Vocabulary Object Detection MSCOCO LocOv (RN50-C4) AP 0.5 28.6 #27 of 32 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections