Datasets › OVAD benchmark

OVAD benchmark (Open-Vocabulary Attribute Detection)

Introduced by María A. Bravo et al. in Open-vocabulary Attribute Detection23 Nov 2022 archive 2025-07-28

Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner. Existing open-vocabulary tasks focus on object classes, whereas research on object attributes is limited due to the lack of a reliable attribute-focused evaluation benchmark. This paper introduces the Open-Vocabulary Attribute Detection (OVAD) task and the corresponding OVAD benchmark. The objective of the novel task and benchmark is to probe object-level attribute information learned by vision-language models. To this end, we created a clean and densely annotated test set covering 117 attribute classes on the 80 object classes of MS COCO. It includes positive and negative annotations, which enables open-vocabulary evaluation. Overall, the benchmark consists of 1.4 million annotations. For reference, we provide a first baseline method for open-vocabulary attribute detection. Moreover, we demonstrate the benchmark's value by studying the attribute detection performance of several foundation models.

Source: Open-vocabulary Attribute Detection

Image Source: https://arxiv.org/pdf/2211.12914v1.pdf

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Open Vocabulary Attribute Detection OVAD-Box benchmark X-VLM mean average precision 28.0 Multi-Grained Vision Language Pre-Training: Aligning... zengyan-97/x-vlm 7 Compare
Open Vocabulary Attribute Detection OVAD benchmark OvarNet (ViT-B16) mean average precision 27.2 OvarNet: Towards Open-vocabulary Object Attribute Recognition KyanChen/OvarNet 5 Compare
Visual Question Answering (VQA) OVAD benchmark BLIP Contains w. Synonyms 45.70 Open-ended VQA benchmarking of Vision-Language models by... lmb-freiburg/ovqa 1 Compare

Papers archive 2025-07-28

12 shown of 12 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 14. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy 1 1 11 Feb 2024 ran 2 of 2 samples (0 unverified)
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models 17 1 30 Jan 2023 ran 4 of 8 samples (4 unverified; 1 pointer-only for licence)
OvarNet: Towards Open-vocabulary Object Attribute Recognition 1 1 23 Jan 2023 not harvested
Reproducible scaling laws for contrastive language-image learning 5 1 14 Dec 2022 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Open-vocabulary Attribute Detection 1 2 23 Nov 2022 ran 3 of 10 samples (7 unverified)
Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection 1 1 7 Jul 2022 ran 2 of 4 samples (2 unverified)
Localized Vision-Language Matching for Open-vocabulary Object Detection 1 1 12 May 2022 not harvested
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation 9 1 28 Jan 2022 not harvested
Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts 1 1 16 Nov 2021 ran 1 of 1 samples (0 unverified)
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation 6 1 16 Jul 2021 ran 3 of 5 samples (2 unverified; 3 pointer-only for licence)
Learning Transferable Visual Models From Natural Language Supervision 82 1 26 Feb 2021 ran 16 of 20 samples (4 unverified; 16 pointer-only for licence)
Open-Vocabulary Object Detection Using Captions 1 1 20 Nov 2020 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • OVAD benchmark
  • OVAD-Box benchmark

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections