Datasets › OVAD benchmark
OVAD benchmark (Open-Vocabulary Attribute Detection)
Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner. Existing open-vocabulary tasks focus on object classes, whereas research on object attributes is limited due to the lack of a reliable attribute-focused evaluation benchmark. This paper introduces the Open-Vocabulary Attribute Detection (OVAD) task and the corresponding OVAD benchmark. The objective of the novel task and benchmark is to probe object-level attribute information learned by vision-language models. To this end, we created a clean and densely annotated test set covering 117 attribute classes on the 80 object classes of MS COCO. It includes positive and negative annotations, which enables open-vocabulary evaluation. Overall, the benchmark consists of 1.4 million annotations. For reference, we provide a first baseline method for open-vocabulary attribute detection. Moreover, we demonstrate the benchmark's value by studying the attribute detection performance of several foundation models.
Source: Open-vocabulary Attribute Detection
Image Source: https://arxiv.org/pdf/2211.12914v1.pdf
Benchmarks archive 2025-07-28
All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Open Vocabulary Attribute Detection | OVAD-Box benchmark | X-VLM mean average precision 28.0 | Multi-Grained Vision Language Pre-Training: Aligning... | zengyan-97/x-vlm | 7 | Compare |
| Open Vocabulary Attribute Detection | OVAD benchmark | OvarNet (ViT-B16) mean average precision 27.2 | OvarNet: Towards Open-vocabulary Object Attribute Recognition | KyanChen/OvarNet | 5 | Compare |
| Visual Question Answering (VQA) | OVAD benchmark | BLIP Contains w. Synonyms 45.70 | Open-ended VQA benchmarking of Vision-Language models by... | lmb-freiburg/ovqa | 1 | Compare |
Papers archive 2025-07-28
12 shown of 12 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 14. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- OVAD benchmark
- OVAD-Box benchmark
2 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections