{"url":"/dataset/ovad-benchmark","name":"OVAD benchmark","full_name":"Open-Vocabulary Attribute Detection","description_markdown":"Vision-language modeling has enabled open-vocabulary tasks where predictions can be queried using any text prompt in a zero-shot manner. Existing open-vocabulary tasks focus on object classes, whereas research on object attributes is limited due to the lack of a reliable attribute-focused evaluation benchmark. This paper introduces the Open-Vocabulary Attribute Detection (OVAD) task and the corresponding OVAD benchmark. The objective of the novel task and benchmark is to probe object-level attribute information learned by vision-language models. To this end, we created a clean and densely annotated test set covering 117 attribute classes on the 80 object classes of MS COCO. It includes positive and negative annotations, which enables open-vocabulary evaluation. Overall, the benchmark consists of 1.4 million annotations. For reference, we provide a first baseline method for open-vocabulary attribute detection. Moreover, we demonstrate the benchmark's value by studying the attribute detection performance of several foundation models.\r\n\r\nSource: [Open-vocabulary Attribute Detection](https://arxiv.org/pdf/2211.12914v1.pdf)\r\n\r\nImage Source: [https://arxiv.org/pdf/2211.12914v1.pdf](https://arxiv.org/pdf/2211.12914v1.pdf)","description_withheld":null,"homepage":"https://ovad-benchmark.github.io/","introduced_date":"2022-11-23","introduced_date_note":null,"introduced_by":{"paper":"/paper/open-vocabulary-attribute-detection","title":"Open-vocabulary Attribute Detection","first_author":"María A. Bravo","url":null},"license":{"name":"Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License","url":"http://creativecommons.org/licenses/by-nc-sa/4.0/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Visual Question Answering (VQA)","url":"/task/visual-question-answering","datasets_with_task":"/datasets/task/visual-question-answering"},{"name":"Language Modelling","url":"/task/language-modelling","datasets_with_task":"/datasets/task/language-modelling"},{"name":"Open Vocabulary Object Detection","url":"/task/open-vocabulary-object-detection","datasets_with_task":"/datasets/task/open-vocabulary-object-detection"},{"name":"Open Vocabulary Attribute Detection","url":"/task/open-vocabulary-attribute-detection","datasets_with_task":"/datasets/task/open-vocabulary-attribute-detection"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["OVAD benchmark","OVAD-Box benchmark"],"data_loaders":[],"num_papers_in_archive":14,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/open-vocabulary-attribute-detection-on-ovad-1","task":"Open Vocabulary Attribute Detection","dataset_variant":"OVAD-Box benchmark","rows":7,"metrics":["mean average precision"],"first_row_in_archive_order":{"model":"X-VLM","paper":"/paper/multi-grained-vision-language-pre-training","metrics":{"mean average precision":"28.0"},"code_links":[{"title":"zengyan-97/x-vlm","url":"https://github.com/zengyan-97/x-vlm"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/open-vocabulary-attribute-detection-on-ovad","task":"Open Vocabulary Attribute Detection","dataset_variant":"OVAD benchmark","rows":5,"metrics":["mean average precision"],"first_row_in_archive_order":{"model":"OvarNet (ViT-B16)","paper":"/paper/ovarnet-towards-open-vocabulary-object","metrics":{"mean average precision":"27.2"},"code_links":[{"title":"KyanChen/OvarNet","url":"https://github.com/KyanChen/OvarNet"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-question-answering-vqa-on-ovad","task":"Visual Question Answering (VQA)","dataset_variant":"OVAD benchmark","rows":1,"metrics":["Contains w. Synonyms","ExactMatch w. Synonyms"],"first_row_in_archive_order":{"model":"BLIP","paper":"/paper/open-ended-vqa-benchmarking-of-vision","metrics":{"Contains w. Synonyms":"45.70","ExactMatch w. Synonyms":"36.99"},"code_links":[{"title":"lmb-freiburg/ovqa","url":"https://github.com/lmb-freiburg/ovqa"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/open-ended-vqa-benchmarking-of-vision","title":"Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy","date":"2024-02-11","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/blip-2-bootstrapping-language-image-pre","title":"BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models","date":"2023-01-30","rows_on_this_dataset":1,"code_links":17,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":4,"samples_unverified":4,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/ovarnet-towards-open-vocabulary-object","title":"OvarNet: Towards Open-vocabulary Object Attribute Recognition","date":"2023-01-23","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/reproducible-scaling-laws-for-contrastive","title":"Reproducible scaling laws for contrastive language-image learning","date":"2022-12-14","rows_on_this_dataset":1,"code_links":5,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/open-vocabulary-attribute-detection","title":"Open-vocabulary Attribute Detection","date":"2022-11-23","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":10,"samples_ran":3,"samples_unverified":7,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/bridging-the-gap-between-object-and-image","title":"Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection","date":"2022-07-07","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":2,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/localized-vision-language-matching-for-open","title":"Localized Vision-Language Matching for Open-vocabulary Object Detection","date":"2022-05-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/blip-bootstrapping-language-image-pre","title":"BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation","date":"2022-01-28","rows_on_this_dataset":1,"code_links":9,"syntology":null},{"paper":"/paper/multi-grained-vision-language-pre-training","title":"Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts","date":"2021-11-16","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/align-before-fuse-vision-and-language","title":"Align before Fuse: Vision and Language Representation Learning with Momentum Distillation","date":"2021-07-16","rows_on_this_dataset":1,"code_links":6,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":3,"samples_unverified":2,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/learning-transferable-visual-models-from","title":"Learning Transferable Visual Models From Natural Language Supervision","date":"2021-02-26","rows_on_this_dataset":1,"code_links":82,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":20,"samples_ran":16,"samples_unverified":4,"pointer_only_for_licence":16,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/open-vocabulary-object-detection-using","title":"Open-Vocabulary Object Detection Using Captions","date":"2020-11-20","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":8,"samples_harvested":53,"samples_ran":34,"samples_unverified":19,"pointer_only_for_licence":23,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}