{"url":"/dataset/description-detection-dataset","name":"Description Detection Dataset","full_name":"Description Detection Dataset","description_markdown":"**Description Detection Dataset** ($D^3$, /dikju:b/) is an attempt at creating a next-generation object detection dataset. Unlike traditional detection datasets, the class names of the objects are no longer simple nouns or noun phrases, but rather complex and descriptive, such as `a dog not being held by a leash`. For each image in the dataset, any object that matches the description is annotated. The dataset provides annotations such as bounding boxes and finely crafted instance masks.It comprises of 422 well-designed descriptions and 24,282 positive object-description pairs.\r\n\r\nThe dataset is meant for the Described Object Detection (DOD) task. OVD detects object based on category name, and each category can have zero to multiple instances; REC grounds one region based on a language description, whether the object truly exits or not; DOD detects all instances on each image in the dataset, based on a flexible reference.","description_withheld":null,"homepage":"https://github.com/shikras/d-cube","introduced_date":"2023-07-24","introduced_date_note":null,"introduced_by":null,"license":{"name":"Creative Commons Attribution-NonCommercial 4.0 International License","url":"https://github.com/shikras/d-cube/blob/main/LICENSE"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Object Detection","url":"/task/object-detection","datasets_with_task":"/datasets/task/object-detection"},{"name":"Referring Expression Comprehension","url":"/task/referring-expression-comprehension","datasets_with_task":"/datasets/task/referring-expression-comprehension"},{"name":"Open Vocabulary Object Detection","url":"/task/open-vocabulary-object-detection","datasets_with_task":"/datasets/task/open-vocabulary-object-detection"},{"name":"Described Object Detection","url":"/task/described-object-detection","datasets_with_task":"/datasets/task/described-object-detection"}],"languages":[],"variants":["Description Detection Dataset"],"data_loaders":[{"repo":"https://github.com/shikras/d-cube","url":"https://github.com/shikras/d-cube/blob/main/doc.md","frameworks":["tf","pytorch","jax"]}],"num_papers_in_archive":10,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/described-object-detection-on-description","task":"Described Object Detection","dataset_variant":"Description Detection Dataset","rows":8,"metrics":["Intra-scenario FULL mAP","Intra-scenario PRES mAP","Intra-scenario ABS mAP"],"first_row_in_archive_order":{"model":"MM-Grounding-DINO","paper":"/paper/an-open-and-comprehensive-pipeline-for","metrics":{"Intra-scenario ABS mAP":"26.0","Intra-scenario FULL mAP":"22.9","Intra-scenario PRES mAP":"21.9"},"code_links":[{"title":"open-mmlab/mmdetection","url":"https://github.com/open-mmlab/mmdetection"},{"title":"cszzshi/SimD","url":"https://github.com/cszzshi/SimD"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/an-open-and-comprehensive-pipeline-for","title":"An Open and Comprehensive Pipeline for Unified Object Grounding and Detection","date":"2024-01-04","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":6,"samples_ran":4,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/sphinx-the-joint-mixing-of-weights-tasks-and","title":"SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models","date":"2023-11-13","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/described-object-detection-liberating-object-1","title":"Described Object Detection: Liberating Object Detection with Flexible Expressions","date":"2023-07-24","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/cora-adapting-clip-for-open-vocabulary","title":"CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-Matching","date":"2023-03-23","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/universal-instance-perception-as-object","title":"Universal Instance Perception as Object Discovery and Retrieval","date":"2023-03-12","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":4,"samples_ran":3,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/coarse-to-fine-vision-language-pre-training","title":"Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone","date":"2022-06-15","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/simple-open-vocabulary-object-detection-with","title":"Simple Open-Vocabulary Object Detection with Vision Transformers","date":"2022-05-12","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/grounded-language-image-pre-training","title":"Grounded Language-Image Pre-training","date":"2021-12-07","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":7,"samples_harvested":24,"samples_ran":18,"samples_unverified":6,"pointer_only_for_licence":10,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}