{"url":"/task/described-object-detection","name":"Described Object Detection","slug":"described-object-detection","description_markdown":"Described Object Detection (DOD) detects all instances on each image in the dataset, based on a flexible reference. It is a superset of Open-Vocabulary Object Detection (OVD) and Referring Expression Comprehension (REC). It expands category names to flexible language expressions for OVD and overcomes the limitation of REC only grounding the pre-existing object. Works related to DOD are tracked in [awesome-DOD list on github](https://github.com/Charles-Xie/awesome-described-object-detection).","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":8,"papers_with_code":8,"benchmarks":1,"benchmark_tables_in_archive":1,"benchmark_tables_shown":1,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":1,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/described-object-detection-on-description","slug":"described-object-detection-on-description","dataset":"Description Detection Dataset","dataset_url":"/dataset/description-detection-dataset","rows_in_archive":8,"metrics":["Intra-scenario FULL mAP","Intra-scenario PRES mAP","Intra-scenario ABS mAP"],"first_row_in_archive_order":{"model":"MM-Grounding-DINO","paper_title":"An Open and Comprehensive Pipeline for Unified Object Grounding and Detection","paper_url":"/paper/an-open-and-comprehensive-pipeline-for","paper_date":"2024-01-04","arxiv_id":"2401.02361","code_links":[{"title":"open-mmlab/mmdetection","url":"https://github.com/open-mmlab/mmdetection"},{"title":"cszzshi/SimD","url":"https://github.com/cszzshi/SimD"}],"syntology":{"n":6,"n_ran":4,"n_unverified":2,"n_pointer_only":0}}}],"datasets":[{"url":"/dataset/description-detection-dataset","name":"Description Detection Dataset","full_name":"Description Detection Dataset","num_papers_in_archive":10}],"subtasks":[],"parent_tasks":[{"url":"/task/object-detection","name":"Object Detection"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":8,"of":8,"tagged_in_all":8,"items":[{"url":"/paper/grounded-language-image-pre-training","title":"Grounded Language-Image Pre-training","date":"2021-12-07","arxiv_id":"2112.03857","repositories_listed":3,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":2}},{"url":"/paper/an-open-and-comprehensive-pipeline-for","title":"An Open and Comprehensive Pipeline for Unified Object Grounding and Detection","date":"2024-01-04","arxiv_id":"2401.02361","repositories_listed":2,"syntology":{"n":6,"n_ran":4,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/simple-open-vocabulary-object-detection-with","title":"Simple Open-Vocabulary Object Detection with Vision Transformers","date":"2022-05-12","arxiv_id":"2205.06230","repositories_listed":2,"syntology":null},{"url":"/paper/sphinx-the-joint-mixing-of-weights-tasks-and","title":"SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models","date":"2023-11-13","arxiv_id":"2311.07575","repositories_listed":1,"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":3}},{"url":"/paper/described-object-detection-liberating-object-1","title":"Described Object Detection: Liberating Object Detection with Flexible Expressions","date":"2023-07-24","arxiv_id":"2307.12813","repositories_listed":1,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":3}},{"url":"/paper/cora-adapting-clip-for-open-vocabulary","title":"CORA: Adapting CLIP for Open-Vocabulary Detection with Region Prompting and Anchor Pre-Matching","date":"2023-03-23","arxiv_id":"2303.13076","repositories_listed":1,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/universal-instance-perception-as-object","title":"Universal Instance Perception as Object Discovery and Retrieval","date":"2023-03-12","arxiv_id":"2303.06674","repositories_listed":1,"syntology":{"n":4,"n_ran":3,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/coarse-to-fine-vision-language-pre-training","title":"Coarse-to-Fine Vision-Language Pre-training with Fusion in the Backbone","date":"2022-06-15","arxiv_id":"2206.07643","repositories_listed":1,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":2}}],"syntology_records":7,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}