{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/context-and-attribute-grounded-dense","title":"Context and Attribute Grounded Dense Captioning","arxiv_id":"1904.01410","date":"2019-04-02","proceeding":"CVPR 2019 6","authors":["Guojun Yin","Lu Sheng","Bin Liu","Nenghai Yu","Xiaogang Wang","Jing Shao"],"abstract":"Dense captioning aims at simultaneously localizing semantic regions and\ndescribing these regions-of-interest (ROIs) with short phrases or sentences in\nnatural language. Previous studies have shown remarkable progresses, but they\nare often vulnerable to the aperture problem that a caption generated by the\nfeatures inside one ROI lacks contextual coherence with its surrounding context\nin the input image. In this work, we investigate contextual reasoning based on\nmulti-scale message propagations from the neighboring contents to the target\nROIs. To this end, we design a novel end-to-end context and attribute grounded\ndense captioning framework consisting of 1) a contextual visual mining module\nand 2) a multi-level attribute grounded description generation module. Knowing\nthat captions often co-occur with the linguistic attributes (such as who, what\nand where), we also incorporate an auxiliary supervision from hierarchical\nlinguistic attributes to augment the distinctiveness of the learned captions.\nExtensive experiments and ablation studies on Visual Genome dataset demonstrate\nthe superiority of the proposed model in comparison to state-of-the-art\nmethods.","url_abs":"http://arxiv.org/abs/1904.01410v1","url_pdf":"http://arxiv.org/pdf/1904.01410v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"dense-captioning","task_name":"Dense Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/dense-captioning-on-visual-genome","task":"Dense Captioning","dataset":"Visual Genome","model":"CAG-Net","rank_in_archive_order":3,"of":4,"metrics":{"mAP":"10.5"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.01410","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}