{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dense-captioning-with-joint-inference-and","title":"Dense Captioning with Joint Inference and Visual Context","arxiv_id":"1611.06949","date":"2016-11-21","proceeding":"CVPR 2017 7","authors":["Linjie Yang","Kevin Tang","Jianchao Yang","Li-Jia Li"],"abstract":"Dense captioning is a newly emerging computer vision topic for understanding\nimages with dense language descriptions. The goal is to densely detect visual\nconcepts (e.g., objects, object parts, and interactions between them) from\nimages, labeling each with a short descriptive phrase. We identify two key\nchallenges of dense captioning that need to be properly addressed when tackling\nthe problem. First, dense visual concept annotations in each image are\nassociated with highly overlapping target regions, making accurate localization\nof each visual concept challenging. Second, the large amount of visual concepts\nmakes it hard to recognize each of them by appearance alone. We propose a new\nmodel pipeline based on two novel ideas, joint inference and context fusion, to\nalleviate these two challenges. We design our model architecture in a\nmethodical manner and thoroughly evaluate the variations in architecture. Our\nfinal model, compact and efficient, achieves state-of-the-art accuracy on\nVisual Genome for dense captioning with a relative gain of 73\\% compared to the\nprevious best algorithm. Qualitative experiments also reveal the semantic\ncapabilities of our model in dense captioning.","url_abs":"http://arxiv.org/abs/1611.06949v2","url_pdf":"http://arxiv.org/pdf/1611.06949v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dense-captioning-with-joint-inference-and","repo_url":"https://github.com/linjieyangsc/densecap","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"dense-captioning","task_name":"Dense Captioning"},{"task_slug":"descriptive","task_name":"Descriptive"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1611.06949","atlas_url":"https://app.syntology.ai/?focus=1611.06949","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}