{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/phrase-localization-and-visual-relationship","title":"Phrase Localization and Visual Relationship Detection with Comprehensive Image-Language Cues","arxiv_id":"1611.06641","date":"2016-11-21","proceeding":"ICCV 2017 10","authors":["Bryan A. Plummer","Arun Mallya","Christopher M. Cervantes","Julia Hockenmaier","Svetlana Lazebnik"],"abstract":"This paper presents a framework for localization or grounding of phrases in\nimages using a large collection of linguistic and visual cues. We model the\nappearance, size, and position of entity bounding boxes, adjectives that\ncontain attribute information, and spatial relationships between pairs of\nentities connected by verbs or prepositions. Special attention is given to\nrelationships between people and clothing or body part mentions, as they are\nuseful for distinguishing individuals. We automatically learn weights for\ncombining these cues and at test time, perform joint inference over all phrases\nin a caption. The resulting system produces state of the art performance on\nphrase localization on the Flickr30k Entities dataset and visual relationship\ndetection on the Stanford VRD dataset.","url_abs":"http://arxiv.org/abs/1611.06641v4","url_pdf":"http://arxiv.org/pdf/1611.06641v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"phrase-localization-and-visual-relationship","repo_url":"https://github.com/BryanPlummer/pl-clc","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":null,"task_name":"Position"},{"task_slug":"relationship-detection","task_name":"Relationship Detection"},{"task_slug":"visual-relationship-detection","task_name":"Visual Relationship Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1611.06641","atlas_url":"https://app.syntology.ai/?focus=1611.06641","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}