{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-relationship-detection-with-language","title":"Visual Relationship Detection with Language Priors","arxiv_id":"1608.00187","date":"2016-07-31","proceeding":null,"authors":["Cewu Lu","Ranjay Krishna","Michael Bernstein","Li Fei-Fei"],"abstract":"Visual relationships capture a wide variety of interactions between pairs of\nobjects in images (e.g. \"man riding bicycle\" and \"man pushing bicycle\").\nConsequently, the set of possible relationships is extremely large and it is\ndifficult to obtain sufficient training examples for all possible\nrelationships. Because of this limitation, previous work on visual relationship\ndetection has concentrated on predicting only a handful of relationships.\nThough most relationships are infrequent, their objects (e.g. \"man\" and\n\"bicycle\") and predicates (e.g. \"riding\" and \"pushing\") independently occur\nmore frequently. We propose a model that uses this insight to train visual\nmodels for objects and predicates individually and later combines them together\nto predict multiple relationships per image. We improve on prior work by\nleveraging language priors from semantic word embeddings to finetune the\nlikelihood of a predicted relationship. Our model can scale to predict\nthousands of types of relationships from a few examples. Additionally, we\nlocalize the objects in the predicted relationships as bounding boxes in the\nimage. We further demonstrate that understanding relationships can improve\ncontent based image retrieval.","url_abs":"http://arxiv.org/abs/1608.00187v1","url_pdf":"http://arxiv.org/pdf/1608.00187v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"content-based-image-retrieval","task_name":"Content-Based Image Retrieval"},{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"relationship-detection","task_name":"Relationship Detection"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"visual-relationship-detection","task_name":"Visual Relationship Detection"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[{"slug":"vrd","name":"VRD","full_name":"Visual Relationship Detection dataset"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/scene-graph-generation-on-vrd","task":"Scene Graph Generation","dataset":"VRD","model":"VRD","rank_in_archive_order":2,"of":2,"metrics":{"Recall@50":"18.16"},"uses_additional_data":false},{"leaderboard":"/sota/visual-relationship-detection-on-vrd-phrase","task":"Visual Relationship Detection","dataset":"VRD Phrase Detection","model":"Lu et. al [[Lu et al.2016]]","rank_in_archive_order":7,"of":7,"metrics":{"R@100":"17.03","R@50":"16.17"},"uses_additional_data":false},{"leaderboard":"/sota/visual-relationship-detection-on-vrd","task":"Visual Relationship Detection","dataset":"VRD Predicate Detection","model":"Lu et. al [[Lu et al.2016]]","rank_in_archive_order":6,"of":7,"metrics":{"R@100":"47.87","R@50":"47.87"},"uses_additional_data":false},{"leaderboard":"/sota/visual-relationship-detection-on-vrd-1","task":"Visual Relationship Detection","dataset":"VRD Relationship Detection","model":"Lu et. al [[Lu et al.2016]]","rank_in_archive_order":8,"of":8,"metrics":{"R@100":"14.70","R@50":"13.86"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1608.00187","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}