{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-variation-structured-reinforcement","title":"Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection","arxiv_id":"1703.03054","date":"2017-03-08","proceeding":"CVPR 2017 7","authors":["Xiaodan Liang","Lisa Lee","Eric P. Xing"],"abstract":"Despite progress in visual perception tasks such as image classification and\ndetection, computers still struggle to understand the interdependency of\nobjects in the scene as a whole, e.g., relations between objects or their\nattributes. Existing methods often ignore global context cues capturing the\ninteractions among different object instances, and can only recognize a handful\nof types by exhaustively training individual detectors for all possible\nrelationships. To capture such global interdependency, we propose a deep\nVariation-structured Reinforcement Learning (VRL) framework to sequentially\ndiscover object relationships and attributes in the whole image. First, a\ndirected semantic action graph is built using language priors to provide a rich\nand compact representation of semantic correlations between object categories,\npredicates, and attributes. Next, we use a variation-structured traversal over\nthe action graph to construct a small, adaptive action set for each step based\non the current state and historical actions. In particular, an ambiguity-aware\nobject mining scheme is used to resolve semantic ambiguity among object\ncategories that the object detector fails to distinguish. We then make\nsequential predictions using a deep RL framework, incorporating global context\ncues and semantic embeddings of previously extracted phrases in the state\nvector. Our experiments on the Visual Relationship Detection (VRD) dataset and\nthe large-scale Visual Genome dataset validate the superiority of VRL, which\ncan achieve significantly better detection results on datasets involving\nthousands of relationship and attribute types. We also demonstrate that VRL is\nable to predict unseen types embedded in our action graph by learning\ncorrelations on shared graph nodes.","url_abs":"http://arxiv.org/abs/1703.03054v1","url_pdf":"http://arxiv.org/pdf/1703.03054v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-variation-structured-reinforcement","repo_url":"https://github.com/nexusapoorvacus/DeepVariationStructuredRL","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"object","task_name":"Object"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"relationship-detection","task_name":"Relationship Detection"},{"task_slug":"visual-relationship-detection","task_name":"Visual Relationship Detection"},{"task_slug":"image-classification","task_name":"image-classification"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-relationship-detection-on-vrd-phrase","task":"Visual Relationship Detection","dataset":"VRD Phrase Detection","model":"Liang et. al [[Liang, Lee, and Xing2017]]","rank_in_archive_order":4,"of":7,"metrics":{"R@100":"22.60","R@50":"21.37"},"uses_additional_data":false},{"leaderboard":"/sota/visual-relationship-detection-on-vrd-1","task":"Visual Relationship Detection","dataset":"VRD Relationship Detection","model":"Liang et. al [[Liang, Lee, and Xing2017]]","rank_in_archive_order":5,"of":8,"metrics":{"R@100":"20.79","R@50":"18.19"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1703.03054","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}