{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-human-object-interactions-by-graph","title":"Learning Human-Object Interactions by Graph Parsing Neural Networks","arxiv_id":"1808.07962","date":"2018-08-23","proceeding":"ECCV 2018 9","authors":["Siyuan Qi","Wenguan Wang","Baoxiong Jia","Jianbing Shen","Song-Chun Zhu"],"abstract":"This paper addresses the task of detecting and recognizing human-object\ninteractions (HOI) in images and videos. We introduce the Graph Parsing Neural\nNetwork (GPNN), a framework that incorporates structural knowledge while being\ndifferentiable end-to-end. For a given scene, GPNN infers a parse graph that\nincludes i) the HOI graph structure represented by an adjacency matrix, and ii)\nthe node labels. Within a message passing inference framework, GPNN iteratively\ncomputes the adjacency matrices and node labels. We extensively evaluate our\nmodel on three HOI detection benchmarks on images and videos: HICO-DET, V-COCO,\nand CAD-120 datasets. Our approach significantly outperforms state-of-art\nmethods, verifying that GPNN is scalable to large datasets and applies to\nspatial-temporal settings. The code is available at\nhttps://github.com/SiyuanQi/gpnn.","url_abs":"http://arxiv.org/abs/1808.07962v1","url_pdf":"http://arxiv.org/pdf/1808.07962v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-human-object-interactions-by-graph","repo_url":"https://github.com/SiyuanQi/gpnn","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"},{"task_slug":"object","task_name":"Object"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-object-interaction-detection-on-hico","task":"Human-Object Interaction Detection","dataset":"HICO-DET","model":"GPNN","rank_in_archive_order":54,"of":55,"metrics":{"mAP":"13.11"},"uses_additional_data":false},{"leaderboard":"/sota/human-object-interaction-detection-on-v-coco","task":"Human-Object Interaction Detection","dataset":"V-COCO","model":"GPNN","rank_in_archive_order":32,"of":34,"metrics":{"AP(S1)":"44.0"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1808.07962","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}