{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/detecting-and-recognizing-human-object","title":"Detecting and Recognizing Human-Object Interactions","arxiv_id":"1704.07333","date":"2017-04-24","proceeding":"CVPR 2018 6","authors":["Georgia Gkioxari","Ross Girshick","Piotr Dollár","Kaiming He"],"abstract":"To understand the visual world, a machine must not only recognize individual\nobject instances but also how they interact. Humans are often at the center of\nsuch interactions and detecting human-object interactions is an important\npractical and scientific problem. In this paper, we address the task of\ndetecting <human, verb, object> triplets in challenging everyday photos. We\npropose a novel model that is driven by a human-centric approach. Our\nhypothesis is that the appearance of a person -- their pose, clothing, action\n-- is a powerful cue for localizing the objects they are interacting with. To\nexploit this cue, our model learns to predict an action-specific density over\ntarget object locations based on the appearance of a detected person. Our model\nalso jointly learns to detect people and objects, and by fusing these\npredictions it efficiently infers interaction triplets in a clean, jointly\ntrained end-to-end system we call InteractNet. We validate our approach on the\nrecently introduced Verbs in COCO (V-COCO) and HICO-DET datasets, where we show\nquantitatively compelling results.","url_abs":"http://arxiv.org/abs/1704.07333v3","url_pdf":"http://arxiv.org/pdf/1704.07333v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"detecting-and-recognizing-human-object","repo_url":"https://github.com/facebookresearch/detectron","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"detecting-and-recognizing-human-object","repo_url":"https://github.com/jiajunhua/facebookresearch-Detectron","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"caffe2","reach":null}],"tasks":[{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"},{"task_slug":"object","task_name":"Object"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-object-interaction-detection-on-hico","task":"Human-Object Interaction Detection","dataset":"HICO-DET","model":"InteractNet","rank_in_archive_order":55,"of":55,"metrics":{"Time Per Frame (ms)":"145","mAP":"9.94"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.07333","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}