{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/contextual-action-recognition-with-rcnn","title":"Contextual Action Recognition with R*CNN","arxiv_id":"1505.01197","date":"2015-05-05","proceeding":"ICCV 2015 12","authors":["Georgia Gkioxari","Ross Girshick","Jitendra Malik"],"abstract":"There are multiple cues in an image which reveal what action a person is\nperforming. For example, a jogger has a pose that is characteristic for\njogging, but the scene (e.g. road, trail) and the presence of other joggers can\nbe an additional source of information. In this work, we exploit the simple\nobservation that actions are accompanied by contextual cues to build a strong\naction recognition system. We adapt RCNN to use more than one region for\nclassification while still maintaining the ability to localize the action. We\ncall our system R*CNN. The action-specific models and the feature maps are\ntrained jointly, allowing for action specific representations to emerge. R*CNN\nachieves 90.2% mean AP on the PASAL VOC Action dataset, outperforming all other\napproaches in the field by a significant margin. Last, we show that R*CNN is\nnot limited to action recognition. In particular, R*CNN can also be used to\ntackle fine-grained tasks such as attribute classification. We validate this\nclaim by reporting state-of-the-art performance on the Berkeley Attributes of\nPeople dataset.","url_abs":"http://arxiv.org/abs/1505.01197v3","url_pdf":"http://arxiv.org/pdf/1505.01197v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"contextual-action-recognition-with-rcnn","repo_url":"https://github.com/gkioxari/RstarCNN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"caffe2","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"contextual-action-recognition-with-rcnn","repo_url":"https://github.com/VivekUdupa/CPSC8810-DeepLearning_vkoodli_shonnah","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"human-object-interaction-detection","task_name":"Human-Object Interaction Detection"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-object-interaction-detection-on-hico-1","task":"Human-Object Interaction Detection","dataset":"HICO","model":"R*CNN","rank_in_archive_order":8,"of":8,"metrics":{"mAP":"28.5"},"uses_additional_data":false},{"leaderboard":"/sota/weakly-supervised-object-detection-on-4","task":"Weakly Supervised Object Detection","dataset":"Charades","model":"R*CNN","rank_in_archive_order":5,"of":6,"metrics":{"MAP":"0.99"},"uses_additional_data":false},{"leaderboard":"/sota/weakly-supervised-object-detection-on-hico","task":"Weakly Supervised Object Detection","dataset":"HICO-DET","model":"R*CNN","rank_in_archive_order":4,"of":4,"metrics":{"MAP":"2.15"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1505.01197","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}