{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/actor-centric-relation-network","title":"Actor-Centric Relation Network","arxiv_id":"1807.10982","date":"2018-07-28","proceeding":"ECCV 2018 9","authors":["Chen Sun","Abhinav Shrivastava","Carl Vondrick","Kevin Murphy","Rahul Sukthankar","Cordelia Schmid"],"abstract":"Current state-of-the-art approaches for spatio-temporal action localization\nrely on detections at the frame level and model temporal context with 3D\nConvNets. Here, we go one step further and model spatio-temporal relations to\ncapture the interactions between human actors, relevant objects and scene\nelements essential to differentiate similar human actions. Our approach is\nweakly supervised and mines the relevant elements automatically with an\nactor-centric relational network (ACRN). ACRN computes and accumulates\npair-wise relation information from actor and global scene features, and\ngenerates relation features for action classification. It is implemented as\nneural networks and can be trained jointly with an existing action detection\nsystem. We show that ACRN outperforms alternative approaches which capture\nrelation information, and that the proposed framework improves upon the\nstate-of-the-art performance on JHMDB and AVA. A visualization of the learned\nrelation features confirms that our approach is able to attend to the relevant\nrelations for each action.","url_abs":"http://arxiv.org/abs/1807.10982v1","url_pdf":"http://arxiv.org/pdf/1807.10982v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"actor-centric-relation-network","repo_url":"https://github.com/open-mmlab/mmaction2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":null,"task_name":"Relation"},{"task_slug":"relation-network","task_name":"Relation Network"},{"task_slug":"spatio-temporal-action-localization","task_name":"Spatio-Temporal Action Localization"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-ava-v21","task":"Action Recognition","dataset":"AVA v2.1","model":"ARCN","rank_in_archive_order":15,"of":15,"metrics":{"mAP (Val)":"17.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.10982","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}