{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/action-machine-rethinking-action-recognition","title":"Action Machine: Rethinking Action Recognition in Trimmed Videos","arxiv_id":"1812.05770","date":"2018-12-14","proceeding":null,"authors":["Jiagang Zhu","Wei Zou","Liang Xu","Yiming Hu","Zheng Zhu","Manyu Chang","Jun-Jie Huang","Guan Huang","Dalong Du"],"abstract":"Existing methods in video action recognition mostly do not distinguish human\nbody from the environment and easily overfit the scenes and objects. In this\nwork, we present a conceptually simple, general and high-performance framework\nfor action recognition in trimmed videos, aiming at person-centric modeling.\nThe method, called Action Machine, takes as inputs the videos cropped by person\nbounding boxes. It extends the Inflated 3D ConvNet (I3D) by adding a branch for\nhuman pose estimation and a 2D CNN for pose-based action recognition, being\nfast to train and test. Action Machine can benefit from the multi-task training\nof action recognition and pose estimation, the fusion of predictions from RGB\nimages and poses. On NTU RGB-D, Action Machine achieves the state-of-the-art\nperformance with top-1 accuracies of 97.2% and 94.3% on cross-view and\ncross-subject respectively. Action Machine also achieves competitive\nperformance on another three smaller action recognition datasets: Northwestern\nUCLA Multiview Action3D, MSR Daily Activity3D and UTD-MHAD. Code will be made\navailable.","url_abs":"http://arxiv.org/abs/1812.05770v2","url_pdf":"http://arxiv.org/pdf/1812.05770v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"multimodal-activity-recognition","task_name":"Multimodal Activity Recognition"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"skeleton-based-action-recognition","task_name":"Skeleton Based Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-ntu-rgbd","task":"Action Recognition","dataset":"NTU RGB+D","model":"Action Machine (RGB only)","rank_in_archive_order":11,"of":28,"metrics":{"Accuracy (CS)":"94.3","Accuracy (CV)":"97.2"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-utd-mhad","task":"Action Recognition","dataset":"UTD-MHAD","model":"Action Machine (RGB only)","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy":"92.5"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-activity-recognition-on-msr-daily-1","task":"Multimodal Activity Recognition","dataset":"MSR Daily Activity3D dataset","model":"Action Machine (RGB only)","rank_in_archive_order":3,"of":6,"metrics":{"Accuracy":"93.0"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-activity-recognition-on-utd-mhad","task":"Multimodal Activity Recognition","dataset":"UTD-MHAD","model":"Action Machine","rank_in_archive_order":4,"of":5,"metrics":{"Accuracy (CS)":"92.5"},"uses_additional_data":false},{"leaderboard":"/sota/skeleton-based-action-recognition-on-n-ucla","task":"Skeleton Based Action Recognition","dataset":"N-UCLA","model":"Action Machine","rank_in_archive_order":22,"of":25,"metrics":{"Accuracy":"92.3%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1812.05770","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}