{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/action-recognition-with-motion","title":"Action Recognition With Motion Diversification and Dynamic Selection","arxiv_id":null,"date":"2022-07-15","proceeding":"TIP 2022 7","authors":["Peiqin Zhuang","Yu Guo","Zhipeng Yu","Luping Zhou","Lei Bai","Ding Liang","Zhiyong Wang","Yali Wang","Wanli Ouyang"],"abstract":"Motion modeling is crucial in modern action\r\nrecognition methods. As motion dynamics like moving tempos\r\nand action amplitude may vary a lot in different video clips,\r\nit poses great challenge on adaptively covering proper motion\r\ninformation. To address this issue, we introduce a Motion\r\nDiversification and Selection (MoDS) module to generate\r\ndiversified spatio-temporal motion features and then select the\r\nsuitable motion representation dynamically for categorizing the\r\ninput video. To be specific, we first propose a spatio-temporal\r\nmotion generation (StMG) module to construct a bank of\r\ndiversified motion features with varying spatial neighborhood\r\nand time range. Then, a dynamic motion selection (DMS)\r\nmodule is leveraged to choose the most discriminative motion\r\nfeature both spatially and temporally from the feature bank.\r\nAs a result, our proposed method can make full use of\r\nthe diversified spatio-temporal motion information, while\r\nmaintaining computational efficiency at the inference stage.\r\nExtensive experiments on five widely-used benchmarks,\r\ndemonstrate the effectiveness of the method and we achieve\r\nstate-of-the-art performance on Something-Something V1 & V2\r\nthat are of large motion variation","url_abs":"https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=9831068","url_pdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=9831068","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"motion-generation","task_name":"Motion Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-something-1","task":"Action Recognition","dataset":"Something-Something V1","model":"MoDS (8+16frames)","rank_in_archive_order":21,"of":74,"metrics":{"Top 1 Accuracy":"56.6"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"MoDS (8+16frames)","rank_in_archive_order":72,"of":123,"metrics":{"Top-1 Accuracy":"67.1"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}