{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-scale-motion-aware-module-for-video","title":"Multi-scale Motion-Aware Module for Video Action Recognition","arxiv_id":null,"date":"2023-02-19","proceeding":"ECCV Workshops 2023 2","authors":["Huai-Wei Peng","Yu-Chee Tseng"],"abstract":"Due to the lengthy computing time for optical flow, recent\r\nworks have proposed to use the correlation operation as an alternative approach to extracting motion features. Although using correlation operations shows significant improvement with negligible FLOPs,\r\nit introduces much more latency per FLOP than convolution operations and increases noticeable latency as a larger searching patch is\r\napplied. Nonetheless, shrinking the searching patch in correlation operation is doomed to degrade its performance owing to the inability to\r\ncapture larger displacements. In this paper, we propose an effective and\r\nlow-latency Multi-Scale Motion-Aware (MSMA) module. It uses smaller\r\nsearching patches at different scales for efficiently extracting motion features from large displacements. It can be installed into and generalizes\r\nwell on different CNN backbones. When installed into TSM ResNet-50,\r\nthe MSMA module introduces ≈ 17.6% more latency on NVIDIA Tesla\r\nV100 GPU, yet, it achieves state-of-the-art performance on SomethingSomething V1 & V2 and Diving-48.","url_abs":"https://link.springer.com/chapter/10.1007/978-3-031-25075-0_40","url_pdf":"https://link.springer.com/chapter/10.1007/978-3-031-25075-0_40","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":null,"task_name":"GPU"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-something-1","task":"Action Recognition","dataset":"Something-Something V1","model":"MSMA (8+16frames)","rank_in_archive_order":12,"of":74,"metrics":{"Top 1 Accuracy":"57.9"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-something","task":"Action Recognition","dataset":"Something-Something V2","model":"MSMA (8+16frames)","rank_in_archive_order":53,"of":123,"metrics":{"Top-1 Accuracy":"68.2"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}