{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/temporal-spatial-mapping-for-action","title":"Temporal-Spatial Mapping for Action Recognition","arxiv_id":"1809.03669","date":"2018-09-11","proceeding":null,"authors":["Xiaolin Song","Cuiling Lan","Wen-Jun Zeng","Junliang Xing","Jingyu Yang","Xiaoyan Sun"],"abstract":"Deep learning models have enjoyed great success for image related computer\nvision tasks like image classification and object detection. For video related\ntasks like human action recognition, however, the advancements are not as\nsignificant yet. The main challenge is the lack of effective and efficient\nmodels in modeling the rich temporal spatial information in a video. We\nintroduce a simple yet effective operation, termed Temporal-Spatial Mapping\n(TSM), for capturing the temporal evolution of the frames by jointly analyzing\nall the frames of a video. We propose a video level 2D feature representation\nby transforming the convolutional features of all frames to a 2D feature map,\nreferred to as VideoMap. With each row being the vectorized feature\nrepresentation of a frame, the temporal-spatial features are compactly\nrepresented, while the temporal dynamic evolution is also well embedded. Based\non the VideoMap representation, we further propose a temporal attention model\nwithin a shallow convolutional neural network to efficiently exploit the\ntemporal-spatial dynamics. The experiment results show that the proposed scheme\nachieves the state-of-the-art performance, with 4.2% accuracy gain over\nTemporal Segment Network (TSN), a competing baseline method, on the challenging\nhuman action benchmark dataset HMDB51.","url_abs":"http://arxiv.org/abs/1809.03669v1","url_pdf":"http://arxiv.org/pdf/1809.03669v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"image-classification","task_name":"image-classification"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-ucf101","task":"Action Recognition","dataset":"UCF101","model":"TSN+TSM","rank_in_archive_order":54,"of":91,"metrics":{"3-fold Accuracy":"94.3"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}