{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hierarchical-self-attention-network-for","title":"Hierarchical Self-Attention Network for Action Localization in Videos","arxiv_id":null,"date":"2019-10-01","proceeding":"ICCV 2019 10","authors":["Rizard Renanda Adhi Pramono"," Yie-Tarng Chen"," Wen-Hsien Fang"],"abstract":"This paper presents a novel Hierarchical Self-Attention Network (HISAN) to generate spatial-temporal tubes for action localization in videos. The essence of HISAN is to combine the two-stream convolutional neural network (CNN) with hierarchical bidirectional self-attention mechanism, which comprises of two levels of bidirectional self-attention to efficaciously capture both of the long-term temporal dependency information and spatial context information to render more precise action localization. Also, a sequence rescoring (SR) algorithm is employed to resolve the dilemma of inconsistent detection scores incurred by occlusion or background clutter. Moreover, a new fusion scheme is invoked, which integrates not only the appearance and motion information from the two-stream network, but also the motion saliency to mitigate the effect of camera motion. Simulations reveal that the new approach achieves competitive performance as the state-of-the-art works in terms of action localization and recognition accuracy on the widespread UCF101-24 and J-HMDB datasets.\r","url_abs":"http://openaccess.thecvf.com/content_ICCV_2019/html/Pramono_Hierarchical_Self-Attention_Network_for_Action_Localization_in_Videos_ICCV_2019_paper.html","url_pdf":"http://openaccess.thecvf.com/content_ICCV_2019/papers/Pramono_Hierarchical_Self-Attention_Network_for_Action_Localization_in_Videos_ICCV_2019_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-localization","task_name":"Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-detection-on-j-hmdb","task":"Action Detection","dataset":"J-HMDB","model":"HISAN (VGG-16)","rank_in_archive_order":3,"of":18,"metrics":{"Frame-mAP 0.5":"76.72","Video-mAP 0.2":"85.97","Video-mAP 0.5":"84.02"},"uses_additional_data":false},{"leaderboard":"/sota/action-detection-on-j-hmdb","task":"Action Detection","dataset":"J-HMDB","model":"HISAN (ResNet-101 + FPN)","rank_in_archive_order":14,"of":18,"metrics":{"Video-mAP 0.2":"87.59","Video-mAP 0.5":"86.49"},"uses_additional_data":false},{"leaderboard":"/sota/action-detection-on-ucf101-24","task":"Action Detection","dataset":"UCF101-24","model":"HISAN (VGG-16)","rank_in_archive_order":10,"of":19,"metrics":{"Frame-mAP 0.5":"73.71","Video-mAP 0.2":"80.42","Video-mAP 0.5":"49.50"},"uses_additional_data":false},{"leaderboard":"/sota/action-detection-on-ucf101-24","task":"Action Detection","dataset":"UCF101-24","model":"HISAN (ResNet-101 + FPN)","rank_in_archive_order":16,"of":19,"metrics":{"Video-mAP 0.2":"82.30","Video-mAP 0.5":"51.47"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}