Papers › Hierarchical Self-Attention Network for Action Localization in Videos
Hierarchical Self-Attention Network for Action Localization in Videos
Rizard Renanda Adhi Pramono, Yie-Tarng Chen, Wen-Hsien Fang
This paper presents a novel Hierarchical Self-Attention Network (HISAN) to generate spatial-temporal tubes for action localization in videos. The essence of HISAN is to combine the two-stream convolutional neural network (CNN) with hierarchical bidirectional self-attention mechanism, which comprises of two levels of bidirectional self-attention to efficaciously capture both of the long-term temporal dependency information and spatial context information to render more precise action localization. Also, a sequence rescoring (SR) algorithm is employed to resolve the dilemma of inconsistent detection scores incurred by occlusion or background clutter. Moreover, a new fusion scheme is invoked, which integrates not only the appearance and motion information from the two-stream network, but also the motion saliency to mitigate the effect of camera motion. Simulations reveal that the new approach achieves competitive performance as the state-of-the-art works in terms of action localization and recognition accuracy on the widespread UCF101-24 and J-HMDB datasets.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Action Detection | J-HMDB | HISAN (VGG-16) | Frame-mAP 0.5 | 76.72 | #3 of 18 | Archive leaderboard | report |
| Action Detection | J-HMDB | HISAN (VGG-16) | Video-mAP 0.2 | 85.97 | #3 of 18 | Archive leaderboard | report |
| Action Detection | J-HMDB | HISAN (VGG-16) | Video-mAP 0.5 | 84.02 | #3 of 18 | Archive leaderboard | report |
| Action Detection | J-HMDB | HISAN (ResNet-101 + FPN) | Video-mAP 0.2 | 87.59 | #14 of 18 | Archive leaderboard | report |
| Action Detection | J-HMDB | HISAN (ResNet-101 + FPN) | Video-mAP 0.5 | 86.49 | #14 of 18 | Archive leaderboard | report |
| Action Detection | UCF101-24 | HISAN (VGG-16) | Frame-mAP 0.5 | 73.71 | #10 of 19 | Archive leaderboard | report |
| Action Detection | UCF101-24 | HISAN (VGG-16) | Video-mAP 0.2 | 80.42 | #10 of 19 | Archive leaderboard | report |
| Action Detection | UCF101-24 | HISAN (VGG-16) | Video-mAP 0.5 | 49.50 | #10 of 19 | Archive leaderboard | report |
| Action Detection | UCF101-24 | HISAN (ResNet-101 + FPN) | Video-mAP 0.2 | 82.30 | #16 of 19 | Archive leaderboard | report |
| Action Detection | UCF101-24 | HISAN (ResNet-101 + FPN) | Video-mAP 0.5 | 51.47 | #16 of 19 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections