Papers › Recognizing Human Actions as the Evolution of Pose Estimation Maps
Recognizing Human Actions as the Evolution of Pose Estimation Maps
Mengyuan Liu, Junsong Yuan
Most video-based action recognition approaches choose to extract features from the whole video to recognize actions. The cluttered background and non-action motions limit the performances of these methods, since they lack the explicit modeling of human body movements. With recent advances of human pose estimation, this work presents a novel method to recognize human action as the evolution of pose estimation maps. Instead of relying on the inaccurate human poses estimated from videos, we observe that pose estimation maps, the byproduct of pose estimation, preserve richer cues of human body to benefit action recognition. Specifically, the evolution of pose estimation maps can be decomposed as an evolution of heatmaps, e.g., probabilistic maps, and an evolution of estimated 2D human poses, which denote the changes of body shape and body pose, respectively. Considering the sparse property of heatmap, we develop spatial rank pooling to aggregate the evolution of heatmaps as a body shape evolution image. As body shape evolution image does not differentiate body parts, we design body guided sampling to aggregate the evolution of poses as a body pose evolution image. The complementary properties between both types of images are explored by deep convolutional neural networks to predict action label. Experiments on NTU RGB+D, UTD-MHAD and PennAction datasets verify the effectiveness of our method, which outperforms most state-of-the-art methods.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Action Recognition | NTU RGB+D | PoseMap (RGB+Pose) | Accuracy (CS) | 91.7 | #23 of 28 | Archive leaderboard | report |
| Action Recognition | NTU RGB+D | PoseMap (RGB+Pose) | Accuracy (CV) | 95.2 | #23 of 28 | Archive leaderboard | report |
| Action Recognition | NTU RGB+D 120 | Body Pose Evolution Map | Accuracy (Cross-Setup) | 64.6 | #21 of 21 | Archive leaderboard | report |
| Action Recognition | NTU RGB+D 120 | Body Pose Evolution Map | Accuracy (Cross-Subject) | 66.9 | #21 of 21 | Archive leaderboard | report |
| Multimodal Activity Recognition | UTD-MHAD | PoseMap | Accuracy (CS) | 94.5 | #1 of 5 | Archive leaderboard | report |
| Skeleton Based Action Recognition | NTU RGB+D 120 | Body Pose Evolution Map | Accuracy (Cross-Setup) | 66.9% | #71 of 83 | Archive leaderboard | report |
| Skeleton Based Action Recognition | NTU RGB+D 120 | Body Pose Evolution Map | Accuracy (Cross-Subject) | 64.6% | #71 of 83 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections