Papers › Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from...
Capturing Humans in Motion: Temporal-Attentive 3D Human Pose and Shape Estimation from Monocular Video
Wen-Li Wei, Jen-Chun Lin, Tyng-Luh Liu, Hong-Yuan Mark Liao
Learning to capture human motion is essential to 3D human pose and shape estimation from monocular video. However, the existing methods mainly rely on recurrent or convolutional operation to model such temporal information, which limits the ability to capture non-local context relations of human motion. To address this problem, we propose a motion pose and shape network (MPS-Net) to effectively capture humans in motion to estimate accurate and temporally coherent 3D human pose and shape from a video. Specifically, we first propose a motion continuity attention (MoCA) module that leverages visual cues observed from human motion to adaptively recalibrate the range that needs attention in the sequence to better capture the motion continuity dependencies. Then, we develop a hierarchical attentive feature integration (HAFI) module to effectively combine adjacent past and future feature representations to strengthen temporal correlation and refine the feature representation of the current frame. By coupling the MoCA and HAFI modules, the proposed MPS-Net excels in estimating 3D human pose and shape in the video. Though conceptually simple, our MPS-Net not only outperforms the state-of-the-art methods on the 3DPW, MPI-INF-3DHP, and Human3.6M benchmark datasets, but also uses fewer network parameters. The video demos can be found at https://mps-net.github.io/MPS-Net/.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| 3D Human Pose Estimation | 3DPW | MPS-Net (T=16) | Acceleration Error | 7.4 | #47 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | 3DPW | MPS-Net (T=16) | FLOPs (G) | 4.45 | #47 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | 3DPW | MPS-Net (T=16) | MPJPE | 84.3 | #47 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | 3DPW | MPS-Net (T=16) | MPVPE | 99.7 | #47 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | 3DPW | MPS-Net (T=16) | Number of parameters (M) | 39.63 | #47 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | 3DPW | MPS-Net (T=16) | PA-MPJPE | 52.1 | #47 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | MPI-INF-3DHP | MPS-Net (T=16) | Acceleration Error | 9.6 | #59 of 108 | Archive leaderboard | report |
| 3D Human Pose Estimation | MPI-INF-3DHP | MPS-Net (T=16) | MPJPE | 96.7 | #59 of 108 | Archive leaderboard | report |
| 3D Human Pose Estimation | MPI-INF-3DHP | MPS-Net (T=16) | PA-MPJPE | 62.8 | #59 of 108 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections