Papers › Temporal-Aware Refinement for Video-based Human Pose and Shape Recovery
Temporal-Aware Refinement for Video-based Human Pose and Shape Recovery
Ming Chen, Yan Zhou, Weihua Jian, Pengfei Wan, Zhongyuan Wang
Though significant progress in human pose and shape recovery from monocular RGB images has been made in recent years, obtaining 3D human motion with high accuracy and temporal consistency from videos remains challenging. Existing video-based methods tend to reconstruct human motion from global image features, which lack detailed representation capability and limit the reconstruction accuracy. In this paper, we propose a Temporal-Aware Refining Network (TAR), to synchronously explore temporal-aware global and local image features for accurate pose and shape recovery. First, a global transformer encoder is introduced to obtain temporal global features from static feature sequences. Second, a bidirectional ConvGRU network takes the sequence of high-resolution feature maps as input, and outputs temporal local feature maps that maintain high resolution and capture the local motion of the human body. Finally, a recurrent refinement module iteratively updates estimated SMPL parameters by leveraging both global and local temporal information to achieve accurate and smooth results. Extensive experiments demonstrate that our TAR obtains more accurate results than previous state-of-the-art methods on popular benchmarks, i.e., 3DPW, MPI-INF-3DHP, and Human3.6M.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| 3D Human Pose Estimation | 3DPW | TAR (N=9) | Acceleration Error | 7.7 | #3 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | 3DPW | TAR (N=9) | MPJPE | 62.7 | #3 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | 3DPW | TAR (N=9) | MPVPE | 74.4 | #3 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | 3DPW | TAR (N=9) | PA-MPJPE | 40.6 | #3 of 119 | Archive leaderboard | report |
| 3D Human Pose Estimation | MPI-INF-3DHP | TAR (N=9) | Acceleration Error | 9.2 | #43 of 108 | Archive leaderboard | report |
| 3D Human Pose Estimation | MPI-INF-3DHP | TAR (N=9) | MPJPE | 85.9 | #43 of 108 | Archive leaderboard | report |
| 3D Human Pose Estimation | MPI-INF-3DHP | TAR (N=9) | PA-MPJPE | 60.5 | #43 of 108 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections