Papers › Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction
Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction
Ruixu Liu, Ju Shen, He Wang, Chen Chen, Sen-ching Cheung, Vijayan Asari
We propose a novel attention-based framework for 3D human pose estimation from a monocular video. Despite the general success of end-to-end deep learning paradigms, our approach is based on two key observations: (1) temporal incoherence and jitter are often yielded from a single frame prediction; (2) error rate can be remarkably reduced by increasing the receptive field in a video. Therefore, we design an attentional mechanism to adaptively identify significant frames and tensor outputs from each deep neural net layer, leading to a more optimal estimation. To achieve large temporal receptive fields, multi-scale dilated convolutions are employed to model long-range dependencies among frames. The architecture is straightforward to implement and can be flexibly adopted for real-time applications. Any off-the-shelf 2D pose estimation system, e.g. Mocap libraries, can be easily integrated in an ad-hoc fashion. We both quantitatively and qualitatively evaluate our method on various standard benchmark datasets (e.g. Human3.6M, HumanEva). Our method considerably outperforms all the state-of-the-art algorithms up to 8% error reduction (average mean per joint position error: 34.7) as compared to the best-reported results. Code is available at: (https://github.com/lrxjason/Attention3DHumanPose)
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| 3D Human Pose Estimation | Human3.6M | Attention3DHumanPose (T=243 CPN) | Average MPJPE (mm) | 45.1 | #45 of 88 | Archive leaderboard | report |
| 3D Human Pose Estimation | Human3.6M | Attention3DHumanPose (T=243 CPN) | Multi-View or Monocular | Multi-View | #45 of 88 | Archive leaderboard | report |
| 3D Human Pose Estimation | Human3.6M | Attention3DHumanPose (T=243 CPN) | Using 2D ground-truth joints | No | #45 of 88 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | Attention3DHumanPose | 2D detector | CPN | #18 of 52 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | Attention3DHumanPose | Average MPJPE (mm) | 45.1 | #18 of 52 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | Attention3DHumanPose | Frames Needed | 243 | #18 of 52 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | Attention3DHumanPose | Need Ground Truth 2D Pose | No | #18 of 52 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | Attention3DHumanPose | Use Video Sequence | Yes | #18 of 52 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections