Papers › Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction

Attention Mechanism Exploits Temporal Contexts: Real-Time 3D Human Pose Reconstruction

1 Jun 2020CVPR 2020 6archive 2025-07-28

Ruixu Liu, Ju Shen, He Wang, Chen Chen, Sen-ching Cheung, Vijayan Asari

We propose a novel attention-based framework for 3D human pose estimation from a monocular video. Despite the general success of end-to-end deep learning paradigms, our approach is based on two key observations: (1) temporal incoherence and jitter are often yielded from a single frame prediction; (2) error rate can be remarkably reduced by increasing the receptive field in a video. Therefore, we design an attentional mechanism to adaptively identify significant frames and tensor outputs from each deep neural net layer, leading to a more optimal estimation. To achieve large temporal receptive fields, multi-scale dilated convolutions are employed to model long-range dependencies among frames. The architecture is straightforward to implement and can be flexibly adopted for real-time applications. Any off-the-shelf 2D pose estimation system, e.g. Mocap libraries, can be easily integrated in an ad-hoc fashion. We both quantitatively and qualitatively evaluate our method on various standard benchmark datasets (e.g. Human3.6M, HumanEva). Our method considerably outperforms all the state-of-the-art algorithms up to 8% error reduction (average mean per joint position error: 34.7) as compared to the best-reported results. Code is available at: (https://github.com/lrxjason/Attention3DHumanPose)

PaperPDFCode

Code

lrxjason/Attention3DHumanPose officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

2D Pose Estimation3D Human Pose EstimationMonocular 3D Human Pose EstimationPose Estimation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Human Pose Estimation Human3.6M Attention3DHumanPose (T=243 CPN) Average MPJPE (mm) 45.1 #45 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M Attention3DHumanPose (T=243 CPN) Multi-View or Monocular Multi-View #45 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M Attention3DHumanPose (T=243 CPN) Using 2D ground-truth joints No #45 of 88 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M Attention3DHumanPose 2D detector CPN #18 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M Attention3DHumanPose Average MPJPE (mm) 45.1 #18 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M Attention3DHumanPose Frames Needed 243 #18 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M Attention3DHumanPose Need Ground Truth 2D Pose No #18 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M Attention3DHumanPose Use Video Sequence Yes #18 of 52 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Tanh Activation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections