Papers › Learning Pyramid-structured Long-range Dependencies for 3D Human Pose Estimation

Learning Pyramid-structured Long-range Dependencies for 3D Human Pose Estimation

3 Jun 2025IEEE Transactions on Multimedia 2025 1arXiv:2506.02853archive 2025-07-28

Mingjie Wei, Xuemei Xie, Yutong Zhong, Guangming Shi

Action coordination in human structure is indispensable for the spatial constraints of 2D joints to recover 3D pose. Usually, action coordination is represented as a long-range dependence among body parts. However, there are two main challenges in modeling long-range dependencies. First, joints should not only be constrained by other individual joints but also be modulated by the body parts. Second, existing methods make networks deeper to learn dependencies between non-linked parts. They introduce uncorrelated noise and increase the model size. In this paper, we utilize a pyramid structure to better learn potential long-range dependencies. It can capture the correlation across joints and groups, which complements the context of the human sub-structure. In an effective cross-scale way, it captures the pyramid-structured long-range dependence. Specifically, we propose a novel Pyramid Graph Attention (PGA) module to capture long-range cross-scale dependencies. It concatenates information from various scales into a compact sequence, and then computes the correlation between scales in parallel. Combining PGA with graph convolution modules, we develop a Pyramid Graph Transformer (PGFormer) for 3D human pose estimation, which is a lightweight multi-scale transformer architecture. It encapsulates human sub-structures into self-attention by pooling. Extensive experiments show that our approach achieves lower error and smaller model size than state-of-the-art methods on Human3.6M and MPI-INF-3DHP datasets. The code is available at https://github.com/MingjieWe/PGFormer.

PaperPDFConference PDFCode

Code

MingjieWe/PGFormer officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Human Pose EstimationGraph AttentionPose Estimation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Human Pose Estimation Human3.6M DiffPyramid (CPN) Average MPJPE (mm) 49.2 #56 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M DiffPyramid (CPN) Multi-View or Monocular Monocular #56 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M DiffPyramid (CPN) Using 2D ground-truth joints No #56 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M PGFormer (CPN) Average MPJPE (mm) 49.5 #60 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M PGFormer (CPN) Multi-View or Monocular Monocular #60 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M PGFormer (CPN) Using 2D ground-truth joints No #60 of 88 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAttentionBPEConvolutionDense ConnectionsDropoutGraph TransformerLabel SmoothingLapEigenLaplacian PELayer NormalizationPGASoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections