Papers › Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and...

Disentangled Diffusion-Based 3D Human Pose Estimation with Hierarchical Spatial and Temporal Denoiser

7 Mar 2024arXiv:2403.04444archive 2025-07-28

Qingyuan Cai, Xuecai Hu, Saihui Hou, Li Yao, Yongzhen Huang

Recently, diffusion-based methods for monocular 3D human pose estimation have achieved state-of-the-art (SOTA) performance by directly regressing the 3D joint coordinates from the 2D pose sequence. Although some methods decompose the task into bone length and bone direction prediction based on the human anatomical skeleton to explicitly incorporate more human body prior constraints, the performance of these methods is significantly lower than that of the SOTA diffusion-based methods. This can be attributed to the tree structure of the human skeleton. Direct application of the disentangled method could amplify the accumulation of hierarchical errors, propagating through each hierarchy. Meanwhile, the hierarchical information has not been fully explored by the previous methods. To address these problems, a Disentangled Diffusion-based 3D Human Pose Estimation method with Hierarchical Spatial and Temporal Denoiser is proposed, termed DDHPose. In our approach: (1) We disentangle the 3D pose and diffuse the bone length and bone direction during the forward process of the diffusion model to effectively model the human pose prior. A disentanglement loss is proposed to supervise diffusion model learning. (2) For the reverse process, we propose Hierarchical Spatial and Temporal Denoiser (HSTDenoiser) to improve the hierarchical modeling of each joint. Our HSTDenoiser comprises two components: the Hierarchical-Related Spatial Transformer (HRST) and the Hierarchical-Related Temporal Transformer (HRTT). HRST exploits joint spatial information and the influence of the parent joint on each joint for spatial modeling, while HRTT utilizes information from both the joint and its hierarchical adjacent joints to explore the hierarchical temporal correlations among joints.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Andyen512/DDHPose officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Human Pose EstimationDisentanglementMonocular 3D Human Pose EstimationMulti-Hypotheses 3D Human Pose EstimationPose Estimation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Monocular 3D Human Pose Estimation Human3.6M DDHPose 2D detector CPN #6 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M DDHPose Average MPJPE (mm) 39.7 #6 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M DDHPose Frames Needed 243 #6 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M DDHPose Need Ground Truth 2D Pose No #6 of 52 Archive leaderboard report
Monocular 3D Human Pose Estimation Human3.6M DDHPose Use Video Sequence Yes #6 of 52 Archive leaderboard report
Multi-Hypotheses 3D Human Pose Estimation Human3.6M DDHPose (H=20, W=10, J-Best) Average MPJPE (mm) 33.62 #1 of 12 Archive leaderboard report
Multi-Hypotheses 3D Human Pose Estimation Human3.6M DDHPose (H=20, W=10, J-Best) Average PMPJPE (mm) 26.48 #1 of 12 Archive leaderboard report
Multi-Hypotheses 3D Human Pose Estimation Human3.6M DDHPose (H=20, W=10, P-Best) Average MPJPE (mm) 39.0 #5 of 12 Archive leaderboard report
Multi-Hypotheses 3D Human Pose Estimation Human3.6M DDHPose (H=20, W=10, P-Best) Average PMPJPE (mm) 31.2 #5 of 12 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDiffusionDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxSpatial TransformerTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections