Papers › Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video
Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video
Xiaowei Zhou, Menglong Zhu, Spyridon Leonardos, Kosta Derpanis, Kostas Daniilidis
This paper addresses the challenge of 3D full-body human pose estimation from a monocular image sequence. Here, two cases are considered: (i) the image locations of the human joints are provided and (ii) the image locations of joints are unknown. In the former case, a novel approach is introduced that integrates a sparsity-driven 3D geometric prior and temporal smoothness. In the latter case, the former case is extended by treating the image locations of the joints as latent variables. A deep fully convolutional network is trained to predict the uncertainty maps of the 2D joint locations. The 3D pose estimates are realized via an Expectation-Maximization algorithm over the entire sequence, where it is shown that the 2D joint location uncertainties can be conveniently marginalized out during inference. Empirical evaluation on the Human3.6M dataset shows that the proposed approaches achieve greater 3D pose estimation accuracy over state-of-the-art baselines. Further, the proposed approach outperforms a publicly available 2D pose estimation baseline on the challenging PennAction dataset.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Monocular 3D Human Pose Estimation | Human3.6M | Sparseness Meets Deepness | Average MPJPE (mm) | 113.01 | #40 of 52 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | Sparseness Meets Deepness | Frames Needed | 300 | #40 of 52 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | Sparseness Meets Deepness | Need Ground Truth 2D Pose | No | #40 of 52 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | Sparseness Meets Deepness | Use Video Sequence | Yes | #40 of 52 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections