Papers › Generalizing Monocular 3D Human Pose Estimation in the Wild

Generalizing Monocular 3D Human Pose Estimation in the Wild

11 Apr 2019arXiv:1904.05512archive 2025-07-28

Luyang Wang, Yan Chen, Zhenhua Guo, Keyuan Qian, Mude Lin, Hongsheng Li, Jimmy S. Ren

The availability of the large-scale labeled 3D poses in the Human3.6M dataset plays an important role in advancing the algorithms for 3D human pose estimation from a still image. We observe that recent innovation in this area mainly focuses on new techniques that explicitly address the generalization issue when using this dataset, because this database is constructed in a highly controlled environment with limited human subjects and background variations. Despite such efforts, we can show that the results of the current methods are still error-prone especially when tested against the images taken in-the-wild. In this paper, we aim to tackle this problem from a different perspective. We propose a principled approach to generate high quality 3D pose ground truth given any in-the-wild image with a person inside. We achieve this by first devising a novel stereo inspired neural network to directly map any 2D pose to high quality 3D counterpart. We then perform a carefully designed geometric searching scheme to further refine the joints. Based on this scheme, we build a large-scale dataset with 400,000 in-the-wild images and their corresponding 3D pose ground truth. This enables the training of a high quality neural network model, without specialized training scheme and auxiliary loss function, which performs favorably against the state-of-the-art 3D pose estimation methods. We also evaluate the generalization ability of our model both quantitatively and qualitatively. Results show that our approach convincingly outperforms the previous methods. We make our dataset and code publicly available.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

llcshappy/Monocular-3D-Human-Pose officialmentioned in papermentioned on GitHubtf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Human Pose Estimation3D Pose EstimationMonocular 3D Human Pose EstimationPose Estimation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Human Pose Estimation Human3.6M Stereoscopic View Synthesis Subnetwork Average MPJPE (mm) 58 #80 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M Stereoscopic View Synthesis Subnetwork Multi-View or Monocular Multi-View #80 of 88 Archive leaderboard report
3D Human Pose Estimation Human3.6M Stereoscopic View Synthesis Subnetwork Using 2D ground-truth joints No #80 of 88 Archive leaderboard report
3D Human Pose Estimation MPI-INF-3DHP Stereoscopic View Synthesis Subnetwork AUC 33.8 #101 of 108 Archive leaderboard report
3D Human Pose Estimation MPI-INF-3DHP Stereoscopic View Synthesis Subnetwork PCK 71.2 #101 of 108 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections