Papers › Predicting Camera Viewpoint Improves Cross-dataset Generalization for 3D Human Pose Estimation
Predicting Camera Viewpoint Improves Cross-dataset Generalization for 3D Human Pose Estimation
Zhe Wang, Daeyun Shin, Charless C. Fowlkes
Monocular estimation of 3d human pose has attracted increased attention with the availability of large ground-truth motion capture datasets. However, the diversity of training data available is limited and it is not clear to what extent methods generalize outside the specific datasets they are trained on. In this work we carry out a systematic study of the diversity and biases present in specific datasets and its effect on cross-dataset generalization across a compendium of 5 pose datasets. We specifically focus on systematic differences in the distribution of camera viewpoints relative to a body-centered coordinate frame. Based on this observation, we propose an auxiliary task of predicting the camera viewpoint in addition to pose. We find that models trained to jointly predict viewpoint and pose systematically show significantly improved cross-dataset generalization.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| 3D Human Pose Estimation | Geometric Pose Affordance | Cross Dataset Generalization | MPJPE | 53.3 | #2 of 2 | Archive leaderboard | report |
| 3D Human Pose Estimation | Surreal | Cross Dataset Generalization | MPJPE | 37.1 | #2 of 6 | Archive leaderboard | report |
| 3D Human Pose Estimation | Surreal | Cross Dataset Generalization | PCK | 97.3 | #2 of 6 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | cross-dataset-evaluation | Average MPJPE (mm) | 52.0 | #25 of 52 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | cross-dataset-evaluation | Frames Needed | 1 | #25 of 52 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | cross-dataset-evaluation | Need Ground Truth 2D Pose | No | #25 of 52 | Archive leaderboard | report |
| Monocular 3D Human Pose Estimation | Human3.6M | cross-dataset-evaluation | Use Video Sequence | No | #25 of 52 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections