{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/predicting-camera-viewpoint-improves-cross","title":"Predicting Camera Viewpoint Improves Cross-dataset Generalization for 3D Human Pose Estimation","arxiv_id":"2004.03143","date":"2020-04-07","proceeding":null,"authors":["Zhe Wang","Daeyun Shin","Charless C. Fowlkes"],"abstract":"Monocular estimation of 3d human pose has attracted increased attention with the availability of large ground-truth motion capture datasets. However, the diversity of training data available is limited and it is not clear to what extent methods generalize outside the specific datasets they are trained on. In this work we carry out a systematic study of the diversity and biases present in specific datasets and its effect on cross-dataset generalization across a compendium of 5 pose datasets. We specifically focus on systematic differences in the distribution of camera viewpoints relative to a body-centered coordinate frame. Based on this observation, we propose an auxiliary task of predicting the camera viewpoint in addition to pose. We find that models trained to jointly predict viewpoint and pose systematically show significantly improved cross-dataset generalization.","url_abs":"https://arxiv.org/abs/2004.03143v1","url_pdf":"https://arxiv.org/pdf/2004.03143v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"monocular-3d-human-pose-estimation","task_name":"Monocular 3D Human Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-geometric-pose-1","task":"3D Human Pose Estimation","dataset":"Geometric Pose Affordance","model":"Cross Dataset Generalization","rank_in_archive_order":2,"of":2,"metrics":{"MPJPE":"53.3"},"uses_additional_data":false},{"leaderboard":"/sota/3d-human-pose-estimation-on-surreal-1","task":"3D Human Pose Estimation","dataset":"Surreal","model":"Cross Dataset Generalization","rank_in_archive_order":2,"of":6,"metrics":{"MPJPE":"37.1","PCK":"97.3"},"uses_additional_data":true},{"leaderboard":"/sota/monocular-3d-human-pose-estimation-on-human3","task":"Monocular 3D Human Pose Estimation","dataset":"Human3.6M","model":"cross-dataset-evaluation","rank_in_archive_order":25,"of":52,"metrics":{"Average MPJPE (mm)":"52.0","Frames Needed":"1","Need Ground Truth 2D Pose":"No","Use Video Sequence":"No"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2004.03143","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}