{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-monocular-3d-human-pose-estimation","title":"Learning Monocular 3D Human Pose Estimation from Multi-view Images","arxiv_id":"1803.04775","date":"2018-03-13","proceeding":"CVPR 2018 6","authors":["Helge Rhodin","Jörg Spörri","Isinsu Katircioglu","Victor Constantin","Frédéric Meyer","Erich Müller","Mathieu Salzmann","Pascal Fua"],"abstract":"Accurate 3D human pose estimation from single images is possible with\nsophisticated deep-net architectures that have been trained on very large\ndatasets. However, this still leaves open the problem of capturing motions for\nwhich no such database exists. Manual annotation is tedious, slow, and\nerror-prone. In this paper, we propose to replace most of the annotations by\nthe use of multiple views, at training time only. Specifically, we train the\nsystem to predict the same pose in all views. Such a consistency constraint is\nnecessary but not sufficient to predict accurate poses. We therefore complement\nit with a supervised loss aiming to predict the correct pose in a small set of\nlabeled images, and with a regularization term that penalizes drift from\ninitial predictions. Furthermore, we propose a method to estimate camera pose\njointly with human pose, which lets us utilize multi-view footage where\ncalibration is difficult, e.g., for pan-tilt or moving handheld cameras. We\ndemonstrate the effectiveness of our approach on established benchmarks, as\nwell as on a new Ski dataset with rotating cameras and expert ski motion, for\nwhich annotations are truly hard to obtain.","url_abs":"http://arxiv.org/abs/1803.04775v2","url_pdf":"http://arxiv.org/pdf/1803.04775v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"},{"task_slug":"monocular-3d-human-pose-estimation","task_name":"Monocular 3D Human Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[{"slug":"ski-pose-ptz-camera","name":"Ski-Pose PTZ-Camera","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.04775","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}