{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-multitask-architecture-for-integrated-2d","title":"Deep Multitask Architecture for Integrated 2D and 3D Human Sensing","arxiv_id":"1701.08985","date":"2017-01-31","proceeding":"CVPR 2017 7","authors":["Alin-Ionut Popa","Mihai Zanfir","Cristian Sminchisescu"],"abstract":"We propose a deep multitask architecture for \\emph{fully automatic 2d and 3d\nhuman sensing} (DMHS), including \\emph{recognition and reconstruction}, in\n\\emph{monocular images}. The system computes the figure-ground segmentation,\nsemantically identifies the human body parts at pixel level, and estimates the\n2d and 3d pose of the person. The model supports the joint training of all\ncomponents by means of multi-task losses where early processing stages\nrecursively feed into advanced ones for increasingly complex calculations,\naccuracy and robustness. The design allows us to tie a complete training\nprotocol, by taking advantage of multiple datasets that would otherwise\nrestrictively cover only some of the model components: complex 2d image data\nwith no body part labeling and without associated 3d ground truth, or complex\n3d data with limited 2d background variability. In detailed experiments based\non several challenging 2d and 3d datasets (LSP, HumanEva, Human3.6M), we\nevaluate the sub-structures of the model, the effect of various types of\ntraining data in the multitask loss, and demonstrate that state-of-the-art\nresults can be achieved at all processing levels. We also show that in the wild\nour monocular RGB architecture is perceptually competitive to a state-of-the\nart (commercial) Kinect system based on RGB-D data.","url_abs":"http://arxiv.org/abs/1701.08985v1","url_pdf":"http://arxiv.org/pdf/1701.08985v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-human-pose-estimation","task_name":"3D Human Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-human-pose-estimation-on-humaneva-i","task":"3D Human Pose Estimation","dataset":"HumanEva-I","model":"DMHSR(J,B,D)","rank_in_archive_order":23,"of":31,"metrics":{"Mean Reconstruction Error (mm)":"33.7"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1701.08985","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}