{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/human-pose-estimation-from-depth-images-via","title":"Human Pose Estimation from Depth Images via Inference Embedded Multi-task Learning","arxiv_id":"1608.03932","date":"2016-08-13","proceeding":null,"authors":["Keze Wang","Shengfu Zhai","Hui Cheng","Xiaodan Liang","Liang Lin"],"abstract":"Human pose estimation (i.e., locating the body parts / joints of a person) is\na fundamental problem in human-computer interaction and multimedia\napplications. Significant progress has been made based on the development of\ndepth sensors, i.e., accessible human pose prediction from still depth images\n[32]. However, most of the existing approaches to this problem involve several\ncomponents/models that are independently designed and optimized, leading to\nsuboptimal performances. In this paper, we propose a novel inference-embedded\nmulti-task learning framework for predicting human pose from still depth\nimages, which is implemented with a deep architecture of neural networks.\nSpecifically, we handle two cascaded tasks: i) generating the heat (confidence)\nmaps of body parts via a fully convolutional network (FCN); ii) seeking the\noptimal configuration of body parts based on the detected body part proposals\nvia an inference built-in MatchNet [10], which measures the appearance and\ngeometric kinematic compatibility of body parts and embodies the dynamic\nprogramming inference as an extra network layer. These two tasks are jointly\noptimized. Our extensive experiments show that the proposed deep model\nsignificantly improves the accuracy of human pose estimation over other several\nstate-of-the-art methods or SDKs. We also release a large-scale dataset for\ncomparison, which includes 100K depth images under challenging scenarios.","url_abs":"http://arxiv.org/abs/1608.03932v1","url_pdf":"http://arxiv.org/pdf/1608.03932v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"pose-prediction","task_name":"Pose Prediction"}],"methods":[],"datasets_introduced":[{"slug":"k2hpd","name":"K2HPD","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1608.03932","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}