{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-scale-structure-aware-network-for-human","title":"Multi-Scale Structure-Aware Network for Human Pose Estimation","arxiv_id":"1803.09894","date":"2018-03-27","proceeding":"ECCV 2018 9","authors":["Lipeng Ke","Ming-Ching Chang","Honggang Qi","Siwei Lyu"],"abstract":"We develop a robust multi-scale structure-aware neural network for human pose\nestimation. This method improves the recent deep conv-deconv hourglass models\nwith four key improvements: (1) multi-scale supervision to strengthen\ncontextual feature learning in matching body keypoints by combining feature\nheatmaps across scales, (2) multi-scale regression network at the end to\nglobally optimize the structural matching of the multi-scale features, (3)\nstructure-aware loss used in the intermediate supervision and at the regression\nto improve the matching of keypoints and respective neighbors to infer a\nhigher-order matching configurations, and (4) a keypoint masking training\nscheme that can effectively fine-tune our network to robustly localize occluded\nkeypoints via adjacent matches. Our method can effectively improve\nstate-of-the-art pose estimation methods that suffer from difficulties in scale\nvarieties, occlusions, and complex multi-person scenarios. This multi-scale\nsupervision tightly integrates with the regression network to effectively (i)\nlocalize keypoints using the ensemble of multi-scale features, and (ii) infer\nglobal pose configuration by maximizing structural consistencies across\nmultiple keypoints and scales. The keypoint masking training enhances these\nadvantages to focus learning on hard occlusion samples. Our method achieves the\nleading position in the MPII challenge leaderboard among the state-of-the-art\nmethods.","url_abs":"http://arxiv.org/abs/1803.09894v3","url_pdf":"http://arxiv.org/pdf/1803.09894v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/pose-estimation-on-mpii-human-pose","task":"Pose Estimation","dataset":"MPII Human Pose","model":"Multi-Scale Structure-Aware Network","rank_in_archive_order":13,"of":46,"metrics":{"PCKh-0.5":"92.1"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.09894","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}