{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-feature-pyramids-for-human-pose","title":"Learning Feature Pyramids for Human Pose Estimation","arxiv_id":"1708.01101","date":"2017-08-03","proceeding":"ICCV 2017 10","authors":["Wei Yang","Shuang Li","Wanli Ouyang","Hongsheng Li","Xiaogang Wang"],"abstract":"Articulated human pose estimation is a fundamental yet challenging task in\ncomputer vision. The difficulty is particularly pronounced in scale variations\nof human body parts when camera view changes or severe foreshortening happens.\nAlthough pyramid methods are widely used to handle scale changes at inference\ntime, learning feature pyramids in deep convolutional neural networks (DCNNs)\nis still not well explored. In this work, we design a Pyramid Residual Module\n(PRMs) to enhance the invariance in scales of DCNNs. Given input features, the\nPRMs learn convolutional filters on various scales of input features, which are\nobtained with different subsampling ratios in a multi-branch network. Moreover,\nwe observe that it is inappropriate to adopt existing methods to initialize the\nweights of multi-branch networks, which achieve superior performance than plain\nnetworks in many tasks recently. Therefore, we provide theoretic derivation to\nextend the current weight initialization scheme to multi-branch network\nstructures. We investigate our method on two standard benchmarks for human pose\nestimation. Our approach obtains state-of-the-art results on both benchmarks.\nCode is available at https://github.com/bearpaw/PyraNet.","url_abs":"http://arxiv.org/abs/1708.01101v1","url_pdf":"http://arxiv.org/pdf/1708.01101v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-feature-pyramids-for-human-pose","repo_url":"https://github.com/bearpaw/PyraNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"torch","reach":{"status":"unanswered"}},{"paper_slug":"learning-feature-pyramids-for-human-pose","repo_url":"https://github.com/wanggrun/Learning-Feature-Pyramids","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"learning-feature-pyramids-for-human-pose","repo_url":"https://github.com/wanggrun/Learning-Feature-Pyramids-For-COCO","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/pose-estimation-on-leeds-sports-poses","task":"Pose Estimation","dataset":"Leeds Sports Poses","model":"Pyramid Residual Modules (PRMs)","rank_in_archive_order":6,"of":18,"metrics":{"PCK":"93.9%"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-mpii-human-pose","task":"Pose Estimation","dataset":"MPII Human Pose","model":"Pyramid Residual Modules (PRMs)","rank_in_archive_order":14,"of":46,"metrics":{"PCKh-0.5":"92.0"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.01101","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}