{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rethinking-on-multi-stage-networks-for-human","title":"Rethinking on Multi-Stage Networks for Human Pose Estimation","arxiv_id":"1901.00148","date":"2019-01-01","proceeding":null,"authors":["Wenbo Li","Zhicheng Wang","Binyi Yin","Qixiang Peng","Yuming Du","Tianzi Xiao","Gang Yu","Hongtao Lu","Yichen Wei","Jian Sun"],"abstract":"Existing pose estimation approaches fall into two categories: single-stage and multi-stage methods. While multi-stage methods are seemingly more suited for the task, their performance in current practice is not as good as single-stage methods. This work studies this issue. We argue that the current multi-stage methods' unsatisfactory performance comes from the insufficiency in various design choices. We propose several improvements, including the single-stage module design, cross stage feature aggregation, and coarse-to-fine supervision. The resulting method establishes the new state-of-the-art on both MS COCO and MPII Human Pose dataset, justifying the effectiveness of a multi-stage architecture. The source code is publicly available for further research.","url_abs":"https://arxiv.org/abs/1901.00148v4","url_pdf":"https://arxiv.org/pdf/1901.00148v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rethinking-on-multi-stage-networks-for-human","repo_url":"https://github.com/megvii-detection/MSPN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"rethinking-on-multi-stage-networks-for-human","repo_url":"https://github.com/chenyilun95/tf-cpn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"rethinking-on-multi-stage-networks-for-human","repo_url":"https://github.com/fenglinglwb/MSPN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"rethinking-on-multi-stage-networks-for-human","repo_url":"https://github.com/hyperionfalling/lightpose","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"rethinking-on-multi-stage-networks-for-human","repo_url":"https://github.com/open-mmlab/mmpose","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"rethinking-on-multi-stage-networks-for-human","repo_url":"https://github.com/xiuyu0000/papers_with_examples/tree/main/mspn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"rethinking-on-multi-stage-networks-for-human","repo_url":"https://github.com/yangyucheng000/MSPN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":{"status":"ok"}}],"tasks":[{"task_slug":"keypoint-detection","task_name":"Keypoint Detection"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/keypoint-detection-on-coco","task":"Keypoint Detection","dataset":"COCO (Common Objects in Context)","model":"MSPN(384x288)","rank_in_archive_order":6,"of":24,"metrics":{"Test AP":"76.1"},"uses_additional_data":false},{"leaderboard":"/sota/keypoint-detection-on-coco-test-challenge","task":"Keypoint Detection","dataset":"COCO test-challenge","model":"MSPN+*","rank_in_archive_order":2,"of":8,"metrics":{"AP":"76.4","AP50":"92.9","AP75":"82.6","APL":"88.6","AR":"82.2","AR50":"96","AR75":"87.7","ARL":"83.2","ARM":"77.5"},"uses_additional_data":false},{"leaderboard":"/sota/keypoint-detection-on-coco-test-dev","task":"Keypoint Detection","dataset":"COCO test-dev","model":"MSPN","rank_in_archive_order":3,"of":16,"metrics":{"AP":"76.1","AP50":"93.4","AP75":"83.8","APL":"81.5","APM":"72.3","AR":"81.6","AR50":"96.3","AR75":"88.1","ARL":"87.1","ARM":"77.5"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-coco-minival","task":"Pose Estimation","dataset":"COCO minival","model":"MSPN","rank_in_archive_order":1,"of":1,"metrics":{"AP":"75.9"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-coco-test-dev","task":"Pose Estimation","dataset":"COCO test-dev","model":"MSPN","rank_in_archive_order":18,"of":47,"metrics":{"AP":"76.1","AP50":"93.4","AP75":"83.8","APL":"81.5","APM":"72.3","AR":"81.6"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-mpii-human-pose","task":"Pose Estimation","dataset":"MPII Human Pose","model":"MSPN","rank_in_archive_order":9,"of":46,"metrics":{"PCKh-0.5":"92.6"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1901.00148","atlas_url":"https://app.syntology.ai/?focus=1901.00148","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}