{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-accurate-multi-person-pose-estimation","title":"Towards Accurate Multi-person Pose Estimation in the Wild","arxiv_id":"1701.01779","date":"2017-01-06","proceeding":"CVPR 2017 7","authors":["George Papandreou","Tyler Zhu","Nori Kanazawa","Alexander Toshev","Jonathan Tompson","Chris Bregler","Kevin Murphy"],"abstract":"We propose a method for multi-person detection and 2-D pose estimation that\nachieves state-of-art results on the challenging COCO keypoints task. It is a\nsimple, yet powerful, top-down approach consisting of two stages.\n  In the first stage, we predict the location and scale of boxes which are\nlikely to contain people; for this we use the Faster RCNN detector. In the\nsecond stage, we estimate the keypoints of the person potentially contained in\neach proposed bounding box. For each keypoint type we predict dense heatmaps\nand offsets using a fully convolutional ResNet. To combine these outputs we\nintroduce a novel aggregation procedure to obtain highly localized keypoint\npredictions. We also use a novel form of keypoint-based Non-Maximum-Suppression\n(NMS), instead of the cruder box-level NMS, and a novel form of keypoint-based\nconfidence score estimation, instead of box-level scoring.\n  Trained on COCO data alone, our final system achieves average precision of\n0.649 on the COCO test-dev set and the 0.643 test-standard sets, outperforming\nthe winner of the 2016 COCO keypoints challenge and other recent state-of-art.\nFurther, by using additional in-house labeled data we obtain an even higher\naverage precision of 0.685 on the test-dev set and 0.673 on the test-standard\nset, more than 5% absolute improvement compared to the previous best performing\nmethod on the same dataset.","url_abs":"http://arxiv.org/abs/1701.01779v2","url_pdf":"http://arxiv.org/pdf/1701.01779v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"human-detection","task_name":"Human Detection"},{"task_slug":"keypoint-detection","task_name":"Keypoint Detection"},{"task_slug":"multi-person-pose-estimation","task_name":"Multi-Person Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/keypoint-detection-on-coco-test-challenge","task":"Keypoint Detection","dataset":"COCO test-challenge","model":"G-RMI*","rank_in_archive_order":6,"of":8,"metrics":{"AP":"69.1","AP50":"85.9","AP75":"75.2","APL":"82.4","AR":"75.1","AR50":"90.7","AR75":"80.7","ARL":"74.5","ARM":"69.7"},"uses_additional_data":false},{"leaderboard":"/sota/keypoint-detection-on-coco-test-dev","task":"Keypoint Detection","dataset":"COCO test-dev","model":"G-RMI","rank_in_archive_order":11,"of":16,"metrics":{"AP50":"85.5","AP75":"71.3","APL":"70.0","APM":"62.3","AR":"69.7","AR50":"88.7","AR75":"75.5","ARL":"77.1","ARM":"64.4"},"uses_additional_data":false},{"leaderboard":"/sota/multi-person-pose-estimation-on-coco","task":"Multi-Person Pose Estimation","dataset":"COCO (Common Objects in Context)","model":"G-RMI*","rank_in_archive_order":9,"of":15,"metrics":{"AP":"0.685"},"uses_additional_data":false},{"leaderboard":"/sota/multi-person-pose-estimation-on-coco","task":"Multi-Person Pose Estimation","dataset":"COCO (Common Objects in Context)","model":"G-RMI","rank_in_archive_order":12,"of":15,"metrics":{"AP":"0.649"},"uses_additional_data":false},{"leaderboard":"/sota/multi-person-pose-estimation-on-coco-test-dev","task":"Multi-Person Pose Estimation","dataset":"COCO test-dev","model":"G-RMI","rank_in_archive_order":13,"of":15,"metrics":{"AP":"64.9","AP50":"85.5","AP75":"71.3","APL":"70.0","APM":"62.3"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-coco-test-dev","task":"Pose Estimation","dataset":"COCO test-dev","model":"G-RMI","rank_in_archive_order":38,"of":47,"metrics":{"AP":"64.9","AP50":"85.5","AP75":"71.3","APL":"70.0","AR":"69.7"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1701.01779","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}