{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-hypothesis-pose-networks-rethinking-top","title":"Multi-Instance Pose Networks: Rethinking Top-Down Pose Estimation","arxiv_id":"2101.11223","date":"2021-01-27","proceeding":"ICCV 2021 10","authors":["Rawal Khirodkar","Visesh Chari","Amit Agrawal","Ambrish Tyagi"],"abstract":"A key assumption of top-down human pose estimation approaches is their expectation of having a single person/instance present in the input bounding box. This often leads to failures in crowded scenes with occlusions. We propose a novel solution to overcome the limitations of this fundamental assumption. Our Multi-Instance Pose Network (MIPNet) allows for predicting multiple 2D pose instances within a given bounding box. We introduce a Multi-Instance Modulation Block (MIMB) that can adaptively modulate channel-wise feature responses for each instance and is parameter efficient. We demonstrate the efficacy of our approach by evaluating on COCO, CrowdPose, and OCHuman datasets. Specifically, we achieve 70.0 AP on CrowdPose and 42.5 AP on OCHuman test sets, a significant improvement of 2.4 AP and 6.5 AP over the prior art, respectively. When using ground truth bounding boxes for inference, MIPNet achieves an improvement of 0.7 AP on COCO, 0.9 AP on CrowdPose, and 9.1 AP on OCHuman validation sets compared to HRNet. Interestingly, when fewer, high confidence bounding boxes are used, HRNet's performance degrades (by 5 AP) on OCHuman, whereas MIPNet maintains a relatively stable performance (drop of 1 AP) for the same inputs.","url_abs":"https://arxiv.org/abs/2101.11223v3","url_pdf":"https://arxiv.org/pdf/2101.11223v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-hypothesis-pose-networks-rethinking-top","repo_url":"https://github.com/rawalkhirodkar/MIPNet","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"2d-human-pose-estimation","task_name":"2D Human Pose Estimation"},{"task_slug":"keypoint-detection","task_name":"Keypoint Detection"},{"task_slug":"multi-person-pose-estimation","task_name":"Multi-Person Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"hrnet","method_name":"HRNet"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/2d-human-pose-estimation-on-ochuman","task":"2D Human Pose Estimation","dataset":"OCHuman","model":"MIPNet (HRNet-W48)","rank_in_archive_order":3,"of":11,"metrics":{"Test AP":"42.5","Validation AP":"42.0"},"uses_additional_data":false},{"leaderboard":"/sota/2d-human-pose-estimation-on-ochuman","task":"2D Human Pose Estimation","dataset":"OCHuman","model":"HRNet-W48","rank_in_archive_order":4,"of":11,"metrics":{"Test AP":"37.2","Validation AP":"37.8"},"uses_additional_data":false},{"leaderboard":"/sota/keypoint-detection-on-coco","task":"Keypoint Detection","dataset":"COCO (Common Objects in Context)","model":"MIPNet(384x288)","rank_in_archive_order":8,"of":24,"metrics":{"Test AP":"75.7","Validation AP":"76.3"},"uses_additional_data":false},{"leaderboard":"/sota/keypoint-detection-on-ochuman","task":"Keypoint Detection","dataset":"OCHuman","model":"MIPNet (HRNet-W48)","rank_in_archive_order":2,"of":10,"metrics":{"Test AP":"42.5","Validation AP":"42.0"},"uses_additional_data":false},{"leaderboard":"/sota/keypoint-detection-on-ochuman","task":"Keypoint Detection","dataset":"OCHuman","model":"HRNet-W48","rank_in_archive_order":3,"of":10,"metrics":{"Test AP":"37.2","Validation AP":"37.8"},"uses_additional_data":false},{"leaderboard":"/sota/multi-person-pose-estimation-on-crowdpose","task":"Multi-Person Pose Estimation","dataset":"CrowdPose","model":"MIPNet (HRNet-W48)","rank_in_archive_order":13,"of":28,"metrics":{"AP Easy":"78.1","AP Hard":"59.4","AP Medium":"71.1","mAP @0.5:0.95":"70.0"},"uses_additional_data":false},{"leaderboard":"/sota/multi-person-pose-estimation-on-ochuman","task":"Multi-Person Pose Estimation","dataset":"OCHuman","model":"MIPNet (gt-bb)","rank_in_archive_order":1,"of":8,"metrics":{"AP50":"89.7","AP75":"80.1","Validation AP":"74.1"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-coco-test-dev","task":"Pose Estimation","dataset":"COCO test-dev","model":"MIPNet","rank_in_archive_order":19,"of":47,"metrics":{"AP":"75.7","AP50":"92.4","AP75":"83.3","APL":"81.2","APM":"71.4","AR":"80.5"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-crowdpose","task":"Pose Estimation","dataset":"CrowdPose","model":"MIPNet (HRNet-W48)","rank_in_archive_order":8,"of":12,"metrics":{"AP":"70.0","AP Hard":"59.4","APM":"71.1"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-ochuman","task":"Pose Estimation","dataset":"OCHuman","model":"MIPNet (HRNet-W48)","rank_in_archive_order":10,"of":19,"metrics":{"Test AP":"42.5","Validation AP":"42.0"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-ochuman","task":"Pose Estimation","dataset":"OCHuman","model":"HRNet-W48","rank_in_archive_order":12,"of":19,"metrics":{"Test AP":"37.2","Validation AP":"37.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2101.11223","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2101.11223"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rawalkhirodkar/MIPNet","reach":null}],"summary":{"ran":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e6fdde43c1dba636","entry":"SELambdaLayer","repo":"rawalkhirodkar/MIPNet","repo_kind":"official","path":"lib/models/pose_hrnet_se_lambda.py","file_url":"https://github.com/rawalkhirodkar/MIPNet/blob/HEAD/lib/models/pose_hrnet_se_lambda.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"e6fdde43c1dba636"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}