{"url":"/sota/pose-estimation-on-ochuman","task":{"name":"Pose Estimation","url":"/task/pose-estimation","note":null},"dataset":{"name":"OCHuman","url":"/dataset/ochuman"},"category":"Computer Vision","categories":["Computer Vision"],"category_note":null,"description":"**Pose Estimation** is a computer vision task where the goal is to detect the position and orientation of a person or an object. Usually, this is done by predicting the location of specific keypoints like hands, head, elbows, etc. in case of Human Pose Estimation.\r\n\r\nA common benchmark for this task is [MPII Human Pose](https://paperswithcode.com/sota/pose-estimation-on-mpii-human-pose)\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [Real-time 2D Multi-Person Pose Estimation on CPU: Lightweight OpenPose](https://github.com/Daniil-Osokin/lightweight-human-pose-estimation.pytorch) )</span>","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["Test AP","Validation AP"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"Test AP":"higher","Validation AP":"higher"}},"counts":{"rows":19,"rows_with_code":18,"rows_with_paper_page":19,"rows_dated":19,"rows_using_additional_data":5},"rows":[{"rank_in_archive_order":1,"model":"ViTPose (ViTAE-G, GT bounding boxes)","metrics":{"Test AP":"93.3","Validation AP":"92.8"},"uses_additional_data":true,"paper_date":"2022-04-26","paper":"/paper/vitpose-simple-vision-transformer-baselines","paper_url":"https://arxiv.org/abs/2204.12484v3","paper_title":"ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation","code":"https://github.com/huggingface/transformers","n_code_links":6,"syntology":{"n_ran":18,"n_unverified":13,"n_samples":31,"n_pointer_only_licence":6}},{"rank_in_archive_order":2,"model":"UniHCP (direct eval)","metrics":{"Test AP":"87.4"},"uses_additional_data":false,"paper_date":"2023-03-06","paper":"/paper/unihcp-a-unified-model-for-human-centric","paper_url":"https://arxiv.org/abs/2303.02936v4","paper_title":"UniHCP: A Unified Model for Human-Centric Perceptions","code":"https://github.com/opengvlab/unihcp","n_code_links":1,"syntology":{"n_ran":7,"n_unverified":6,"n_samples":13,"n_pointer_only_licence":0}},{"rank_in_archive_order":3,"model":"PoseBH-H","metrics":{"Test AP":"87.0","Validation AP":"86.0"},"uses_additional_data":true,"paper_date":"2025-05-23","paper":"/paper/posebh-prototypical-multi-dataset-training","paper_url":"https://arxiv.org/abs/2505.17475v1","paper_title":"PoseBH: Prototypical Multi-Dataset Training Beyond Human Pose Estimation","code":"https://github.com/uyoung-jeong/PoseBH","n_code_links":1,"syntology":null},{"rank_in_archive_order":4,"model":"RTMPose(RTMPose-l, GT bounding boxes)","metrics":{"Test AP":"80.3","Validation AP":"80.5"},"uses_additional_data":true,"paper_date":"2023-03-13","paper":"/paper/rtmpose-real-time-multi-person-pose","paper_url":"https://arxiv.org/abs/2303.07399v2","paper_title":"RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose","code":"https://github.com/open-mmlab/mmpose","n_code_links":1,"syntology":null},{"rank_in_archive_order":5,"model":"BBox-Mask-Pose 2x","metrics":{"Test AP":"48.3","Validation AP":"48.6"},"uses_additional_data":true,"paper_date":"2024-12-02","paper":"/paper/detection-pose-estimation-and-segmentation-1","paper_url":"https://arxiv.org/abs/2412.01562v1","paper_title":"Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous Circle","code":"https://github.com/MiraPurkrabek/BBoxMaskPose","n_code_links":1,"syntology":null},{"rank_in_archive_order":6,"model":"BUCTD (CID-W32)","metrics":{"Test AP":"47.2","Validation AP":"47.7"},"uses_additional_data":false,"paper_date":"2023-06-13","paper":"/paper/rethinking-pose-estimation-in-crowds","paper_url":"https://arxiv.org/abs/2306.07879v2","paper_title":"Rethinking pose estimation in crowds: overcoming the detection information-bottleneck and ambiguity","code":"https://github.com/amathislab/BUCTD","n_code_links":1,"syntology":{"n_ran":4,"n_unverified":2,"n_samples":6,"n_pointer_only_licence":0}},{"rank_in_archive_order":7,"model":"HQNet (ViT-L)","metrics":{"Test AP":"45.6"},"uses_additional_data":false,"paper_date":"2023-12-09","paper":"/paper/you-only-learn-one-query-learning-unified","paper_url":"https://arxiv.org/abs/2312.05525v3","paper_title":"You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-Person Multi-Task Human-Centric Perception","code":"https://github.com/lishuhuai527/coco-unihuman","n_code_links":1,"syntology":null},{"rank_in_archive_order":8,"model":"CID (HRNet-W48)","metrics":{"Test AP":"45.0","Validation AP":"46.1"},"uses_additional_data":false,"paper_date":"2022-01-01","paper":"/paper/contextual-instance-decoupling-for-robust","paper_url":"http://openaccess.thecvf.com//content/CVPR2022/html/Wang_Contextual_Instance_Decoupling_for_Robust_Multi-Person_Pose_Estimation_CVPR_2022_paper.html","paper_title":"Contextual Instance Decoupling for Robust Multi-Person Pose Estimation","code":"https://github.com/kennethwdk/cid","n_code_links":1,"syntology":null},{"rank_in_archive_order":9,"model":"MaskPose-b","metrics":{"Test AP":"45.0","Validation AP":"45.3"},"uses_additional_data":true,"paper_date":"2024-12-02","paper":"/paper/detection-pose-estimation-and-segmentation-1","paper_url":"https://arxiv.org/abs/2412.01562v1","paper_title":"Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous Circle","code":"https://github.com/MiraPurkrabek/BBoxMaskPose","n_code_links":1,"syntology":null},{"rank_in_archive_order":10,"model":"MIPNet (HRNet-W48)","metrics":{"Test AP":"42.5","Validation AP":"42.0"},"uses_additional_data":false,"paper_date":"2021-01-27","paper":"/paper/multi-hypothesis-pose-networks-rethinking-top","paper_url":"https://arxiv.org/abs/2101.11223v3","paper_title":"Multi-Instance Pose Networks: Rethinking Top-Down Pose Estimation","code":"https://github.com/rawalkhirodkar/MIPNet","n_code_links":1,"syntology":{"n_ran":1,"n_unverified":0,"n_samples":1,"n_pointer_only_licence":0}},{"rank_in_archive_order":11,"model":"HQNet (ResNet-50)","metrics":{"Test AP":"40.0"},"uses_additional_data":false,"paper_date":"2023-12-09","paper":"/paper/you-only-learn-one-query-learning-unified","paper_url":"https://arxiv.org/abs/2312.05525v3","paper_title":"You Only Learn One Query: Learning Unified Human Query for Single-Stage Multi-Person Multi-Task Human-Centric Perception","code":"https://github.com/lishuhuai527/coco-unihuman","n_code_links":1,"syntology":null},{"rank_in_archive_order":12,"model":"HRNet-W48","metrics":{"Test AP":"37.2","Validation AP":"37.8"},"uses_additional_data":false,"paper_date":"2021-01-27","paper":"/paper/multi-hypothesis-pose-networks-rethinking-top","paper_url":"https://arxiv.org/abs/2101.11223v3","paper_title":"Multi-Instance Pose Networks: Rethinking Top-Down Pose Estimation","code":"https://github.com/rawalkhirodkar/MIPNet","n_code_links":1,"syntology":{"n_ran":1,"n_unverified":0,"n_samples":1,"n_pointer_only_licence":0}},{"rank_in_archive_order":13,"model":"HGG (AE+)","metrics":{"Test AP":"36.0","Validation AP":"41.8"},"uses_additional_data":false,"paper_date":"2020-07-23","paper":"/paper/differentiable-hierarchical-graph-grouping","paper_url":"https://arxiv.org/abs/2007.11864v1","paper_title":"Differentiable Hierarchical Graph Grouping for Multi-Person Pose Estimation","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":14,"model":"ResNet-152","metrics":{"Test AP":"33.3","Validation AP":"41.0"},"uses_additional_data":false,"paper_date":"2018-04-17","paper":"/paper/simple-baselines-for-human-pose-estimation","paper_url":"http://arxiv.org/abs/1804.06208v2","paper_title":"Simple Baselines for Human Pose Estimation and Tracking","code":"https://github.com/PaddlePaddle/PaddleDetection","n_code_links":27,"syntology":null},{"rank_in_archive_order":15,"model":"Associative Embedding+","metrics":{"Test AP":"32.8","Validation AP":"40.0"},"uses_additional_data":false,"paper_date":"2016-11-16","paper":"/paper/associative-embedding-end-to-end-learning-for","paper_url":"http://arxiv.org/abs/1611.05424v2","paper_title":"Associative Embedding: End-to-End Learning for Joint Detection and Grouping","code":"https://github.com/open-mmlab/mmpose","n_code_links":5,"syntology":{"n_ran":0,"n_unverified":8,"n_samples":8,"n_pointer_only_licence":0}},{"rank_in_archive_order":16,"model":"RMPE","metrics":{"Test AP":"30.7","Validation AP":"38.8"},"uses_additional_data":false,"paper_date":"2016-12-01","paper":"/paper/rmpe-regional-multi-person-pose-estimation","paper_url":"http://arxiv.org/abs/1612.00137v5","paper_title":"RMPE: Regional Multi-person Pose Estimation","code":"https://github.com/MVIG-SJTU/AlphaPose","n_code_links":14,"syntology":null},{"rank_in_archive_order":17,"model":"ResNet-50","metrics":{"Test AP":"29.5","Validation AP":"32.1"},"uses_additional_data":false,"paper_date":"2018-04-17","paper":"/paper/simple-baselines-for-human-pose-estimation","paper_url":"http://arxiv.org/abs/1804.06208v2","paper_title":"Simple Baselines for Human Pose Estimation and Tracking","code":"https://github.com/PaddlePaddle/PaddleDetection","n_code_links":27,"syntology":null},{"rank_in_archive_order":18,"model":"Associative Embedding","metrics":{"Test AP":"29.5","Validation AP":"32.1"},"uses_additional_data":false,"paper_date":"2016-11-16","paper":"/paper/associative-embedding-end-to-end-learning-for","paper_url":"http://arxiv.org/abs/1611.05424v2","paper_title":"Associative Embedding: End-to-End Learning for Joint Detection and Grouping","code":"https://github.com/open-mmlab/mmpose","n_code_links":5,"syntology":{"n_ran":0,"n_unverified":8,"n_samples":8,"n_pointer_only_licence":0}},{"rank_in_archive_order":19,"model":"TransPose-H","metrics":{"Validation AP":"62.3"},"uses_additional_data":false,"paper_date":"2020-12-28","paper":"/paper/transpose-towards-explainable-human-pose","paper_url":"https://arxiv.org/abs/2012.14214v5","paper_title":"TransPose: Keypoint Localization via Transformer","code":"https://github.com/yangsenius/TransPose","n_code_links":1,"syntology":null}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":7,"rows_with_any_sample_ran":5,"distinct_papers_with_graph_line":5,"distinct_papers_with_any_sample_ran":4,"samples_over_distinct_papers":{"n_ran":30,"n_unverified":29,"n_samples":59,"n_pointer_only_licence":6,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":31,"n_unverified":37,"n_samples":68,"n_pointer_only_licence":6,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}