{"url":"/task/2d-human-pose-estimation","name":"2D Human Pose Estimation","slug":"2d-human-pose-estimation","description_markdown":"What is Human Pose Estimation?\r\nHuman pose estimation is the process of estimating the configuration of the body (pose) from a single, typically monocular, image. Background. Human pose estimation is one of the key problems in computer vision that has been studied for well over 15 years.  The reason for its importance is the\r\nabundance of applications that can benefit from such a technology. For example,\r\nhuman pose estimation allows for higher-level reasoning in the context of human-computer interaction and activity recognition; it is also one of the basic building blocks for marker-less motion capture (MoCap) technology. MoCap technology is useful for applications ranging from character animation to clinical analysis of gait pathologies.","categories":[{"name":"Computer Vision","url":"/area/computer-vision"},{"name":"Knowledge Base","url":"/area/knowledge-base"},{"name":"Reasoning","url":"/area/reasoning"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":118,"papers_with_code":68,"benchmarks":10,"benchmark_tables_in_archive":10,"benchmark_tables_shown":10,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":25,"subtasks":6,"parent_tasks":0},"benchmarks":[{"leaderboard":"/sota/2d-human-pose-estimation-on-coco-wholebody-1","slug":"2d-human-pose-estimation-on-coco-wholebody-1","dataset":"COCO-WholeBody","dataset_url":"/dataset/coco-wholebody","rows_in_archive":15,"metrics":["WB","body","foot","face","hand"],"first_row_in_archive_order":{"model":"RTMW-x","paper_title":"RTMW: Real-Time Multi-Person 2D and 3D Whole-body Pose Estimation","paper_url":"/paper/rtmw-real-time-multi-person-2d-and-3d-whole","paper_date":"2024-07-11","arxiv_id":"2407.08634","code_links":[{"title":"open-mmlab/mmpose","url":"https://github.com/open-mmlab/mmpose"}],"syntology":null}},{"leaderboard":"/sota/2d-human-pose-estimation-on-ochuman","slug":"2d-human-pose-estimation-on-ochuman","dataset":"OCHuman","dataset_url":"/dataset/ochuman","rows_in_archive":11,"metrics":["Test AP","Validation AP"],"first_row_in_archive_order":{"model":"BBox-Mask-Pose 2x","paper_title":"Detection, Pose Estimation and Segmentation for Multiple Bodies: Closing the Virtuous Circle","paper_url":"/paper/detection-pose-estimation-and-segmentation-1","paper_date":"2024-12-02","arxiv_id":"2412.01562","code_links":[{"title":"MiraPurkrabek/BBoxMaskPose","url":"https://github.com/MiraPurkrabek/BBoxMaskPose"}],"syntology":null}},{"leaderboard":"/sota/2d-human-pose-estimation-on-human-art","slug":"2d-human-pose-estimation-on-human-art","dataset":"Human-Art","dataset_url":"/dataset/human-art","rows_in_archive":10,"metrics":["AP","AP (gt bbox)","Validation AP"],"first_row_in_archive_order":{"model":"UniPose","paper_title":"X-Pose: Detecting Any Keypoints","paper_url":"/paper/unipose-detecting-any-keypoints","paper_date":"2023-10-12","arxiv_id":"2310.08530","code_links":[{"title":"idea-research/x-pose","url":"https://github.com/idea-research/x-pose"},{"title":"IDEA-Research/UniPose","url":"https://github.com/IDEA-Research/UniPose"}],"syntology":{"n":13,"n_ran":11,"n_unverified":2,"n_pointer_only":13}}},{"leaderboard":"/sota/2d-human-pose-estimation-on-jhmdb-2d-poses","slug":"2d-human-pose-estimation-on-jhmdb-2d-poses","dataset":"JHMDB (2D poses only)","dataset_url":"/dataset/jhmdb","rows_in_archive":5,"metrics":["PCK"],"first_row_in_archive_order":{"model":"DeciWatch","paper_title":"DeciWatch: A Simple Baseline for 10x Efficient 2D and 3D Pose Estimation","paper_url":"/paper/deciwatch-a-simple-baseline-for-10x-efficient","paper_date":"2022-03-16","arxiv_id":"2203.08713","code_links":[{"title":"cure-lab/DeciWatch","url":"https://github.com/cure-lab/DeciWatch"}],"syntology":{"n":8,"n_ran":6,"n_unverified":2,"n_pointer_only":0}}},{"leaderboard":"/sota/2d-human-pose-estimation-on-alibaba-cluster","slug":"2d-human-pose-estimation-on-alibaba-cluster","dataset":"Alibaba Cluster Trace","dataset_url":"/dataset/alibaba-cluster-trace","rows_in_archive":1,"metrics":["10-20% Mask PSNR"],"first_row_in_archive_order":{"model":"mitsimpo","paper_title":"Alibaba at IJCNLP-2017 Task 1: Embedding Grammatical Features into LSTMs for Chinese Grammatical Error Diagnosis Task","paper_url":"/paper/alibaba-at-ijcnlp-2017-task-1-embedding","paper_date":"2017-12-01","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/2d-human-pose-estimation-on-exlpose-ll-e","slug":"2d-human-pose-estimation-on-exlpose-ll-e","dataset":"ExLPose-LL-E","dataset_url":null,"rows_in_archive":1,"metrics":["AP"],"first_row_in_archive_order":{"model":"DA-LLPose","paper_title":"Domain-Adaptive 2D Human Pose Estimation via Dual Teachers in Extremely Low-Light Conditions","paper_url":"/paper/domain-adaptive-2d-human-pose-estimation-via","paper_date":"2024-07-22","arxiv_id":"2407.15451","code_links":[{"title":"ayh015-dev/da-llpose","url":"https://github.com/ayh015-dev/da-llpose"}],"syntology":{"n":13,"n_ran":11,"n_unverified":2,"n_pointer_only":13}}},{"leaderboard":"/sota/2d-human-pose-estimation-on-exlpose-ll-h","slug":"2d-human-pose-estimation-on-exlpose-ll-h","dataset":"ExLPose-LL-H","dataset_url":null,"rows_in_archive":1,"metrics":["AP"],"first_row_in_archive_order":{"model":"DA-LLPose","paper_title":"Domain-Adaptive 2D Human Pose Estimation via Dual Teachers in Extremely Low-Light Conditions","paper_url":"/paper/domain-adaptive-2d-human-pose-estimation-via","paper_date":"2024-07-22","arxiv_id":"2407.15451","code_links":[{"title":"ayh015-dev/da-llpose","url":"https://github.com/ayh015-dev/da-llpose"}],"syntology":{"n":13,"n_ran":11,"n_unverified":2,"n_pointer_only":13}}},{"leaderboard":"/sota/2d-human-pose-estimation-on-exlpose-ll-n","slug":"2d-human-pose-estimation-on-exlpose-ll-n","dataset":"ExLPose-LL-N","dataset_url":null,"rows_in_archive":1,"metrics":["AP"],"first_row_in_archive_order":{"model":"DA-LLPose","paper_title":"Domain-Adaptive 2D Human Pose Estimation via Dual Teachers in Extremely Low-Light Conditions","paper_url":"/paper/domain-adaptive-2d-human-pose-estimation-via","paper_date":"2024-07-22","arxiv_id":"2407.15451","code_links":[{"title":"ayh015-dev/da-llpose","url":"https://github.com/ayh015-dev/da-llpose"}],"syntology":{"n":13,"n_ran":11,"n_unverified":2,"n_pointer_only":13}}},{"leaderboard":"/sota/2d-human-pose-estimation-on-exlpose-ocn","slug":"2d-human-pose-estimation-on-exlpose-ocn","dataset":"ExLPose-OCN-RICOH3","dataset_url":null,"rows_in_archive":1,"metrics":["AP"],"first_row_in_archive_order":{"model":"DA-LLPose","paper_title":"Domain-Adaptive 2D Human Pose Estimation via Dual Teachers in Extremely Low-Light Conditions","paper_url":"/paper/domain-adaptive-2d-human-pose-estimation-via","paper_date":"2024-07-22","arxiv_id":"2407.15451","code_links":[{"title":"ayh015-dev/da-llpose","url":"https://github.com/ayh015-dev/da-llpose"}],"syntology":{"n":13,"n_ran":11,"n_unverified":2,"n_pointer_only":13}}},{"leaderboard":"/sota/2d-human-pose-estimation-on-exlpose-ocn-a7m3","slug":"2d-human-pose-estimation-on-exlpose-ocn-a7m3","dataset":"ExLPose-OCN-A7M3","dataset_url":null,"rows_in_archive":1,"metrics":["AP"],"first_row_in_archive_order":{"model":"DA-LLPose","paper_title":"Domain-Adaptive 2D Human Pose Estimation via Dual Teachers in Extremely Low-Light Conditions","paper_url":"/paper/domain-adaptive-2d-human-pose-estimation-via","paper_date":"2024-07-22","arxiv_id":"2407.15451","code_links":[{"title":"ayh015-dev/da-llpose","url":"https://github.com/ayh015-dev/da-llpose"}],"syntology":{"n":13,"n_ran":11,"n_unverified":2,"n_pointer_only":13}}}],"datasets":[{"url":"/dataset/jhmdb","name":"JHMDB","full_name":"Joint-annotated Human Motion Data Base","num_papers_in_archive":249},{"url":"/dataset/agora","name":"AGORA","full_name":"","num_papers_in_archive":68},{"url":"/dataset/ochuman","name":"OCHuman","full_name":null,"num_papers_in_archive":66},{"url":"/dataset/uav-human","name":"UAV-Human","full_name":"","num_papers_in_archive":47},{"url":"/dataset/coco-wholebody","name":"COCO-WholeBody","full_name":"","num_papers_in_archive":33},{"url":"/dataset/alibaba-cluster-trace","name":"Alibaba Cluster Trace","full_name":"","num_papers_in_archive":8},{"url":"/dataset/human-art","name":"Human-Art","full_name":"","num_papers_in_archive":7},{"url":"/dataset/deepsport-dataset","name":"DeepSport Dataset","full_name":"","num_papers_in_archive":5},{"url":"/dataset/peoplesanspeople","name":"PeopleSansPeople","full_name":"PeopleSansPeople: A Synthetic Data Generator for Human-Centric Computer Vision","num_papers_in_archive":4},{"url":"/dataset/relative-human","name":"Relative Human","full_name":"","num_papers_in_archive":4},{"url":"/dataset/sportspose","name":"SportsPose","full_name":"SportsPose - A Dynamic 3D sports pose dataset","num_papers_in_archive":4},{"url":"/dataset/anime-drawings-dataset","name":"Anime Drawings Dataset","full_name":"","num_papers_in_archive":3},{"url":"/dataset/fitness-aqa","name":"Fitness-AQA","full_name":"Fitness Action Quality Assessment [ECCV 2022]","num_papers_in_archive":3},{"url":"/dataset/popart","name":"PoPArt","full_name":"Poses of People in Art: A Data Set for Human Pose Estimation in Digital Art History","num_papers_in_archive":3},{"url":"/dataset/bizarre-pose-dataset","name":"Bizarre Pose Dataset","full_name":"Bizarre Pose Dataset of Illustrated Characters","num_papers_in_archive":2},{"url":"/dataset/c2a-dataset-human-detection-in-disaster","name":"C2A: Human Detection in Disaster Scenarios","full_name":"Combination to Application","num_papers_in_archive":2},{"url":"/dataset/jrdb-pose","name":"JRDB-Pose","full_name":"","num_papers_in_archive":2},{"url":"/dataset/mmvr","name":"MMVR","full_name":"Millimeter-wave Multi-View Radar (MMVR) Dataset","num_papers_in_archive":2},{"url":"/dataset/cropcoco","name":"CropCOCO","full_name":"","num_papers_in_archive":1},{"url":"/dataset/freeman","name":"FreeMan","full_name":"","num_papers_in_archive":1},{"url":"/dataset/halpe-fullbody","name":"Halpe-FullBody","full_name":"","num_papers_in_archive":1},{"url":"/dataset/mpii-human-pose-descriptions","name":"MPII Human Pose Descriptions","full_name":"","num_papers_in_archive":1},{"url":"/dataset/nvidia-synthetic-head-dataset","name":"NVIDIA Synthetic Head Dataset","full_name":"","num_papers_in_archive":1},{"url":"/dataset/repogen","name":"RePoGen","full_name":"","num_papers_in_archive":1},{"url":"/dataset/infiniterep","name":"InfiniteRep","full_name":"InfiniteRep","num_papers_in_archive":0}],"subtasks":[{"url":"/task/3d-face-animation","name":"3D Face Animation"},{"url":"/task/action-anticipation","name":"Action Anticipation"},{"url":"/task/articles","name":"Articles"},{"url":"/task/community-question-answering","name":"Community Question Answering"},{"url":"/task/semi-supervised-human-pose-estimation","name":"Semi-Supervised Human Pose Estimation"},{"url":"/task/style-transfer","name":"Style Transfer"}],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":68,"tagged_in_all":118,"items":[{"url":"/paper/realtime-multi-person-2d-pose-estimation","title":"Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields","date":"2016-11-24","arxiv_id":"1611.08050","repositories_listed":61,"syntology":{"n":23,"n_ran":4,"n_unverified":19,"n_pointer_only":4}},{"url":"/paper/deep-high-resolution-representation-learning","title":"Deep High-Resolution Representation Learning for Human Pose Estimation","date":"2019-02-25","arxiv_id":"1902.09212","repositories_listed":39,"syntology":{"n":25,"n_ran":9,"n_unverified":16,"n_pointer_only":0}},{"url":"/paper/simple-baselines-for-human-pose-estimation","title":"Simple Baselines for Human Pose Estimation and Tracking","date":"2018-04-17","arxiv_id":"1804.06208","repositories_listed":27,"syntology":null},{"url":"/paper/bottom-up-higher-resolution-networks-for","title":"HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation","date":"2019-08-27","arxiv_id":"1908.10357","repositories_listed":19,"syntology":{"n":27,"n_ran":11,"n_unverified":16,"n_pointer_only":0}},{"url":"/paper/rmpe-regional-multi-person-pose-estimation","title":"RMPE: Regional Multi-person Pose Estimation","date":"2016-12-01","arxiv_id":"1612.00137","repositories_listed":14,"syntology":null},{"url":"/paper/simple-pose-rethinking-and-improving-a-bottom","title":"Simple Pose: Rethinking and Improving a Bottom-up Approach for Multi-Person Pose Estimation","date":"2019-11-24","arxiv_id":"1911.10529","repositories_listed":8,"syntology":null},{"url":"/paper/blazepose-on-device-real-time-body-pose","title":"BlazePose: On-device Real-time Body Pose tracking","date":"2020-06-17","arxiv_id":"2006.10204","repositories_listed":7,"syntology":null},{"url":"/paper/pose2seg-detection-free-human-instance","title":"Pose2Seg: Detection Free Human Instance Segmentation","date":"2018-03-28","arxiv_id":"1803.10683","repositories_listed":7,"syntology":{"n":16,"n_ran":0,"n_unverified":16,"n_pointer_only":0}},{"url":"/paper/vitpose-simple-vision-transformer-baselines","title":"ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation","date":"2022-04-26","arxiv_id":"2204.12484","repositories_listed":6,"syntology":{"n":31,"n_ran":18,"n_unverified":13,"n_pointer_only":6}},{"url":"/paper/near-optimal-representation-learning-for","title":"Near-Optimal Representation Learning for Hierarchical Reinforcement Learning","date":"2018-10-02","arxiv_id":"1810.01257","repositories_listed":6,"syntology":null},{"url":"/paper/associative-embedding-end-to-end-learning-for","title":"Associative Embedding: End-to-End Learning for Joint Detection and Grouping","date":"2016-11-16","arxiv_id":"1611.05424","repositories_listed":5,"syntology":{"n":8,"n_ran":0,"n_unverified":8,"n_pointer_only":0}},{"url":"/paper/explicit-box-detection-unifies-end-to-end","title":"Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation","date":"2023-02-03","arxiv_id":"2302.01593","repositories_listed":3,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":1}},{"url":"/paper/sapiens-foundation-for-human-vision-models","title":"Sapiens: Foundation for Human Vision Models","date":"2024-08-22","arxiv_id":"2408.12569","repositories_listed":2,"syntology":{"n":11,"n_ran":9,"n_unverified":2,"n_pointer_only":0}},{"url":"/paper/unipose-detecting-any-keypoints","title":"X-Pose: Detecting Any Keypoints","date":"2023-10-12","arxiv_id":"2310.08530","repositories_listed":2,"syntology":{"n":13,"n_ran":11,"n_unverified":2,"n_pointer_only":13}},{"url":"/paper/improving-2d-human-pose-estimation-across","title":"Improving 2D Human Pose Estimation in Rare Camera Views with Synthetic Data","date":"2023-07-13","arxiv_id":"2307.06737","repositories_listed":2,"syntology":null},{"url":"/paper/vitpose-vision-transformer-foundation-model","title":"ViTPose++: Vision Transformer for Generic Body Pose Estimation","date":"2022-12-07","arxiv_id":"2212.04246","repositories_listed":2,"syntology":{"n":3,"n_ran":2,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/ppt-token-pruned-pose-transformer-for","title":"PPT: token-Pruned Pose Transformer for monocular and multi-view human pose estimation","date":"2022-09-16","arxiv_id":"2209.08194","repositories_listed":2,"syntology":null},{"url":"/paper/xmem-long-term-video-object-segmentation-with","title":"XMem: Long-Term Video Object Segmentation with an Atkinson-Shiffrin Memory Model","date":"2022-07-14","arxiv_id":"2207.07115","repositories_listed":2,"syntology":{"n":3,"n_ran":1,"n_unverified":2,"n_pointer_only":2}},{"url":"/paper/smoothnet-a-plug-and-play-network-for","title":"SmoothNet: A Plug-and-Play Network for Refining Human Poses in Videos","date":"2021-12-27","arxiv_id":"2112.13715","repositories_listed":2,"syntology":{"n":1,"n_ran":0,"n_unverified":1,"n_pointer_only":0}},{"url":"/paper/event-neural-networks","title":"Event Neural Networks","date":"2021-12-02","arxiv_id":"2112.00891","repositories_listed":2,"syntology":{"n":7,"n_ran":7,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/estimating-parkinsonism-severity-in-natural","title":"Estimating Parkinsonism Severity in Natural Gait Videos of Older Adults with Dementia","date":"2021-05-07","arxiv_id":"2105.03464","repositories_listed":2,"syntology":null},{"url":"/paper/scalable-bottom-up-hierarchical-clustering","title":"Scalable Hierarchical Agglomerative Clustering","date":"2020-10-22","arxiv_id":"2010.11821","repositories_listed":2,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/whole-body-human-pose-estimation-in-the-wild","title":"Whole-Body Human Pose Estimation in the Wild","date":"2020-07-23","arxiv_id":"2007.11858","repositories_listed":2,"syntology":null},{"url":"/paper/efficienthrnet-efficient-scaling-for","title":"EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation","date":"2020-07-16","arxiv_id":"2007.08090","repositories_listed":2,"syntology":null},{"url":"/paper/pifpaf-composite-fields-for-human-pose","title":"PifPaf: Composite Fields for Human Pose Estimation","date":"2019-03-15","arxiv_id":"1903.06593","repositories_listed":2,"syntology":null},{"url":"/paper/learning-from-synthetic-humans","title":"Learning from Synthetic Humans","date":"2017-01-05","arxiv_id":"1701.01370","repositories_listed":2,"syntology":{"n":2,"n_ran":1,"n_unverified":1,"n_pointer_only":2}},{"url":"/paper/preconditioned-stochastic-gradient-descent","title":"Preconditioned Stochastic Gradient Descent","date":"2015-12-14","arxiv_id":"1512.04202","repositories_listed":2,"syntology":{"n":6,"n_ran":5,"n_unverified":1,"n_pointer_only":6}},{"url":"/paper/2d-human-pose-estimation-new-benchmark-and","title":"2D Human Pose Estimation: New Benchmark and State of the Art Analysis","date":"2014-06-01","arxiv_id":null,"repositories_listed":2,"syntology":null},{"url":"/paper/poseidon-a-vit-based-architecture-for-multi","title":"Poseidon: A ViT-based Architecture for Multi-Frame Pose Estimation with Adaptive Frame Weighting and Multi-Scale Feature Fusion","date":"2025-01-14","arxiv_id":"2501.08446","repositories_listed":1,"syntology":null},{"url":"/paper/probpose-a-probabilistic-approach-to-2d-human","title":"ProbPose: A Probabilistic Approach to 2D Human Pose Estimation","date":"2024-12-03","arxiv_id":"2412.02254","repositories_listed":1,"syntology":null}],"syntology_records":16,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-25T09:33:49+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}