Browse State-of-the-Art › 2D Human Pose Estimation
2D Human Pose Estimation
68 papers with code · 10 benchmarks · 25 datasets archive 2025-07-28
What is Human Pose Estimation? Human pose estimation is the process of estimating the configuration of the body (pose) from a single, typically monocular, image. Background. Human pose estimation is one of the key problems in computer vision that has been studied for well over 15 years. The reason for its importance is the abundance of applications that can benefit from such a technology. For example, human pose estimation allows for higher-level reasoning in the context of human-computer interaction and activity recognition; it is also one of the basic building blocks for marker-less motion capture (MoCap) technology. MoCap technology is useful for applications ranging from character animation to clinical analysis of gait pathologies.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
10 leaderboard tables shown for this task, 10 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
25 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
6 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 68 papers with code (118 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
24 Nov 2016 61 repositories listed Syntology ran 4 of 23 samples · 19 unverified · 4 pointer-only (licence)We present an approach to efficiently detect the 2D pose of multiple people in an image.
-
25 Feb 2019 39 repositories listed Syntology ran 8 of 25 samples · 17 unverifiedWe start from a high-resolution subnetwork as the first stage, gradually add high-to-low resolution subnetworks one by one to form more stages, and connect the mutli-resolution subnetworks in parallel.
-
17 Apr 2018 27 repositories listedThere has been significant progress on pose estimation and increasing interests on pose tracking in recent years.
-
27 Aug 2019 19 repositories listed Syntology ran 5 of 27 samples · 22 unverifiedHigherHRNet even surpasses all top-down methods on CrowdPose test (67.
-
1 Dec 2016 14 repositories listedIn this paper, we propose a novel regional multi-person pose estimation (RMPE) framework to facilitate pose estimation in the presence of inaccurate human bounding boxes.
-
24 Nov 2019 8 repositories listedWe rethink a well-know bottom-up approach for multi-person pose estimation and propose an improved one.
-
17 Jun 2020 7 repositories listedWe present BlazePose, a lightweight convolutional neural network architecture for human pose estimation that is tailored for real-time inference on mobile devices.
-
28 Mar 2018 7 repositories listed Syntology ran 0 of 16 samples · 16 unverifiedWe demonstrate that our pose-based framework can achieve better accuracy than the state-of-art detection-based approach on the human instance segmentation problem, and can moreover better handle occlusion.
-
26 Apr 2022 6 repositories listed Syntology ran 18 of 31 samples · 13 unverified · 6 pointer-only (licence)In this paper, we show the surprisingly good capabilities of plain vision transformers for pose estimation from various aspects, namely simplicity in model structure, scalability in model size, flexibility in training…
-
2 Oct 2018 6 repositories listedWe study the problem of representation learning in goal-conditioned hierarchical reinforcement learning.
-
16 Nov 2016 5 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedWe introduce associative embedding, a novel method for supervising convolutional neural networks for the task of detection and grouping.
-
3 Feb 2023 3 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)This paper presents a novel end-to-end framework with Explicit box Detection for multi-person Pose estimation, called ED-Pose, where it unifies the contextual learning between human-level (global) and keypoint-level…
-
22 Aug 2024 2 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedWe present Sapiens, a family of models for four fundamental human-centric vision tasks -- 2D pose estimation, body-part segmentation, depth estimation, and surface normal prediction.
-
12 Oct 2023 2 repositories listed Syntology ran 11 of 13 samples · 2 unverified · 13 pointer-only (licence)This work aims to address an advanced keypoint detection problem: how to accurately detect any keypoints in complex real-world scenarios, which involves massive, messy, and open-ended objects as well as their associated…
-
13 Jul 2023 2 repositories listedMethods and datasets for human pose estimation focus predominantly on side- and front-view scenarios.
-
7 Dec 2022 2 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedIn this paper, we show the surprisingly good properties of plain vision transformers for body pose estimation from various aspects, namely simplicity in model structure, scalability in model size, flexibility in…
-
16 Sep 2022 2 repositories listedIn this paper, we propose the token-Pruned Pose Transformer (PPT) for 2D human pose estimation, which can locate a rough human mask and performs self-attention only within selected tokens.
-
14 Jul 2022 2 repositories listed Syntology ran 1 of 3 samples · 2 unverified · 2 pointer-only (licence)We present XMem, a video object segmentation architecture for long videos with unified feature memory stores inspired by the Atkinson-Shiffrin memory model.
-
27 Dec 2021 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWith a simple yet effective motion-aware fully-connected network, SmoothNet improves the temporal smoothness of existing pose estimators significantly and enhances the estimation accuracy of those challenging frames as…
-
2 Dec 2021 2 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedVideo data is often repetitive; for example, the contents of adjacent frames are usually strongly correlated.
-
7 May 2021 2 repositories listedThis work leverages novel spatial-temporal graph convolutional network (ST-GCN) architectures and training procedures to predict clinical scores of parkinsonism in gait from video of individuals with dementia.
-
22 Oct 2020 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedThe applicability of agglomerative clustering, for inferring both hierarchical and flat clustering, is limited by its scalability.
-
23 Jul 2020 2 repositories listedThis paper investigates the task of 2D human whole-body pose estimation, which aims to localize dense landmarks on the entire human body including face, hands, body, and feet.
-
16 Jul 2020 2 repositories listedThe largest model is able to come within 4.
-
15 Mar 2019 2 repositories listedWe propose a new bottom-up method for multi-person 2D human pose estimation that is particularly well suited for urban mobility such as self-driving cars and delivery robots.
-
5 Jan 2017 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)In this work we present SURREAL (Synthetic hUmans foR REAL tasks): a new large-scale dataset with synthetically-generated but realistic images of people rendered from 3D sequences of human motion capture data.
-
14 Dec 2015 2 repositories listed Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)When stochastic gradient is used, it can naturally damp the gradient noise to stabilize SGD.
-
1 Jun 2014 2 repositories listedHuman pose estimation has made significant progress during the last years.
-
14 Jan 2025 1 repository listedHuman pose estimation, a vital task in computer vision, involves detecting and localising human joints in images and videos.
-
3 Dec 2024 1 repository listedCurrent Human Pose Estimation methods have achieved significant improvements.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections