Browse State-of-the-Art › Monocular 3D Human Pose Estimation
Monocular 3D Human Pose Estimation
70 papers with code · 1 benchmark · 5 datasets archive 2025-07-28
This task targets at 3D human pose estimation with a single RGB camera.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Human3.6M (52 rows) | MotionBERT (Finetune) | MotionBERT: A Unified Perspective on Learning Human Motion Representations | code | Syntology ran 5 of 11 samples · 6 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 70 papers with code (101 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
1 Feb 2018 22 repositories listedIn this work, we establish dense correspondences between RGB image and a surface-based representation of the human body, a task we refer to as dense human pose estimation.
-
8 May 2017 14 repositories listed Syntology ran 5 of 7 samples · 2 unverifiedFollowing the success of deep convolutional networks, state-of-the-art methods for 3d human pose estimation have focused on deep end-to-end systems that predict 3d joint locations given raw image pixels.
-
1 Jan 2017 11 repositories listedWe propose a unified formulation for the problem of 3D human pose estimation from a single raw RGB image that reasons jointly about 2D joint estimation and 3D pose reconstruction to improve both tasks.
-
28 Nov 2018 10 repositories listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)We start with predicted 2D keypoints for unlabeled video, then estimate 3D poses and finally back-project to the input 2D keypoints.
-
18 Dec 2017 10 repositories listedThe main objective is to minimize the reprojection loss of keypoints, which allow our model to be trained using images in-the-wild that only have ground truth 2D annotations.
-
8 Apr 2017 6 repositories listedWe propose a weakly-supervised transfer learning method that uses mixed 2D and 3D labels in a unified deep neutral network that presents two-stage cascaded structure.
-
11 Dec 2019 5 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Human motion is fundamental to understanding behavior.
-
6 Apr 2019 5 repositories listed Syntology ran 2 of 8 samples · 6 unverifiedIn this paper, we study the problem of learning Graph Convolutional Networks (GCNs) for regression.
-
26 Jul 2019 4 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedAlthough significant improvement has been achieved recently in 3D human pose estimation, most of the previous methods only treat a single-person case.
-
1 Jul 2019 4 repositories listedThe first stage is a convolutional neural network (CNN) that estimates 2D and 3D pose features along with identity assignments for all visible joints of all individuals.
-
18 Mar 2021 3 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 8 pointer-only (licence)Transformer architectures have become the model of choice in natural language processing and are now being introduced into computer vision tasks such as image classification, object detection, and semantic segmentation.
-
12 Oct 2022 2 repositories listed Syntology ran 0 of 17 samples · 17 unverifiedThe state-of-the-art for monocular 3D human pose estimation in videos is dominated by the paradigm of 2D-to-3D pose uplifting.
-
20 Jun 2022 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We capture a new dataset called RICH for "Real scenes, Interaction, Contact and Humans."
-
16 Dec 2021 2 repositories listedFirst, ICON infers detailed clothed-human normals (front/back) conditioned on the SMPL(-X) normals.
-
23 May 2021 2 repositories listedHowever, recent models depend on supervised training with 3D pose ground truth data or known pose priors for their target domains.
-
8 May 2019 2 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedImage-based features are attached to the mesh vertices and the Graph-CNN is responsible to process them on the mesh structure, while the regression target for each vertex is its 3D location.
-
11 Apr 2019 2 repositories listed Syntology ran 1 of 11 samples · 10 unverifiedWe argue that 3D human pose estimation from a monocular input is an inverse problem where multiple feasible solutions can exist.
-
17 Aug 2018 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Direct prediction of 3D body pose and shape remains a challenge even for highly parameterized deep learning models.
-
10 Jan 2017 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)With a comprehensive set of experiments, we show how this data can be used to train discriminative models that produce results with an unprecedented level of detail: our models predict 31 segments and 91 landmark…
-
17 Jun 2025 1 repository listedExisting monocular 3D pose estimation methods primarily rely on joint positional features, while overlooking intrinsic directional and angular correlations within the skeleton.
-
2 Apr 2025 1 repository listedTo address these limitations, our work introduces a groundbreaking motion pre-training method based on contextualized representation learning.
-
3 Jan 2025 1 repository listedHowever, previous methods ignore the intricate dependence within the 2D pose sequence and learn single temporal correlation.
-
7 Aug 2024 1 repository listedWithin this bidirectional global-local spatio-temporal SSM block, we introduce a reordering strategy to enhance the local modeling capability of the SSM.
-
31 Mar 2024 1 repository listed Syntology ran 6 of 10 samples · 4 unverified · 10 pointer-only (licence)This paper presents a novel Kinematics and Trajectory Prior Knowledge-Enhanced Transformer (KTPFormer), which overcomes the weakness in existing transformer-based methods for 3D human pose estimation that the derivation…
-
7 Mar 2024 1 repository listedTo address these problems, a Disentangled Diffusion-based 3D Human Pose Estimation method with Hierarchical Spatial and Temporal Denoiser is proposed, termed DDHPose.
-
18 Jan 2024 1 repository listedHowever, there is still ample room for improvement as these methods often overlook the exploration of correlation between the 2D and 3D joint-level features.
-
19 Dec 2023 1 repository listedThe lifting of 3D structure and camera from 2D landmarks is at the cornerstone of the entire discipline of computer vision.
-
11 Dec 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)By constraining the outputs to lie on the human pose manifold, ManiPose guarantees the consistency of all hypothetical poses, in contrast to previous works.
-
25 Oct 2023 1 repository listed Syntology ran 14 of 16 samples · 2 unverifiedOur proposed GCNFormer module exploits the local relationship between adjacent joints, outputting a new representation that is complementary to the transformer output.
-
10 Aug 2023 1 repository listedNotably, our model achieves state-of-the-art performance on all action categories in the Human3.
Syntology lines on 15 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections