Browse State-of-the-Art › 3D Hand Pose Estimation
3D Hand Pose Estimation
87 papers with code · 7 benchmarks · 20 datasets archive 2025-07-28
Image: Zimmerman et l
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task, 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| FreiHAND (33 rows) | ExtPose | ExtPose: Robust and Coherent Pose Estimation by Extending ViTs | — | — | Compare |
| HO-3D v2 (24 rows) | ExtPose (T=16) | ExtPose: Robust and Coherent Pose Estimation by Extending ViTs | — | — | Compare |
| H3WB (15 rows) | SemGAN | 3D WholeBody Pose Estimation based on Semantic Graph Attention... | — | — | Compare |
| DexYCB (11 rows) | HOISDF | HOISDF: Constraining 3D Hand-Object Pose Estimation with Global... | code | Syntology ran 16 of 23 samples · 7 unverified | Compare |
| HInt: Hand Interactions in the wild (10 rows) | ExtPose* | ExtPose: Robust and Coherent Pose Estimation by Extending ViTs | — | — | Compare |
| HO-3D v3 (8 rows) | Hamba | Hamba: Single-view 3D Hand Reconstruction with Graph-guided... | code | Syntology ran 7 of 11 samples · 4 unverified | Compare |
| InterHand2.6M (1 row) | Epipolar Transformers | Epipolar Transformers | code | Syntology ran 0 of 5 samples · 5 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
20 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
3 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 87 papers with code (178 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
18 Dec 2017 10 repositories listedThe main objective is to minimize the reprojection loss of keypoints, which allow our model to be trained using images in-the-wild that only have ground truth 2D annotations.
-
3 May 2017 8 repositories listedLow-cost consumer depth cameras and deep learning have enabled reasonable 3D hand pose estimation from single depth images.
-
24 Mar 2023 6 repositories listed Syntology ran 1 of 5 samples · 4 unverified · 5 pointer-only (licence)To this end, we introduce a novel token mixing operator, RepMixer, a building block of FastViT, that uses structural reparameterization to lower the memory access cost by removing skip-connections in the network.
-
20 Nov 2017 5 repositories listedTo overcome these weaknesses, we firstly cast the 3D hand and human pose estimation problem from a single depth map into a voxel-to-voxel prediction that uses a 3D voxelized grid and estimates the per-voxel likelihood…
-
4 Apr 2020 4 repositories listed Syntology ran 6 of 20 samples · 14 unverifiedWe introduce a simple and effective network architecture for monocular 3D hand pose estimation consisting of an image encoder followed by a mesh convolutional decoder that is trained through a direct 3D hand mesh…
-
2 Jul 2019 4 repositories listedThis dataset is currently made of 77, 558 frames, 68 sequences, 10 persons, and 10 objects.
-
28 Aug 2017 4 repositories listedDeepPrior is a simple approach based on Deep Learning that predicts the joint 3D locations of a hand given a depth map.
-
1 Apr 2021 3 repositories listedWe present a graph-convolution-reinforced transformer, named Mesh Graphormer, for 3D human pose and mesh reconstruction from a single image.
-
11 Apr 2019 3 repositories listedPrevious work has made significant progress towards reconstruction of hand poses and object shapes in isolation.
-
7 Mar 2024 2 repositories listed Syntology ran 10 of 16 samples · 6 unverified · 16 pointer-only (licence)These two stereo constraints are used in a complementary manner to generate pseudo-labels, allowing reliable adaptation.
-
12 Sep 2021 2 repositories listedIn contrast, data synthesis can easily ensure those diversities separately.
-
10 Jun 2021 2 repositories listedEncouraged by the success of contrastive learning on image classification tasks, we propose a new self-supervised method for the structured regression task of 3D hand pose estimation.
-
9 Apr 2021 2 repositories listedWe introduce DexYCB, a new dataset for capturing hand grasping of objects.
-
18 Feb 2021 2 repositories listed3D hand pose estimation and shape recovery are challenging tasks in computer vision.
-
21 Aug 2020 2 repositories listedTherefore, we firstly propose (1) a large-scale dataset, InterHand2.
-
20 Aug 2020 2 repositories listedMost of the recent deep learning-based 3D human pose and mesh estimation methods regress the pose and shape parameters of human mesh models, such as SMPL and MANO, from an input image.
-
8 May 2019 2 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedImage-based features are attached to the mesh vertices and the Graph-CNN is responsible to process them on the mesh structure, while the regression target for each vertex is its 3D location.
-
3 Mar 2019 2 repositories listedThis work addresses a novel and challenging problem of estimating the full 3D hand shape and pose from a single RGB image.
-
25 Feb 2019 2 repositories listedIn this paper, we present a HAnd Mesh Recovery (HAMR) framework to tackle the problem of reconstructing the full 3D mesh of a human hand from a single RGB image.
-
9 Feb 2019 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We present in this work the first end-to-end deep learning based method that predicts both 3D hand shape and pose from RGB images in the wild.
-
10 Jun 2025 1 repository listedEstimating the 3D hand articulation from a single color image is a continuously investigated problem with applications in Augmented Reality (AR), Virtual Reality (VR), Human-Computer Interaction (HCI), and robotics.
-
25 Mar 2025 1 repository listed Syntology ran 6 of 6 samples · 0 unverified · 6 pointer-only (licence)We demonstrate that synthetic hand data can achieve the same level of accuracy as real data when integrating our identified components, paving the path to use synthetic data alone for hand pose estimation.
-
21 Feb 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand characteristics, dubbed SimHand.
-
18 Sep 2024 1 repository listedIn recent years, 3D hand pose estimation methods have garnered significant attention due to their extensive applications in human-computer interaction, virtual reality, and robotics.
-
19 Aug 2024 1 repository listedThe 3D hand pose, together with information from object detection, is processed by a transformer-based action recognition network, resulting in an accuracy of 91.
-
30 Jul 2024 1 repository listedTo address this challenge, this paper proposes the Denoising Adaptive Graph Transformer, HandDAGT, for hand pose estimation.
-
AttentionHand: Text-driven Controllable Hand Image Generation for 3D Hand Reconstruction in the Wild25 Jul 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedTo overcome these issues, we propose AttentionHand, a novel method for text-driven controllable hand image generation.
-
12 Jul 2024 1 repository listed Syntology ran 7 of 11 samples · 4 unverified · 11 pointer-only (licence)Specifically, we design a Graph-guided State Space (GSS) block that learns the graph-structured relations and spatial sequences of joints and uses 88.
-
4 Apr 2024 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedExtracting keypoint locations from input hand frames, known as 3D hand pose estimation, is a critical task in various human-computer interaction applications.
-
27 Mar 2024 1 repository listed Syntology ran 10 of 10 samples · 0 unverifiedReconstructing 3D hand mesh robustly from a single image is very challenging, due to the lack of diversity in existing real-world datasets.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections