Browse State-of-the-Art › Pose Estimation

Pose Estimation

1,679 papers with code · 31 benchmarks · 124 datasets archive 2025-07-28

Computer Vision

Pose Estimation is a computer vision task where the goal is to detect the position and orientation of a person or an object. Usually, this is done by predicting the location of specific keypoints like hands, head, elbows, etc. in case of Human Pose Estimation.

A common benchmark for this task is MPII Human Pose

( Image credit: Real-time 2D Multi-Person Pose Estimation on CPU: Lightweight OpenPose )

Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.

Benchmarks archive 2025-07-28

31 leaderboard tables shown for this task, 31 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 31 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
COCO test-dev (47 rows) ViTPose (ViTAE-G, ensemble) ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation code Syntology ran 18 of 31 samples · 13 unverified Compare
MPII Human Pose (46 rows) PCT (swin-l, test set) Human Pose as Compositional Tokens code Syntology ran 2 of 2 samples · 0 unverified Compare
OCHuman (19 rows) ViTPose (ViTAE-G, GT bounding boxes) ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation code Syntology ran 18 of 31 samples · 13 unverified Compare
Leeds Sports Poses (18 rows) OmniPose OmniPose: A Multi-Scale Framework for Multi-Person Pose Estimation code — Compare
CrowdPose (12 rows) BUCTD-W48 (w/cond. input from PETR, and generative sampling) Rethinking pose estimation in crowds: overcoming the detection... code Syntology ran 4 of 6 samples · 2 unverified Compare
COCO val2017 (11 rows) CCNet (ViTPose-B_GT-bbox_256x192) On the Calibration of Human Pose Estimation — — Compare
AIC (10 rows) Hulk(Finetune, ViT-L) Hulk: A Universal Knowledge Translator for Human-Centric Tasks code Syntology ran 12 of 24 samples · 12 unverified Compare
COCO (Common Objects in Context) (10 rows) OmniPose (WASPv2) OmniPose: A Multi-Scale Framework for Multi-Person Pose Estimation code — Compare
ITOP front-view (7 rows) AdaPose Sequential 3D Human Pose Estimation Using Adaptive Point Cloud... code — Compare
InLoc (6 rows) GIM-DKM GIM: Learning Generalizable Image Matcher From Internet Videos code Syntology ran 12 of 17 samples · 5 unverified Compare
ITOP top-view (5 rows) DECA-D3 DECA: Deep viewpoint-Equivariant human pose estimation using... code — Compare
J-HMDB (5 rows) SimpleBaseline + HANet Kinematic-aware Hierarchical Attention Network for Human Pose... code — Compare
MPII Single Person (5 rows) 4xRSN-50 Learning Delicate Local Representations for Multi-Person Pose Estimation code Syntology ran 2 of 3 samples · 1 unverified Compare
UPenn Action (5 rows) OmniPose OmniPose: A Multi-Scale Framework for Multi-Person Pose Estimation code — Compare
DensePose-COCO (4 rows) DP-RCNN-DeepLab (ResNet-101) Continuous Surface Embeddings code Syntology ran 5 of 6 samples · 1 unverified Compare
SALSA (4 rows) SubdivNet Subdivision-Based Mesh Convolution Networks code Syntology ran 0 of 5 samples · 5 unverified Compare
300W (Full) (3 rows) SPIGA Shape Preserving Facial Landmarks with Graph Attention Networks code — Compare
BRACE (2 rows) HRNet fine-tuned on BRACE Deep High-Resolution Representation Learning for Human Pose Estimation code Syntology ran 8 of 25 samples · 17 unverified Compare
COCO 2017 val (2 rows) LOGO-CAP (Ours) HRNet-W48 — — — Compare
FLIC Elbows (2 rows) Stacked Hourglass Networks Stacked Hourglass Networks for Human Pose Estimation code Syntology ran 2 of 25 samples · 23 unverified Compare
FLIC Wrists (2 rows) Stacked Hourglass Networks Stacked Hourglass Networks for Human Pose Estimation code Syntology ran 2 of 25 samples · 23 unverified Compare
UAV-Human (2 rows) AlphaPose RMPE: Regional Multi-person Pose Estimation code — Compare
!(()&&!|*|*| (1 row) Nate Neural Ordinary Differential Equations code Syntology ran 89 of 124 samples · 35 unverified Compare
3DPW (1 row) HybridCap HybridCap: Inertia-aid Monocular Capture of Challenging Human Motions — — Compare
ApolloCar3D (1 row) GSNet GSNet: Joint Vehicle Pose and Shape Reconstruction with... code — Compare
COCO minival (1 row) MSPN Rethinking on Multi-Stage Networks for Human Pose Estimation code — Compare
KITTI 2015 (1 row) GeoNet GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose code Syntology ran 0 of 8 samples · 8 unverified Compare
MERL-RAV (1 row) SPIGA Shape Preserving Facial Landmarks with Graph Attention Networks code — Compare
MPII (1 row) OmniPose (WASPv2) OmniPose: A Multi-Scale Framework for Multi-Person Pose Estimation code — Compare
MS-COCO (1 row) UniHCP (finetune) UniHCP: A Unified Model for Human-Centric Perceptions code Syntology ran 7 of 13 samples · 6 unverified Compare
Pix3D (1 row) Mid-Level based Object Pose Estimation using Mid-level Visual Representations code — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

124 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 124 until expanded.

COCO (Common Objects in Context)KITTIMPII3DPWAMASS10,000 People - Human Pose Recognition DataDensePoseJHMDBPASCAL3D+300WLSPYCB-VideoMPII Human PosePix3DLFPWSUN3DNCLTPenn ActionCrowdPoseAachen Day-NightInLocOCHumanTotalCaptureInterHand2.6MUAV-HumanFLICGazeFollowJTAMuCo-3DHPCOCO-WholeBodyMannequinChallengeAnimal KingdomIKEA ASMINRIA PersonKeypointNetITOPAnimal-Pose DatasetModaNetSALSAApolloCar3D3D Hand PoseFirst-Person Hand Action BenchmarkMSRA HandSVIROAICMVOREgoDexterMoViSLOPER4DUnrealEgoUnite the PeoplexR-EgoPoseCarFusionEgoCapHUMAN4DATRWGPAH3WBIMUPoserBiwi Kinect Head PoseHandNetKITTI Odometry BenchmarkMonoPerfCap DatasetSynthHandsFATYCBInEOAT DatasetYoga-82BRACEDREAM-datasetPedXFewSOLFraunhofer IPA Bin-PickingHuPRMERL-RAVSportsPoseAmateur DrawingsCHAIRS datasetComposable activities datasetCORSMALFitness-AQAHOPE-ImageHOPE-VideoICVL Hand PosturePoPArtVBRBASEPRODDrunkard's DatasetMBW - Zoo DatasetMOTFrontMPHOI-72Parkinson's Pose Estimation DatasetPoserRendered Handpose DatasetRetinal MicrosurgerySLAM2REFUBC3V DatasetVRMocap: VR Mocap Dataset for Pose ReconstructionBigHand2.2M BenchmarkCIPConSLAMCOPE-119Data for "Image-based Backbone Reconstruction for Non-Slender Soft Robots"DensePose-TrackDesert LocustFine-grained 3D PoseHalpe-FullBodyHumanoidRobotPoseImmediacy DatasetiMoCapMacaquePoseNToPOmniLabRealArt-6SIDODSMOTStore datasetStreet View Image, Pose, and 3D Cities DatasetSymmetric SolidsVinegar FlyVR Mocap Dataset for Pose/Orientation PredictionGrévy’s ZebraICT-3DHPInfiniteRepPose Estimation Lunar Robot

Subtasks archive 2025-07-28

17 subtasks in the archive's task tree.

Parent tasks archive 2025-07-28

Most implemented papers archive 2025-07-28

30 shown of 1,679 papers with code (4,228 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

Syntology lines on 22 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections