{"url":"/sota/3d-human-pose-estimation-on-humaneva-i","task":{"name":"3D Human Pose Estimation","url":"/task/3d-human-pose-estimation","note":null},"dataset":{"name":"HumanEva-I","url":null},"category":"Computer Vision","categories":["Computer Vision"],"category_note":null,"description":"**3D Human Pose Estimation** is a computer vision task that involves estimating the 3D positions and orientations of body joints and bones from 2D images or videos. The goal is to reconstruct the 3D pose of a person in real-time, which can be used in a variety of applications, such as virtual reality, human-computer interaction, and motion analysis.","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["Mean Reconstruction Error (mm)"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"Mean Reconstruction Error (mm)":"lower"}},"counts":{"rows":31,"rows_with_code":16,"rows_with_paper_page":31,"rows_dated":31,"rows_using_additional_data":0},"rows":[{"rank_in_archive_order":1,"model":"GLA-GCN (T=27, GT)","metrics":{"Mean Reconstruction Error (mm)":"9.2"},"uses_additional_data":false,"paper_date":"2023-07-12","paper":"/paper/gla-gcn-global-local-adaptive-graph","paper_url":"https://arxiv.org/abs/2307.05853v2","paper_title":"GLA-GCN: Global-local Adaptive Graph Convolutional Network for 3D Human Pose Estimation from Monocular Video","code":"https://github.com/bruceyo/GLA-GCN","n_code_links":1,"syntology":null},{"rank_in_archive_order":2,"model":"StridedTransformer (T=27 GT)","metrics":{"Mean Reconstruction Error (mm)":"12.2"},"uses_additional_data":false,"paper_date":"2021-03-26","paper":"/paper/lifting-transformer-for-3d-human-pose","paper_url":"https://arxiv.org/abs/2103.14304v8","paper_title":"Exploiting Temporal Contexts with Strided Transformer for 3D Human Pose Estimation","code":"https://github.com/Vegetebird/StridedTransformer-Pose3D","n_code_links":1,"syntology":{"n_ran":3,"n_unverified":7,"n_samples":10,"n_pointer_only_licence":0}},{"rank_in_archive_order":3,"model":"Spatio-Temporal Network (T=128)","metrics":{"Mean Reconstruction Error (mm)":"13.5"},"uses_additional_data":false,"paper_date":"2020-04-07","paper":"/paper/3d-human-pose-estimation-using-spatio-1","paper_url":"https://arxiv.org/abs/2004.11822v1","paper_title":"3D Human Pose Estimation using Spatio-Temporal Networks with Explicit Occlusion Training","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":4,"model":"Occlusion-Aware Networks","metrics":{"Mean Reconstruction Error (mm)":"14.3"},"uses_additional_data":false,"paper_date":"2019-10-01","paper":"/paper/occlusion-aware-networks-for-3d-human-pose","paper_url":"http://openaccess.thecvf.com/content_ICCV_2019/html/Cheng_Occlusion-Aware_Networks_for_3D_Human_Pose_Estimation_in_Video_ICCV_2019_paper.html","paper_title":"Occlusion-Aware Networks for 3D Human Pose Estimation in Video","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":5,"model":"HEMlets Pose","metrics":{"Mean Reconstruction Error (mm)":"15.2"},"uses_additional_data":false,"paper_date":"2019-10-26","paper":"/paper/hemlets-pose-learning-part-centric-heatmap-1","paper_url":"https://arxiv.org/abs/1910.12032v1","paper_title":"HEMlets Pose: Learning Part-Centric Heatmap Triplets for Accurate 3D Human Pose Estimation","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":6,"model":"Attention (T=27 MA)","metrics":{"Mean Reconstruction Error (mm)":"15.4"},"uses_additional_data":false,"paper_date":"2021-03-04","paper":"/paper/enhanced-3d-human-pose-estimation-from-videos","paper_url":"https://arxiv.org/abs/2103.03170v1","paper_title":"Enhanced 3D Human Pose Estimation from Videos by using Attention-Based Neural Network with Dilated Convolutions","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":7,"model":"MixSTE (T=43, FT)","metrics":{"Mean Reconstruction Error (mm)":"16.1"},"uses_additional_data":false,"paper_date":"2022-03-02","paper":"/paper/mixste-seq2seq-mixed-spatio-temporal-encoder","paper_url":"https://arxiv.org/abs/2203.00859v4","paper_title":"MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video","code":"https://github.com/JinluZhang1126/MixSTE","n_code_links":1,"syntology":{"n_ran":4,"n_unverified":0,"n_samples":4,"n_pointer_only_licence":4}},{"rank_in_archive_order":8,"model":"Ordinal Depth Supervision","metrics":{"Mean Reconstruction Error (mm)":"18.3"},"uses_additional_data":false,"paper_date":"2018-05-10","paper":"/paper/ordinal-depth-supervision-for-3d-human-pose","paper_url":"http://arxiv.org/abs/1805.04095v1","paper_title":"Ordinal Depth Supervision for 3D Human Pose Estimation","code":"https://github.com/geopavlakos/ordinal-pose3d","n_code_links":1,"syntology":null},{"rank_in_archive_order":9,"model":"StridedTransformer (T=27 MRCNN)","metrics":{"Mean Reconstruction Error (mm)":"18.9"},"uses_additional_data":false,"paper_date":"2021-03-26","paper":"/paper/lifting-transformer-for-3d-human-pose","paper_url":"https://arxiv.org/abs/2103.14304v8","paper_title":"Exploiting Temporal Contexts with Strided Transformer for 3D Human Pose Estimation","code":"https://github.com/Vegetebird/StridedTransformer-Pose3D","n_code_links":1,"syntology":{"n_ran":3,"n_unverified":7,"n_samples":10,"n_pointer_only_licence":0}},{"rank_in_archive_order":10,"model":"RTPCA","metrics":{"Mean Reconstruction Error (mm)":"19.1"},"uses_additional_data":false,"paper_date":"2023-09-04","paper":"/paper/refined-temporal-pyramidal-compression-and","paper_url":"https://arxiv.org/abs/2309.01365v3","paper_title":"Refined Temporal Pyramidal Compression-and-Amplification Transformer for 3D Human Pose Estimation","code":"https://github.com/hbing-l/rtpca","n_code_links":1,"syntology":null},{"rank_in_archive_order":11,"model":"DG-Net (T=4)","metrics":{"Mean Reconstruction Error (mm)":"19.5"},"uses_additional_data":false,"paper_date":"2021-09-15","paper":"/paper/learning-dynamical-human-joint-affinity-for","paper_url":"https://arxiv.org/abs/2109.07353v1","paper_title":"Learning Dynamical Human-Joint Affinity for 3D Pose Estimation in Videos","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":12,"model":"GAST","metrics":{"Mean Reconstruction Error (mm)":"21.2"},"uses_additional_data":false,"paper_date":"2020-03-11","paper":"/paper/gast-net-graph-attention-spatio-temporal","paper_url":"https://arxiv.org/abs/2003.14179v4","paper_title":"A Graph Attention Spatio-temporal Convolutional Network for 3D Human Pose Estimation in Video","code":"https://github.com/fabro66/GAST-Net-3DPoseEstimation","n_code_links":1,"syntology":null},{"rank_in_archive_order":13,"model":"PoseFormer","metrics":{"Mean Reconstruction Error (mm)":"21.6"},"uses_additional_data":false,"paper_date":"2021-03-18","paper":"/paper/3d-human-pose-estimation-with-spatial-and","paper_url":"https://arxiv.org/abs/2103.10455v3","paper_title":"3D Human Pose Estimation with Spatial and Temporal Transformers","code":"https://github.com/zczcwh/PoseFormer","n_code_links":3,"syntology":{"n_ran":4,"n_unverified":4,"n_samples":8,"n_pointer_only_licence":8}},{"rank_in_archive_order":14,"model":"Sequence-to-sequence network","metrics":{"Mean Reconstruction Error (mm)":"22"},"uses_additional_data":false,"paper_date":"2017-11-23","paper":"/paper/exploiting-temporal-information-for-3d-pose","paper_url":"http://arxiv.org/abs/1711.08585v4","paper_title":"Exploiting temporal information for 3D pose estimation","code":"https://github.com/rayat137/Pose_3D","n_code_links":1,"syntology":{"n_ran":2,"n_unverified":0,"n_samples":2,"n_pointer_only_licence":2}},{"rank_in_archive_order":15,"model":"3D Pose Grammar Network","metrics":{"Mean Reconstruction Error (mm)":"22.9"},"uses_additional_data":false,"paper_date":"2019-06-01","paper":"/paper/learning-pose-grammar-for-monocular-3d-pose","paper_url":"http://www.stat.ucla.edu/~jxie/personalpage_file/publications/3dpose_pami19.pdf","paper_title":"Learning Pose Grammar for Monocular 3D Pose Estimation","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":16,"model":"Pose Grammar","metrics":{"Mean Reconstruction Error (mm)":"22.9"},"uses_additional_data":false,"paper_date":"2017-10-17","paper":"/paper/learning-pose-grammar-to-encode-human-body","paper_url":"http://arxiv.org/abs/1710.06513v6","paper_title":"Learning Pose Grammar to Encode Human Body Configuration for 3D Pose Estimation","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":17,"model":"Ours (Oracle)","metrics":{"Mean Reconstruction Error (mm)":"23.9"},"uses_additional_data":false,"paper_date":"2019-04-02","paper":"/paper/monocular-3d-human-pose-estimation-by-1","paper_url":"https://arxiv.org/abs/1904.01324v2","paper_title":"Monocular 3D Human Pose Estimation by Generation and Ordinal Ranking","code":"https://github.com/ssfootball04/generative_pose","n_code_links":1,"syntology":{"n_ran":0,"n_unverified":3,"n_samples":3,"n_pointer_only_licence":0}},{"rank_in_archive_order":18,"model":"c2f-vol","metrics":{"Mean Reconstruction Error (mm)":"24.3"},"uses_additional_data":false,"paper_date":"2016-11-23","paper":"/paper/coarse-to-fine-volumetric-prediction-for","paper_url":"http://arxiv.org/abs/1611.07828v2","paper_title":"Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose","code":"https://github.com/geopavlakos/c2f-vol-train","n_code_links":4,"syntology":null},{"rank_in_archive_order":19,"model":"ConvFormer (T=43)","metrics":{"Mean Reconstruction Error (mm)":"24.3"},"uses_additional_data":false,"paper_date":"2023-04-04","paper":"/paper/convformer-parameter-reduction-in-transformer","paper_url":"https://arxiv.org/abs/2304.02147v1","paper_title":"ConvFormer: Parameter Reduction in Transformer Models for 3D Human Pose Estimation by Leveraging Dynamic Multi-Headed Convolutional Attention","code":"https://github.com/ajda1992/convformer","n_code_links":1,"syntology":null},{"rank_in_archive_order":20,"model":"SIM (SH detections)","metrics":{"Mean Reconstruction Error (mm)":"24.6"},"uses_additional_data":false,"paper_date":"2017-05-08","paper":"/paper/a-simple-yet-effective-baseline-for-3d-human","paper_url":"http://arxiv.org/abs/1705.03098v2","paper_title":"A simple yet effective baseline for 3d human pose estimation","code":"https://github.com/open-mmlab/mmpose","n_code_links":14,"syntology":{"n_ran":5,"n_unverified":2,"n_samples":7,"n_pointer_only_licence":0}},{"rank_in_archive_order":21,"model":"EDM","metrics":{"Mean Reconstruction Error (mm)":"26.9"},"uses_additional_data":false,"paper_date":"2016-11-28","paper":"/paper/3d-human-pose-estimation-from-a-single-image","paper_url":"http://arxiv.org/abs/1611.09010v1","paper_title":"3D Human Pose Estimation from a Single Image via Distance Matrix Regression","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":22,"model":"Recurrent 3D Pose Sequence Machines","metrics":{"Mean Reconstruction Error (mm)":"30.8"},"uses_additional_data":false,"paper_date":"2017-07-31","paper":"/paper/recurrent-3d-pose-sequence-machines","paper_url":"http://arxiv.org/abs/1707.09695v1","paper_title":"Recurrent 3D Pose Sequence Machines","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":23,"model":"DMHSR(J,B,D)","metrics":{"Mean Reconstruction Error (mm)":"33.7"},"uses_additional_data":false,"paper_date":"2017-01-31","paper":"/paper/deep-multitask-architecture-for-integrated-2d","paper_url":"http://arxiv.org/abs/1701.08985v1","paper_title":"Deep Multitask Architecture for Integrated 2D and 3D Human Sensing","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":24,"model":"Dual-source approach","metrics":{"Mean Reconstruction Error (mm)":"38.9"},"uses_additional_data":false,"paper_date":"2015-09-22","paper":"/paper/a-dual-source-approach-for-3d-pose-estimation","paper_url":"http://arxiv.org/abs/1509.06720v2","paper_title":"A Dual-Source Approach for 3D Pose Estimation from a Single Image","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":25,"model":"TGP","metrics":{"Mean Reconstruction Error (mm)":"39.1"},"uses_additional_data":false,"paper_date":"2010-03-01","paper":"/paper/twin-gaussian-processes-for-structured","paper_url":"https://doi.org/10.1007/s11263-008-0204-y","paper_title":"Twin gaussian processes for structured prediction","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":26,"model":"DSRF","metrics":{"Mean Reconstruction Error (mm)":"40.3"},"uses_additional_data":false,"paper_date":"2014-09-01","paper":"/paper/depth-sweep-regression-forests-for-estimating","paper_url":"http://dx.doi.org/10.5244/C.28.80","paper_title":"Depth sweep regression forests for estimating 3d human pose from images","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":27,"model":"Simo-Serra et al.","metrics":{"Mean Reconstruction Error (mm)":"56.7"},"uses_additional_data":false,"paper_date":"2013-06-01","paper":"/paper/a-joint-model-for-2d-and-3d-pose-estimation","paper_url":"http://openaccess.thecvf.com/content_cvpr_2013/html/Simo-Serra_A_Joint_Model_2013_CVPR_paper.html","paper_title":"A Joint Model for 2D and 3D Pose Estimation from a Single Image","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":28,"model":"Ours","metrics":{"Mean Reconstruction Error (mm)":"64"},"uses_additional_data":false,"paper_date":"2018-08-17","paper":"/paper/neural-body-fitting-unifying-deep-learning","paper_url":"http://arxiv.org/abs/1808.05942v1","paper_title":"Neural Body Fitting: Unifying Deep Learning and Model-Based Human Pose and Shape Estimation","code":"https://github.com/andrewjong/SwapNet","n_code_links":2,"syntology":{"n_ran":1,"n_unverified":0,"n_samples":1,"n_pointer_only_licence":1}},{"rank_in_archive_order":29,"model":"Wang et al.","metrics":{"Mean Reconstruction Error (mm)":"71.3"},"uses_additional_data":false,"paper_date":"2014-06-09","paper":"/paper/robust-estimation-of-3d-human-poses-from-a","paper_url":"http://arxiv.org/abs/1406.2282v1","paper_title":"Robust Estimation of 3D Human Poses from a Single Image","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":30,"model":"SMPLify (dense)","metrics":{"Mean Reconstruction Error (mm)":"74.5"},"uses_additional_data":false,"paper_date":"2017-01-10","paper":"/paper/unite-the-people-closing-the-loop-between-3d","paper_url":"http://arxiv.org/abs/1701.02468v3","paper_title":"Unite the People: Closing the Loop Between 3D and 2D Human Representations","code":"https://github.com/MandyMo/pytorch_HMR","n_code_links":2,"syntology":{"n_ran":1,"n_unverified":0,"n_samples":1,"n_pointer_only_licence":1}},{"rank_in_archive_order":31,"model":"SMPLify","metrics":{"Mean Reconstruction Error (mm)":"79.9"},"uses_additional_data":false,"paper_date":"2016-07-27","paper":"/paper/keep-it-smpl-automatic-estimation-of-3d-human","paper_url":"http://arxiv.org/abs/1607.08128v1","paper_title":"Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image","code":"https://github.com/Jtoo/fitting_human_smpl_model","n_code_links":2,"syntology":null}],"since_archive":{"claim":"Results that newer papers report for their own method, placed here by Syntology. A model pointed at the cell in the paper's own table; the number was read from that cell and checked against this leaderboard's metric, dataset, split and scale; an independent check that saw this leaderboard's other rows and every other leaderboard on the same dataset accepted it. Not reviewed by the paper's authors or by the archive's editors, and not ranked against the archive rows.","extraction_file_present":true,"measurement":{"test_papers":883,"papers_with_output":881,"judged_true":108,"judged":110,"wilson95_lower":0.9361,"measured_on":"2026-09-24","frozen_commit":"0e3de0df94"},"measurement_note":"blind adjudication of accepted entries on a held-out split of archive papers, rules frozen before the test","coverage":{"sentence":"Syntology has checked 6,264 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","complete":false,"papers_newer_than_archive":9581,"papers_checked":6264,"papers_extracted_not_yet_verified":0,"boards_without_verdict":2,"papers_not_yet_extracted":3316},"order":"newest first by month (arXiv date, else the arXiv-id month), then arXiv id descending","columns":[],"entries":[]},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":9,"rows_with_any_sample_ran":8,"distinct_papers_with_graph_line":8,"distinct_papers_with_any_sample_ran":7,"samples_over_distinct_papers":{"n_ran":20,"n_unverified":16,"n_samples":36,"n_pointer_only_licence":16,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":23,"n_unverified":23,"n_samples":46,"n_pointer_only_licence":16,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}