Papers › Kinematic-Structure-Preserved Representation for Unsupervised 3D Human Pose Estimation

Kinematic-Structure-Preserved Representation for Unsupervised 3D Human Pose Estimation

24 Jun 2020arXiv:2006.14107archive 2025-07-28

Jogendra Nath Kundu, Siddharth Seth, Rahul M. V, Mugalodi Rakesh, R. Venkatesh Babu, Anirban Chakraborty

Estimation of 3D human pose from monocular image has gained considerable attention, as a key step to several human-centric applications. However, generalizability of human pose estimation models developed using supervision on large-scale in-studio datasets remains questionable, as these models often perform unsatisfactorily on unseen in-the-wild environments. Though weakly-supervised models have been proposed to address this shortcoming, performance of such models relies on availability of paired supervision on some related tasks, such as 2D pose or multi-view image pairs. In contrast, we propose a novel kinematic-structure-preserved unsupervised 3D pose estimation framework, which is not restrained by any paired or unpaired weak supervisions. Our pose estimation framework relies on a minimal set of prior knowledge that defines the underlying kinematic 3D structure, such as skeletal joint connectivity information with bone-length ratios in a fixed canonical scale. The proposed model employs three consecutive differentiable transformations named as forward-kinematics, camera-projection and spatial-map transformation. This design not only acts as a suitable bottleneck stimulating effective pose disentanglement but also yields interpretable latent pose representations avoiding training of an explicit latent embedding to pose mapper. Furthermore, devoid of unstable adversarial setup, we re-utilize the decoder to formalize an energy-based loss, which enables us to learn from in-the-wild videos, beyond laboratory settings. Comprehensive experiments demonstrate our state-of-the-art unsupervised and weakly-supervised pose estimation performance on both Human3.6M and MPI-INF-3DHP datasets. Qualitative results on unseen environments further establish our superior generalization ability.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Human Pose Estimation3D Pose EstimationDisentanglementPose EstimationUnsupervised 3D Human Pose Estimation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Unsupervised 3D Human Pose Estimation Human3.6M Kinematic-Structure-Preserved Representation PA-MPJPE 89.4 #12 of 12 Archive leaderboard report
Unsupervised 3D Human Pose Estimation MPI-INF-3DHP Kinematic-Structure-Preserved Representation AUC 43.4 #2 of 4 Archive leaderboard report
Unsupervised 3D Human Pose Estimation MPI-INF-3DHP Kinematic-Structure-Preserved Representation MPJPE 99.2 #2 of 4 Archive leaderboard report
Unsupervised 3D Human Pose Estimation MPI-INF-3DHP Kinematic-Structure-Preserved Representation PCK 79.2 #2 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections