{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-3d-pose-estimation-for","title":"Unsupervised 3D Pose Estimation for Hierarchical Dance Video Recognition","arxiv_id":"2109.09166","date":"2021-09-19","proceeding":"ICCV 2021 10","authors":["Xiaodan Hu","Narendra Ahuja"],"abstract":"Dance experts often view dance as a hierarchy of information, spanning low-level (raw images, image sequences), mid-levels (human poses and bodypart movements), and high-level (dance genre). We propose a Hierarchical Dance Video Recognition framework (HDVR). HDVR estimates 2D pose sequences, tracks dancers, and then simultaneously estimates corresponding 3D poses and 3D-to-2D imaging parameters, without requiring ground truth for 3D poses. Unlike most methods that work on a single person, our tracking works on multiple dancers, under occlusions. From the estimated 3D pose sequence, HDVR extracts body part movements, and therefrom dance genre. The resulting hierarchical dance representation is explainable to experts. To overcome noise and interframe correspondence ambiguities, we enforce spatial and temporal motion smoothness and photometric continuity over time. We use an LSTM network to extract 3D movement subsequences from which we recognize the dance genre. For experiments, we have identified 154 movement types, of 16 body parts, and assembled a new University of Illinois Dance (UID) Dataset, containing 1143 video clips of 9 genres covering 30 hours, annotated with movement and genre labels. Our experimental results demonstrate that our algorithms outperform the state-of-the-art 3D pose estimation methods, which also enhances our dance recognition performance.","url_abs":"https://arxiv.org/abs/2109.09166v1","url_pdf":"https://arxiv.org/pdf/2109.09166v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-3d-pose-estimation-for","repo_url":"https://github.com/garfield-kh/posetriplet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"3d-pose-estimation","task_name":"3D Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"unsupervised-3d-human-pose-estimation","task_name":"Unsupervised 3D Human Pose Estimation"},{"task_slug":"video-recognition","task_name":"Video Recognition"},{"task_slug":"weakly-supervised-3d-human-pose-estimation","task_name":"Weakly-supervised 3D Human Pose Estimation"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-3d-human-pose-estimation-on","task":"Unsupervised 3D Human Pose Estimation","dataset":"Human3.6M","model":"HDVR","rank_in_archive_order":3,"of":12,"metrics":{"MPJPE":"82.1"},"uses_additional_data":false},{"leaderboard":"/sota/weakly-supervised-3d-human-pose-estimation-on","task":"Weakly-supervised 3D Human Pose Estimation","dataset":"Human3.6M","model":"HDVR","rank_in_archive_order":3,"of":33,"metrics":{"Average MPJPE (mm)":"47.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2109.09166","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}