{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cross-modal-deep-variational-hand-pose","title":"Cross-modal Deep Variational Hand Pose Estimation","arxiv_id":"1803.11404","date":"2018-03-30","proceeding":"CVPR 2018 6","authors":["Adrian Spurr","Jie Song","Seonwook Park","Otmar Hilliges"],"abstract":"The human hand moves in complex and high-dimensional ways, making estimation\nof 3D hand pose configurations from images alone a challenging task. In this\nwork we propose a method to learn a statistical hand model represented by a\ncross-modal trained latent space via a generative deep neural network. We\nderive an objective function from the variational lower bound of the VAE\nframework and jointly optimize the resulting cross-modal KL-divergence and the\nposterior reconstruction objective, naturally admitting a training regime that\nleads to a coherent latent space across multiple modalities such as RGB images,\n2D keypoint detections or 3D hand configurations. Additionally, it grants a\nstraightforward way of using semi-supervision. This latent space can be\ndirectly used to estimate 3D hand poses from RGB images, outperforming the\nstate-of-the art in different settings. Furthermore, we show that our proposed\nmethod can be used without changes on depth images and performs comparably to\nspecialized methods. Finally, the model is fully generative and can synthesize\nconsistent pairs of hand configurations across modalities. We evaluate our\nmethod on both RGB and depth datasets and analyze the latent space\nqualitatively.","url_abs":"http://arxiv.org/abs/1803.11404v1","url_pdf":"http://arxiv.org/pdf/1803.11404v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cross-modal-deep-variational-hand-pose","repo_url":"https://github.com/spurra/vae-hands-3d","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"hand-pose-estimation","task_name":"Hand Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.11404","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}