{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-good-practices-for-deep-3d-hand-pose","title":"Towards Good Practices for Deep 3D Hand Pose Estimation","arxiv_id":"1707.07248","date":"2017-07-23","proceeding":null,"authors":["Hengkai Guo","Guijin Wang","Xinghao Chen","Cairong Zhang"],"abstract":"3D hand pose estimation from single depth image is an important and\nchallenging problem for human-computer interaction. Recently deep convolutional\nnetworks (ConvNet) with sophisticated design have been employed to address it,\nbut the improvement over traditional random forest based methods is not so\napparent. To exploit the good practice and promote the performance for hand\npose estimation, we propose a tree-structured Region Ensemble Network (REN) for\ndirectly 3D coordinate regression. It first partitions the last convolution\noutputs of ConvNet into several grid regions. The results from separate\nfully-connected (FC) regressors on each regions are then integrated by another\nFC layer to perform the estimation. By exploitation of several training\nstrategies including data augmentation and smooth $L_1$ loss, proposed REN can\nsignificantly improve the performance of ConvNet to localize hand joints. The\nexperimental results demonstrate that our approach achieves the best\nperformance among state-of-the-art algorithms on three public hand pose\ndatasets. We also experiment our methods on fingertip detection and human pose\ndatasets and obtain state-of-the-art accuracy.","url_abs":"http://arxiv.org/abs/1707.07248v1","url_pdf":"http://arxiv.org/pdf/1707.07248v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-hand-pose-estimation","task_name":"3D Hand Pose Estimation"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"fingertip-detection","task_name":"Fingertip Detection"},{"task_slug":"hand-pose-estimation","task_name":"Hand Pose Estimation"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/hand-pose-estimation-on-icvl-hands","task":"Hand Pose Estimation","dataset":"ICVL Hands","model":"Tree Region Ensemble Network","rank_in_archive_order":13,"of":15,"metrics":{"Average 3D Error":"7.31"},"uses_additional_data":false},{"leaderboard":"/sota/hand-pose-estimation-on-nyu-hands","task":"Hand Pose Estimation","dataset":"NYU Hands","model":"REN","rank_in_archive_order":17,"of":17,"metrics":{"Average 3D Error":"15.6"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-itop-front-view","task":"Pose Estimation","dataset":"ITOP front-view","model":"REN","rank_in_archive_order":6,"of":7,"metrics":{"Mean mAP":"84.9"},"uses_additional_data":false},{"leaderboard":"/sota/pose-estimation-on-itop-top-view","task":"Pose Estimation","dataset":"ITOP top-view","model":"REN","rank_in_archive_order":5,"of":5,"metrics":{"Mean mAP":"75.5"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1707.07248","atlas_url":"https://app.syntology.ai/?focus=1707.07248","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}