{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/joint-voxel-and-coordinate-regression-for","title":"Joint Voxel and Coordinate Regression for Accurate 3D Facial Landmark Localization","arxiv_id":"1801.09242","date":"2018-01-28","proceeding":null,"authors":["Hongwen Zhang","Qi Li","Zhenan Sun"],"abstract":"3D face shape is more expressive and viewpoint-consistent than its 2D\ncounterpart. However, 3D facial landmark localization in a single image is\nchallenging due to the ambiguous nature of landmarks under 3D perspective.\nExisting approaches typically adopt a suboptimal two-step strategy, performing\n2D landmark localization followed by depth estimation. In this paper, we\npropose the Joint Voxel and Coordinate Regression (JVCR) method for 3D facial\nlandmark localization, addressing it more effectively in an end-to-end fashion.\nFirst, a compact volumetric representation is proposed to encode the per-voxel\nlikelihood of positions being the 3D landmarks. The dimensionality of such a\nrepresentation is fixed regardless of the number of target landmarks, so that\nthe curse of dimensionality could be avoided. Then, a stacked hourglass network\nis adopted to estimate the volumetric representation from coarse to fine,\nfollowed by a 3D convolution network that takes the estimated volume as input\nand regresses 3D coordinates of the face shape. In this way, the 3D structural\nconstraints between landmarks could be learned by the neural network in a more\nefficient manner. Moreover, the proposed pipeline enables end-to-end training\nand improves the robustness and accuracy of 3D facial landmark localization.\nThe effectiveness of our approach is validated on the 3DFAW and AFLW2000-3D\ndatasets. Experimental results show that the proposed method achieves\nstate-of-the-art performance in comparison with existing methods.","url_abs":"http://arxiv.org/abs/1801.09242v1","url_pdf":"http://arxiv.org/pdf/1801.09242v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-facial-landmark-localization","task_name":"3D Facial Landmark Localization"},{"task_slug":"depth-estimation","task_name":"Depth Estimation"},{"task_slug":"face-alignment","task_name":"Face Alignment"},{"task_slug":"facial-landmark-detection","task_name":"Facial Landmark Detection"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[{"method_slug":"3d-convolution","method_name":"3D Convolution"},{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-facial-landmark-localization-on-3dfaw","task":"3D Facial Landmark Localization","dataset":"3DFAW","model":"JVCR","rank_in_archive_order":1,"of":1,"metrics":{"CVGTCE":"3.46","GTE":"4.35"},"uses_additional_data":false},{"leaderboard":"/sota/3d-facial-landmark-localization-on-aflw2000","task":"3D Facial Landmark Localization","dataset":"AFLW2000-3D","model":"JVCR","rank_in_archive_order":1,"of":1,"metrics":{"GTE":"7.28"},"uses_additional_data":false},{"leaderboard":"/sota/facial-landmark-detection-on-aflw2000-3d","task":"Facial Landmark Detection","dataset":"AFLW2000-3D","model":"JVCR","rank_in_archive_order":1,"of":1,"metrics":{"GTE":"7.28"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1801.09242","atlas_url":"https://app.syntology.ai/?focus=1801.09242","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}