{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/faster-than-real-time-facial-alignment-a-3d","title":"Faster Than Real-time Facial Alignment: A 3D Spatial Transformer Network Approach in Unconstrained Poses","arxiv_id":"1707.05653","date":"2017-07-18","proceeding":"ICCV 2017 10","authors":["Chandrasekhar Bhagavatula","Chenchen Zhu","Khoa Luu","Marios Savvides"],"abstract":"Facial alignment involves finding a set of landmark points on an image with a\nknown semantic meaning. However, this semantic meaning of landmark points is\noften lost in 2D approaches where landmarks are either moved to visible\nboundaries or ignored as the pose of the face changes. In order to extract\nconsistent alignment points across large poses, the 3D structure of the face\nmust be considered in the alignment step. However, extracting a 3D structure\nfrom a single 2D image usually requires alignment in the first place. We\npresent our novel approach to simultaneously extract the 3D shape of the face\nand the semantically consistent 2D alignment through a 3D Spatial Transformer\nNetwork (3DSTN) to model both the camera projection matrix and the warping\nparameters of a 3D model. By utilizing a generic 3D model and a Thin Plate\nSpline (TPS) warping function, we are able to generate subject specific 3D\nshapes without the need for a large 3D shape basis. In addition, our proposed\nnetwork can be trained in an end-to-end framework on entirely synthetic data\nfrom the 300W-LP dataset. Unlike other 3D methods, our approach only requires\none pass through the network resulting in a faster than real-time alignment.\nEvaluations of our model on the Annotated Facial Landmarks in the Wild (AFLW)\nand AFLW2000-3D datasets show our method achieves state-of-the-art performance\nover other 3D approaches to alignment.","url_abs":"http://arxiv.org/abs/1707.05653v2","url_pdf":"http://arxiv.org/pdf/1707.05653v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"face-alignment","task_name":"Face Alignment"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/face-alignment-on-aflw2000-3d","task":"Face Alignment","dataset":"AFLW2000-3D","model":"3DSTN","rank_in_archive_order":11,"of":14,"metrics":{"Balanced NME (2D Sparse Alignment)":"4.49%"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.05653","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}