{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/self-supervised-3d-face-reconstruction-via-1","title":"Self-Supervised 3D Face Reconstruction via Conditional Estimation","arxiv_id":"2110.04800","date":"2021-10-10","proceeding":"ICCV 2021 10","authors":["Yandong Wen","Weiyang Liu","Bhiksha Raj","Rita Singh"],"abstract":"We present a conditional estimation (CEST) framework to learn 3D facial parameters from 2D single-view images by self-supervised training from videos. CEST is based on the process of analysis by synthesis, where the 3D facial parameters (shape, reflectance, viewpoint, and illumination) are estimated from the face image, and then recombined to reconstruct the 2D face image. In order to learn semantically meaningful 3D facial parameters without explicit access to their labels, CEST couples the estimation of different 3D facial parameters by taking their statistical dependency into account. Specifically, the estimation of any 3D facial parameter is not only conditioned on the given image, but also on the facial parameters that have already been derived. Moreover, the reflectance symmetry and consistency among the video frames are adopted to improve the disentanglement of facial parameters. Together with a novel strategy for incorporating the reflectance symmetry and consistency, CEST can be efficiently trained with in-the-wild video clips. Both qualitative and quantitative experiments demonstrate the effectiveness of CEST.","url_abs":"https://arxiv.org/abs/2110.04800v1","url_pdf":"https://arxiv.org/pdf/2110.04800v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-face-reconstruction","task_name":"3D Face Reconstruction"},{"task_slug":"disentanglement","task_name":"Disentanglement"},{"task_slug":"face-reconstruction","task_name":"Face Reconstruction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-face-reconstruction-on-realy","task":"3D Face Reconstruction","dataset":"REALY","model":"CEST","rank_in_archive_order":16,"of":24,"metrics":{"@cheek":"1.456 (±0.485)","@forehead":"2.384 (±0.578)","@mouth":"1.448 (±0.406)","@nose":"2.779 (±0.835)","all":"2.017"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2110.04800","atlas_url":"https://app.syntology.ai/?focus=2110.04800","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}