{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/style-aggregated-network-for-facial-landmark","title":"Style Aggregated Network for Facial Landmark Detection","arxiv_id":"1803.04108","date":"2018-03-12","proceeding":"CVPR 2018 6","authors":["Xuanyi Dong","Yan Yan","Wanli Ouyang","Yi Yang"],"abstract":"Recent advances in facial landmark detection achieve success by learning\ndiscriminative features from rich deformation of face shapes and poses. Besides\nthe variance of faces themselves, the intrinsic variance of image styles, e.g.,\ngrayscale vs. color images, light vs. dark, intense vs. dull, and so on, has\nconstantly been overlooked. This issue becomes inevitable as increasing web\nimages are collected from various sources for training neural networks. In this\nwork, we propose a style-aggregated approach to deal with the large intrinsic\nvariance of image styles for facial landmark detection. Our method transforms\noriginal face images to style-aggregated images by a generative adversarial\nmodule. The proposed scheme uses the style-aggregated image to maintain face\nimages that are more robust to environmental changes. Then the original face\nimages accompanying with style-aggregated ones play a duet to train a landmark\ndetector which is complementary to each other. In this way, for each face, our\nmethod takes two images as input, i.e., one in its original style and the other\nin the aggregated style. In experiments, we observe that the large variance of\nimage styles would degenerate the performance of facial landmark detectors.\nMoreover, we show the robustness of our method to the large variance of image\nstyles by comparing to a variant of our approach, in which the generative\nadversarial module is removed, and no style-aggregated images are used. Our\napproach is demonstrated to perform well when compared with state-of-the-art\nalgorithms on benchmark datasets AFLW and 300-W. Code is publicly available on\nGitHub: https://github.com/D-X-Y/SAN","url_abs":"http://arxiv.org/abs/1803.04108v4","url_pdf":"http://arxiv.org/pdf/1803.04108v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"style-aggregated-network-for-facial-landmark","repo_url":"https://github.com/D-X-Y/SAN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"face-alignment","task_name":"Face Alignment"},{"task_slug":"facial-landmark-detection","task_name":"Facial Landmark Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/face-alignment-on-300w","task":"Face Alignment","dataset":"300W","model":"SAN","rank_in_archive_order":36,"of":48,"metrics":{"NME_inter-ocular (%, Challenge)":"6.60","NME_inter-ocular (%, Common)":"3.34","NME_inter-ocular (%, Full)":"3.98"},"uses_additional_data":false},{"leaderboard":"/sota/face-alignment-on-aflw-19","task":"Face Alignment","dataset":"AFLW-19","model":"SAN","rank_in_archive_order":17,"of":23,"metrics":{"AUC_box@0.07 (%, Full)":"54.0","NME_box (%, Full)":"4.04","NME_diag (%, Frontal)":"1.85","NME_diag (%, Full)":"1.91"},"uses_additional_data":false},{"leaderboard":"/sota/facial-landmark-detection-on-300w","task":"Facial Landmark Detection","dataset":"300W","model":"SAN GT","rank_in_archive_order":11,"of":15,"metrics":{"NME":"3.98"},"uses_additional_data":false},{"leaderboard":"/sota/facial-landmark-detection-on-aflw-front","task":"Facial Landmark Detection","dataset":"AFLW-Front","model":"SAN","rank_in_archive_order":3,"of":3,"metrics":{"Mean NME ":"1.85"},"uses_additional_data":false},{"leaderboard":"/sota/facial-landmark-detection-on-aflw-full","task":"Facial Landmark Detection","dataset":"AFLW-Full","model":"SAN","rank_in_archive_order":3,"of":5,"metrics":{"Mean NME ":"1.91"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.04108","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}