{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/supervision-by-registration-an-unsupervised","title":"Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors","arxiv_id":"1807.00966","date":"2018-07-03","proceeding":"CVPR 2018 6","authors":["Xuanyi Dong","Shoou-I Yu","Xinshuo Weng","Shih-En Wei","Yi Yang","Yaser Sheikh"],"abstract":"In this paper, we present supervision-by-registration, an unsupervised\napproach to improve the precision of facial landmark detectors on both images\nand video. Our key observation is that the detections of the same landmark in\nadjacent frames should be coherent with registration, i.e., optical flow.\nInterestingly, the coherency of optical flow is a source of supervision that\ndoes not require manual labeling, and can be leveraged during detector\ntraining. For example, we can enforce in the training loss function that a\ndetected landmark at frame$_{t-1}$ followed by optical flow tracking from\nframe$_{t-1}$ to frame$_t$ should coincide with the location of the detection\nat frame$_t$. Essentially, supervision-by-registration augments the training\nloss function with a registration loss, thus training the detector to have\noutput that is not only close to the annotations in labeled images, but also\nconsistent with registration on large amounts of unlabeled videos. End-to-end\ntraining with the registration loss is made possible by a differentiable\nLucas-Kanade operation, which computes optical flow registration in the forward\npass, and back-propagates gradients that encourage temporal coherency in the\ndetector. The output of our method is a more precise image-based facial\nlandmark detector, which can be applied to single images or video. With\nsupervision-by-registration, we demonstrate (1) improvements in facial landmark\ndetection on both images (300W, ALFW) and video (300VW, Youtube-Celebrities),\nand (2) significant reduction of jittering in video detections.","url_abs":"http://arxiv.org/abs/1807.00966v2","url_pdf":"http://arxiv.org/pdf/1807.00966v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"supervision-by-registration-an-unsupervised","repo_url":"https://github.com/facebookresearch/supervision-by-registration","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"facial-landmark-detection","task_name":"Facial Landmark Detection"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/facial-landmark-detection-on-300-vw-c","task":"Facial Landmark Detection","dataset":"300-VW (C)","model":"CPM+SBR+PAM","rank_in_archive_order":1,"of":2,"metrics":{"AUC0.08 private":"59.39"},"uses_additional_data":false},{"leaderboard":"/sota/facial-landmark-detection-on-300-vw-c","task":"Facial Landmark Detection","dataset":"300-VW (C)","model":"CPM+SBR","rank_in_archive_order":2,"of":2,"metrics":{"AUC0.08 private":"58.22"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1807.00966","atlas_url":"https://app.syntology.ai/?focus=1807.00966","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}