{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-region-and-multi-label-learning-for","title":"Deep Region and Multi-Label Learning for Facial Action Unit Detection","arxiv_id":null,"date":"2016-06-01","proceeding":"CVPR 2016 6","authors":["Kaili Zhao","Wen-Sheng Chu","Honggang Zhang"],"abstract":"Region learning (RL) and multi-label learning (ML) have recently attracted increasing attentions in the field of facial Action Unit (AU) detection. Knowing that AUs are active on sparse facial regions, RL aims to identify these regions for a better specificity. On the other hand, a strong statistical evidence of AU correlations suggests that ML is a natural way to model the detection task. In this paper, we propose Deep Region and Multi-label Learning (DRML), a unified deep network that simultaneously addresses these two problems. One crucial aspect in DRML is a novel region layer that uses feed-forward functions to induce important facial regions, forcing the learned weights to capture structural information of the face. Our region layer serves as an alternative design between locally connected layers (i.e., confined kernels to individual pixels) and conventional convolution layers (i.e., shared kernels across an entire image). Unlike previous studies that solve RL and ML alternately, DRML by construction addresses both problems, allowing the two seemingly irrelevant problems to interact more directly. The complete network is end-to-end trainable, and automatically learns representations robust to variations inherent within a local region. Experiments on BP4D and DISFA benchmarks show that DRML performs the highest average F1-score and AUC within and across datasets in comparison with alternative methods. ","url_abs":"http://openaccess.thecvf.com/content_cvpr_2016/html/Zhao_Deep_Region_and_CVPR_2016_paper.html","url_pdf":"http://openaccess.thecvf.com/content_cvpr_2016/papers/Zhao_Deep_Region_and_CVPR_2016_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-region-and-multi-label-learning-for","repo_url":"https://github.com/AlexHex7/DRML_pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"deep-region-and-multi-label-learning-for","repo_url":"https://github.com/zkl20061823/DRML","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-unit-detection","task_name":"Action Unit Detection"},{"task_slug":"facial-action-unit-detection","task_name":"Facial Action Unit Detection"},{"task_slug":"multi-label-learning","task_name":"Multi-Label Learning"},{"task_slug":"specificity","task_name":"Specificity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/facial-action-unit-detection-on-bp4d","task":"Facial Action Unit Detection","dataset":"BP4D","model":"DRML","rank_in_archive_order":10,"of":10,"metrics":{"Average AUC":"56.0","Average F1":"48.3"},"uses_additional_data":false},{"leaderboard":"/sota/facial-action-unit-detection-on-disfa","task":"Facial Action Unit Detection","dataset":"DISFA","model":"DRML","rank_in_archive_order":8,"of":8,"metrics":{"Average AUC":"52.3","Average F1":"26.7"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}