{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/expression-affect-action-unit-recognition-aff","title":"Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace","arxiv_id":"1910.04855","date":"2019-09-25","proceeding":null,"authors":["Dimitrios Kollias","Stefanos Zafeiriou"],"abstract":"Affective computing has been largely limited in terms of available data resources. The need to collect and annotate diverse in-the-wild datasets has become apparent with the rise of deep learning models, as the default approach to address any computer vision task. Some in-the-wild databases have been recently proposed. However: i) their size is small, ii) they are not audiovisual, iii) only a small part is manually annotated, iv) they contain a small number of subjects, or v) they are not annotated for all main behavior tasks (valence-arousal estimation, action unit detection and basic expression classification). To address these, we substantially extend the largest available in-the-wild database (Aff-Wild) to study continuous emotions such as valence and arousal. Furthermore, we annotate parts of the database with basic expressions and action units. As a consequence, for the first time, this allows the joint study of all three types of behavior states. We call this database Aff-Wild2. We conduct extensive experiments with CNN and CNN-RNN architectures that use visual and audio modalities; these networks are trained on Aff-Wild2 and their performance is then evaluated on 10 publicly available emotion databases. We show that the networks achieve state-of-the-art performance for the emotion recognition tasks. Additionally, we adapt the ArcFace loss function in the emotion recognition context and use it for training two new networks on Aff-Wild2 and then re-train them in a variety of diverse expression recognition databases. The networks are shown to improve the existing state-of-the-art. The database, emotion recognition models and source code are available at http://ibug.doc.ic.ac.uk/resources/aff-wild2.","url_abs":"https://arxiv.org/abs/1910.04855v1","url_pdf":"https://arxiv.org/pdf/1910.04855v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-unit-detection","task_name":"Action Unit Detection"},{"task_slug":"arousal-estimation","task_name":"Arousal Estimation"},{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"facial-expression-recognition","task_name":"Facial Expression Recognition (FER)"},{"task_slug":"multi-task-learning","task_name":"Multi-Task Learning"}],"methods":[{"method_slug":"arcface","method_name":"ArcFace"}],"datasets_introduced":[{"slug":"aff-wild2","name":"Aff-Wild2","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/facial-expression-recognition-on-affectnet","task":"Facial Expression Recognition (FER)","dataset":"AffectNet","model":"MT-ArcRes","rank_in_archive_order":12,"of":50,"metrics":{"Accuracy (8 emotion)":"63"},"uses_additional_data":false},{"leaderboard":"/sota/facial-expression-recognition-on-raf-db","task":"Facial Expression Recognition (FER)","dataset":"RAF-DB","model":"MT-ArcVGG","rank_in_archive_order":35,"of":35,"metrics":{"Avg. Accuracy":"76"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1910.04855","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}