{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/leveraging-recent-advances-in-deep-learning","title":"Leveraging Recent Advances in Deep Learning for Audio-Visual Emotion Recognition","arxiv_id":"2103.09154","date":"2021-03-16","proceeding":null,"authors":["Liam Schoneveld","Alice Othmani","Hazem Abdelkawy"],"abstract":"Emotional expressions are the behaviors that communicate our emotional state or attitude to others. They are expressed through verbal and non-verbal communication. Complex human behavior can be understood by studying physical features from multiple modalities; mainly facial, vocal and physical gestures. Recently, spontaneous multi-modal emotion recognition has been extensively studied for human behavior analysis. In this paper, we propose a new deep learning-based approach for audio-visual emotion recognition. Our approach leverages recent advances in deep learning like knowledge distillation and high-performing deep architectures. The deep feature representations of the audio and visual modalities are fused based on a model-level fusion strategy. A recurrent neural network is then used to capture the temporal dynamics. Our proposed approach substantially outperforms state-of-the-art approaches in predicting valence on the RECOLA dataset. Moreover, our proposed visual facial expression feature extraction network outperforms state-of-the-art results on the AffectNet and Google Facial Expression Comparison datasets.","url_abs":"https://arxiv.org/abs/2103.09154v2","url_pdf":"https://arxiv.org/pdf/2103.09154v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"facial-expression-recognition","task_name":"Facial Expression Recognition (FER)"},{"task_slug":"knowledge-distillation","task_name":"Knowledge Distillation"}],"methods":[{"method_slug":"knowledge-distillation","method_name":"Knowledge Distillation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/facial-expression-recognition-on-affectnet","task":"Facial Expression Recognition (FER)","dataset":"AffectNet","model":"Distilled student","rank_in_archive_order":18,"of":50,"metrics":{"Accuracy (7 emotion)":"65.4","Accuracy (8 emotion)":"61.60"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2103.09154","atlas_url":"https://app.syntology.ai/?focus=2103.09154","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}