{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-multimodal-emotion-recognition","title":"End-to-End Multimodal Emotion Recognition using Deep Neural Networks","arxiv_id":"1704.08619","date":"2017-04-27","proceeding":null,"authors":["Panagiotis Tzirakis","George Trigeorgis","Mihalis A. Nicolaou","Björn Schuller","Stefanos Zafeiriou"],"abstract":"Automatic affect recognition is a challenging task due to the various\nmodalities emotions can be expressed with. Applications can be found in many\ndomains including multimedia retrieval and human computer interaction. In\nrecent years, deep neural networks have been used with great success in\ndetermining emotional states. Inspired by this success, we propose an emotion\nrecognition system using auditory and visual modalities. To capture the\nemotional content for various styles of speaking, robust features need to be\nextracted. To this purpose, we utilize a Convolutional Neural Network (CNN) to\nextract features from the speech, while for the visual modality a deep residual\nnetwork (ResNet) of 50 layers. In addition to the importance of feature\nextraction, a machine learning algorithm needs also to be insensitive to\noutliers while being able to model the context. To tackle this problem, Long\nShort-Term Memory (LSTM) networks are utilized. The system is then trained in\nan end-to-end fashion where - by also taking advantage of the correlations of\nthe each of the streams - we manage to significantly outperform the traditional\napproaches based on auditory and visual handcrafted features for the prediction\nof spontaneous and natural emotions on the RECOLA database of the AVEC 2016\nresearch challenge on emotion recognition.","url_abs":"http://arxiv.org/abs/1704.08619v1","url_pdf":"http://arxiv.org/pdf/1704.08619v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-multimodal-emotion-recognition","repo_url":"https://github.com/asfathermou/human-computer-interaction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"end-to-end-multimodal-emotion-recognition","repo_url":"https://github.com/tzirakis/Multimodal-Emotion-Recognition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"multimodal-emotion-recognition","task_name":"Multimodal Emotion Recognition"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.08619","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}