{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-utterance-level-affect-analysis","title":"Multimodal Utterance-level Affect Analysis using Visual, Audio and Text Features","arxiv_id":"1805.00625","date":"2018-05-02","proceeding":null,"authors":["Didan Deng","Yuqian Zhou","Jimin Pi","Bertram E. Shi"],"abstract":"The integration of information across multiple modalities and across time is\na promising way to enhance the emotion recognition performance of affective\nsystems. Much previous work has focused on instantaneous emotion recognition.\nThe 2018 One-Minute Gradual-Emotion Recognition (OMG-Emotion) challenge, which\nwas held in conjunction with the IEEE World Congress on Computational\nIntelligence, encouraged participants to address long-term emotion recognition\nby integrating cues from multiple modalities, including facial expression,\naudio and language. Intuitively, a multi-modal inference network should be able\nto leverage information from each modality and their correlations to improve\nrecognition over that achievable by a single modality network. We describe here\na multi-modal neural architecture that integrates visual information over time\nusing an LSTM, and combines it with utterance level audio and text cues to\nrecognize human sentiment from multimodal clips. Our model outperforms the\nunimodal baseline, achieving the concordance correlation coefficients (CCC) of\n0.400 on the arousal task, and 0.353 on the valence task.","url_abs":"http://arxiv.org/abs/1805.00625v2","url_pdf":"http://arxiv.org/pdf/1805.00625v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multimodal-utterance-level-affect-analysis","repo_url":"https://github.com/HKUST-NISL/OMG-Emotion-Challenge-2018","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"multimodal-utterance-level-affect-analysis","repo_url":"https://github.com/toxtli/AutomEditor","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.00625","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}