{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unimse-towards-unified-multimodal-sentiment","title":"UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition","arxiv_id":"2211.11256","date":"2022-11-21","proceeding":null,"authors":["Guimin Hu","Ting-En Lin","Yi Zhao","Guangming Lu","Yuchuan Wu","Yongbin Li"],"abstract":"Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors. From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a longer period. However, most existing works study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two. In this paper, we propose a multimodal sentiment knowledge-sharing framework (UniMSE) that unifies MSA and ERC tasks from features, labels, and models. We perform modality fusion at the syntactic and semantic levels and introduce contrastive learning between modalities and samples to better capture the difference and consistency between sentiments and emotions. Experiments on four public benchmark datasets, MOSI, MOSEI, MELD, and IEMOCAP, demonstrate the effectiveness of the proposed method and achieve consistent improvements compared with state-of-the-art methods.","url_abs":"https://arxiv.org/abs/2211.11256v1","url_pdf":"https://arxiv.org/pdf/2211.11256v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unimse-towards-unified-multimodal-sentiment","repo_url":"https://github.com/lemei/unimse","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"},{"task_slug":"emotion-recognition","task_name":"Emotion Recognition"},{"task_slug":"emotion-recognition-in-conversation","task_name":"Emotion Recognition in Conversation"},{"task_slug":"multimodal-sentiment-analysis","task_name":"Multimodal Sentiment Analysis"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[{"method_slug":"contrastive-learning","method_name":"Contrastive Learning"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/emotion-recognition-in-conversation-on","task":"Emotion Recognition in Conversation","dataset":"IEMOCAP","model":"UniMSE","rank_in_archive_order":13,"of":59,"metrics":{"Accuracy":"70.56","Weighted-F1":"70.66"},"uses_additional_data":false},{"leaderboard":"/sota/emotion-recognition-in-conversation-on-meld","task":"Emotion Recognition in Conversation","dataset":"MELD","model":"UniMSE","rank_in_archive_order":30,"of":68,"metrics":{"Accuracy":"65.09","Weighted-F1":"65.51"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-sentiment-analysis-on-cmu-mosei-1","task":"Multimodal Sentiment Analysis","dataset":"CMU-MOSEI","model":"UniMSE","rank_in_archive_order":3,"of":15,"metrics":{"Accuracy":"87.50","F1":"87.46","MAE":"0.523"},"uses_additional_data":true},{"leaderboard":"/sota/multimodal-sentiment-analysis-on-cmu-mosi","task":"Multimodal Sentiment Analysis","dataset":"CMU-MOSI","model":"UniMSE","rank_in_archive_order":3,"of":12,"metrics":{"Acc-2":"86.9","Acc-7":"48.68","Corr":"0.809","F1":"86.42","MAE":"0.691"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-sentiment-analysis-on-mosi","task":"Multimodal Sentiment Analysis","dataset":"MOSI","model":"UniMSE","rank_in_archive_order":3,"of":11,"metrics":{"Accuracy":"86.9","F1 score":"86.42"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2211.11256","atlas_url":"https://app.syntology.ai/?focus=2211.11256","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}