Papers › UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

21 Nov 2022arXiv:2211.11256archive 2025-07-28

Guimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu, Yuchuan Wu, Yongbin Li

Multimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors. From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a longer period. However, most existing works study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two. In this paper, we propose a multimodal sentiment knowledge-sharing framework (UniMSE) that unifies MSA and ERC tasks from features, labels, and models. We perform modality fusion at the syntactic and semantic levels and introduce contrastive learning between modalities and samples to better capture the difference and consistency between sentiments and emotions. Experiments on four public benchmark datasets, MOSI, MOSEI, MELD, and IEMOCAP, demonstrate the effectiveness of the proposed method and achieve consistent improvements compared with state-of-the-art methods.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

lemei/unimse officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Contrastive LearningEmotion RecognitionEmotion Recognition in ConversationMultimodal Sentiment AnalysisSentiment Analysis

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Emotion Recognition in Conversation IEMOCAP UniMSE Accuracy 70.56 #13 of 59 Archive leaderboard report
Emotion Recognition in Conversation IEMOCAP UniMSE Weighted-F1 70.66 #13 of 59 Archive leaderboard report
Emotion Recognition in Conversation MELD UniMSE Accuracy 65.09 #30 of 68 Archive leaderboard report
Emotion Recognition in Conversation MELD UniMSE Weighted-F1 65.51 #30 of 68 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSEI UniMSE Accuracy 87.50 #3 of 15 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSEI UniMSE F1 87.46 #3 of 15 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSEI UniMSE MAE 0.523 #3 of 15 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSI UniMSE Acc-2 86.9 #3 of 12 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSI UniMSE Acc-7 48.68 #3 of 12 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSI UniMSE Corr 0.809 #3 of 12 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSI UniMSE F1 86.42 #3 of 12 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSI UniMSE MAE 0.691 #3 of 12 Archive leaderboard report
Multimodal Sentiment Analysis MOSI UniMSE Accuracy 86.9 #3 of 11 Archive leaderboard report
Multimodal Sentiment Analysis MOSI UniMSE F1 score 86.42 #3 of 11 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Contrastive Learning

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections