Papers › Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimodal Emotion Recognition

Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimodal Emotion Recognition

18 Nov 2023arXiv:2311.11009archive 2025-07-28

Dongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu Okumura

Multimodal emotion recognition aims to recognize emotions for each utterance of multiple modalities, which has received increasing attention for its application in human-machine interaction. Current graph-based methods fail to simultaneously depict global contextual features and local diverse uni-modal features in a dialogue. Furthermore, with the number of graph layers increasing, they easily fall into over-smoothing. In this paper, we propose a method for joint modality fusion and graph contrastive learning for multimodal emotion recognition (Joyful), where multimodality fusion, contrastive learning, and emotion recognition are jointly optimized. Specifically, we first design a new multimodal fusion mechanism that can provide deep interaction and fusion between the global contextual and uni-modal specific features. Then, we introduce a graph contrastive learning framework with inter-view and intra-view contrastive losses to learn more distinguishable representations for samples with different sentiments. Extensive experiments on three benchmark datasets indicate that Joyful achieved state-of-the-art (SOTA) performance compared to all baselines.

PaperPDFCode

Code

wykstc/MERC-main officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Contrastive LearningEmotion RecognitionEmotion Recognition in ConversationFace SwappingMultimodal Emotion Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Emotion Recognition in Conversation IEMOCAP-4 Joyful Weighted F1 85.70 #2 of 8 Archive leaderboard report
Face Swapping HOD Work 0-shot MRR Good #1 of 1 Archive leaderboard report
Multimodal Emotion Recognition IEMOCAP Joyful Accuracy 71.0 #2 of 2 Archive leaderboard report
Multimodal Emotion Recognition IEMOCAP Joyful Weighted F1 70.50 #2 of 2 Archive leaderboard report
Multimodal Emotion Recognition IEMOCAP-4 Joyful Accuracy 85.60 #2 of 11 Archive leaderboard report
Multimodal Emotion Recognition IEMOCAP-4 Joyful Weighted F1 85.70 #2 of 11 Archive leaderboard report
Multimodal Emotion Recognition MELD Joyful Accuracy 62.53 #3 of 3 Archive leaderboard report
Multimodal Emotion Recognition MELD Joyful Weighted F1 61.77 #3 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Contrastive LearningGraph Contrastive Coding

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections