Browse State-of-the-Art › Multimodal Emotion Recognition
Multimodal Emotion Recognition
80 papers with code · 7 benchmarks · 11 datasets archive 2025-07-28
This is a leaderboard for multimodal emotion recognition on the IEMOCAP dataset. The modality abbreviations are A: Acoustic T: Text V: Visual
Please include the modality in the bracket after the model name.
All models must use standard five emotion categories and are evaluated in standard leave-one-session-out (LOSO). See the papers for references.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task, 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| IEMOCAP-4 (11 rows) | GraphSmile | Tracing Intricate Cues in Dialogue: Joint Graph Structure and... | code | — | Compare |
| MELD (3 rows) | GraphSmile | Tracing Intricate Cues in Dialogue: Joint Graph Structure and... | code | — | Compare |
| IEMOCAP (2 rows) | GraphSmile | Tracing Intricate Cues in Dialogue: Joint Graph Structure and... | code | — | Compare |
| CMU-MOSEI-Sentiment (1 row) | GraphSmile | Tracing Intricate Cues in Dialogue: Joint Graph Structure and... | code | — | Compare |
| CMU-MOSEI-Sentiment-3 (1 row) | GraphSmile | Tracing Intricate Cues in Dialogue: Joint Graph Structure and... | code | — | Compare |
| Expressive hands and faces dataset (EHF). (1 row) | SMPLify-X | Multi-Modal Emotion recognition on IEMOCAP Dataset using Deep Learning | code | — | Compare |
| MELD-Sentiment (1 row) | GraphSmile | Tracing Intricate Cues in Dialogue: Joint Graph Structure and... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
11 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 80 papers with code (180 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
12 Apr 2019 5 repositories listedIn this work, we adopt a feature-engineering based approach to tackle the task of speech emotion recognition.
-
10 Oct 2018 4 repositories listedSpeech emotion recognition is a challenging task, and extensive reliance has been placed on models that use audio features in building well-performing classifiers.
-
18 Apr 2023 3 repositories listedThe first Multimodal Emotion Recognition Challenge (MER 2023) was successfully held at ACM Multimedia.
-
17 Apr 2019 3 repositories listedTherefore, in this paper, based on audio and text, we consider the task of multimodal sentiment analysis and propose a novel fusion strategy including both multi-feature fusion and multi-modality fusion to improve the…
-
11 Dec 2024 2 repositories listedWhile Multimodal Large Language Models (MLLMs) demonstrate robust general capabilities, they face considerable challenges in the field of affective computing, particularly in detecting subtle facial expressions and…
-
17 Jul 2024 2 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)Our results indicate that multimodal textualization provides lower accuracy than feature-based models on C-EXPR-DB, where text transcripts are captured in the wild.
-
17 Jun 2024 2 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedAccurate emotion perception is crucial for various applications, including human-computer interaction, education, and counseling.
-
26 Apr 2024 2 repositories listed Syntology ran 5 of 9 samples · 4 unverified · 9 pointer-only (licence)However, this process may lead to inaccurate annotations, such as ignoring non-majority or non-candidate labels.
-
5 May 2022 2 repositories listedEmotions are an inherent part of human interactions, and consequently, it is imperative to develop AI systems that understand and recognize human emotions.
-
15 Jun 2020 2 repositories listedHumans are able to comprehend information from multiple domains for e.
-
1 Nov 2018 2 repositories listedEmotion detection in conversations is a necessary step for a number of applications, including opinion mining over chat history, social media threads, debates, argumentation mining, understanding consumer feedback in…
-
16 Apr 2018 2 repositories listedEmotion recognition has become an important field of research in Human Computer Interactions as we improve upon the techniques for modelling the various aspects of behaviour.
-
1 Jul 2017 2 repositories listedMultimodal sentiment analysis is a developing area of research, which involves the identification of sentiments in videos.
-
27 Apr 2017 2 repositories listedThe system is then trained in an end-to-end fashion where - by also taking advantage of the correlations of the each of the streams - we manage to significantly outperform the traditional approaches based on auditory…
-
12 Jun 2025 1 repository listedTo address these issues, we propose a novel robust MER framework, Causal Inference Distiller (CIDer), and introduce a new task, Random Modality Feature Missing (RMFM), to generalize the definition of modality missing.
-
10 May 2025 1 repository listedSpecifically, for the redundant features, we make one modality perform intra-modal feature selection through a self-attention mechanism, so that the selected features can adaptively and efficiently interact with another…
-
21 Mar 2025 1 repository listedThis article presents our results for the eighth Affective Behavior Analysis in-the-wild (ABAW) competition.
-
7 Mar 2025 1 repository listed Syntology ran 0 of 3 samples · 3 unverified · 3 pointer-only (licence)In this work, we present the first application of Reinforcement Learning with Verifiable Reward (RLVR) to an Omni-multimodal large language model in the context of emotion recognition, a task where both visual and audio…
-
19 Feb 2025 1 repository listedSpecifically, we introduce a contrastive disentangled distribution mechanism within the emotion space to model the multimodal data, allowing for the extraction of semantic features and uncertainty.
-
1 Feb 2025 1 repository listedAblation studies further validate the contributions of each module, highlighting the significance of advanced feature extraction and fusion strategies in enhancing emotion recognition performance.
-
29 Nov 2024 1 repository listedTo address these issues, we propose a Spectral Domain Reconstruction Graph Neural Network (SDR-GNN) for incomplete multimodal learning in conversational emotion recognition.
-
13 Sep 2024 1 repository listedThen, a hypercomplex fusion module learns inter-modal relations among the embeddings of the different modalities.
-
13 Sep 2024 1 repository listedIn this paper, we introduce PHemoNet, a fully hypercomplex network for multimodal emotion recognition from physiological signals.
-
11 Sep 2024 1 repository listedThe goal of this survey is to explore the current landscape of multimodal affective research, identify development trends, and highlight the similarities and differences across various tasks, offering a comprehensive…
-
23 Aug 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedThe Multimodal Emotion Recognition challenge MER2024 focuses on recognizing emotions using audio, language, and visual signals.
-
20 Aug 2024 1 repository listed Syntology ran 5 of 8 samples · 3 unverifiedThis paper presents our winning approach for the MER-NOISE and MER-OV tracks of the MER2024 Challenge on multimodal emotion recognition.
-
18 Aug 2024 1 repository listedTo address the limitation in multimodal emotion recognition (MER) performance arising from inter-modal information fusion, we propose a novel MER framework based on multitask learning where fusion occurs after…
-
16 Aug 2024 1 repository listedHowever, PKD methods based on structural similarity are primarily confined to learning from a single joint teacher representation, which limits their robustness, accuracy, and ability to learn from diverse multimodal…
-
31 Jul 2024 1 repository listedFurthermore, GraphSmile is effortlessly applied to multimodal sentiment analysis in conversation (MSAC), forging a unified multimodal affective model capable of executing MERC and MSAC tasks.
-
26 Jul 2024 1 repository listedDespite its potential, multimodal emotion recognition faces significant challenges, particularly in synchronization, feature extraction, and fusion of diverse data sources.
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections