Browse State-of-the-Art › Multimodal Sentiment Analysis
Multimodal Sentiment Analysis
93 papers with code · 5 benchmarks · 7 datasets archive 2025-07-28
Multimodal sentiment analysis is the task of performing sentiment analysis with multiple data sources - e.g. a camera feed of someone's face and their recorded speech.
( Image credit: ICON: Interactive Conversational Memory Network for Multimodal Emotion Detection )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CMU-MOSEI (15 rows) | SeMUL-PCD | Multi-label Emotion Analysis in Conversation via Multimodal... | — | — | Compare |
| CMU-MOSI (12 rows) | MMML | Multimodal Multi-loss Fusion Network for Sentiment Analysis | code | — | Compare |
| MOSI (11 rows) | MMML | Multimodal Multi-loss Fusion Network for Sentiment Analysis | code | — | Compare |
| B-T4SA (7 rows) | AutoML-Based Fusion Approach | An AutoML-based Approach to Multimodal Image Sentiment Analysis | — | — | Compare |
| CH-SIMS (2 rows) | MMML | Multimodal Multi-loss Fusion Network for Sentiment Analysis | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 93 papers with code (202 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Nov 2018 5 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedHumans convey their intentions through the usage of both verbal and nonverbal behaviors during face-to-face communication.
-
1 Jun 2019 4 repositories listed Syntology ran 8 of 17 samples · 9 unverified · 10 pointer-only (licence)Human language is often multimodal, which comprehends a mixture of natural language, facial gestures, and acoustic behaviors.
-
10 Oct 2018 4 repositories listedSpeech emotion recognition is a challenging task, and extensive reliance has been placed on models that use audio features in building well-performing classifiers.
-
28 Sep 2022 3 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedIn this work, we present the Textless Vision-Language Transformer (TVLT), where homogeneous transformer blocks take raw visual and audio inputs for vision-and-language representation learning with minimal…
-
23 Mar 2022 3 repositories listedThe platform features a fully modular video sentiment analysis framework consisting of data management, feature extraction, model training, and result analysis modules.
-
17 Apr 2019 3 repositories listedTherefore, in this paper, based on audio and text, we consider the task of multimodal sentiment analysis and propose a novel fusion strategy including both multi-feature fusion and multi-modality fusion to improve the…
-
31 May 2018 3 repositories listedPrevious research in this field has exploited the expressiveness of tensors for multimodal representation.
-
11 Dec 2024 2 repositories listedWhile Multimodal Large Language Models (MLLMs) demonstrate robust general capabilities, they face considerable challenges in the field of affective computing, particularly in detecting subtle facial expressions and…
-
24 Jul 2024 2 repositories listedSentiment Reasoning is an auxiliary task in sentiment analysis where the model predicts both the sentiment label and generates the rationale behind it based on the input transcript.
-
4 Sep 2023 2 repositories listed Syntology ran 6 of 8 samples · 2 unverifiedSentiment analysis is a crucial task that aims to understand people's emotional states and predict emotional categories based on multimodal information.
-
1 Sep 2021 2 repositories listed Syntology ran 4 of 7 samples · 3 unverifiedIn this work, we propose a framework named MultiModal InfoMax (MMIM), which hierarchically maximizes the Mutual Information (MI) in unimodal input pairs (inter-modality) and between multimodal fusion result and unimodal…
-
28 Jul 2021 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedMultimodal sentiment analysis aims to extract and integrate semantic information collected from multiple modalities to recognize the expressed emotions and sentiment in multimodal data.
-
9 Feb 2021 2 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedOn MOSI and MOSEI datasets, our method surpasses the current state-of-the-art methods.
-
7 May 2020 2 repositories listed Syntology ran 2 of 10 samples · 8 unverifiedIn this paper, we aim to learn effective modality representations to aid the process of fusion.
-
19 Dec 2018 2 repositories listedOur method is based on the key insight that translation from a source to a target modality provides a method of learning joint representations using only the source modality as input.
-
25 May 2018 2 repositories listedWe propose a novel approach to multimodal sentiment analysis using deep neural networks combining visual analysis and natural language processing.
-
3 Feb 2018 2 repositories listedIn this paper, we propose the Gated Multimodal Embedding LSTM with Temporal Attention (GME-LSTM(A)) model that is composed of 2 modules.
-
3 Feb 2018 2 repositories listedAI must understand each modality and the interactions between them that shape human communication.
-
23 Jul 2017 2 repositories listedMultimodal sentiment analysis is an increasingly popular research area, which extends the conventional language-based definition of sentiment analysis to a multimodal setup where other relevant modalities accompany…
-
1 Jul 2017 2 repositories listedMultimodal sentiment analysis is a developing area of research, which involves the identification of sentiments in videos.
-
7 Jun 2025 1 repository listedSeveral machine learning algorithms have been developed for the prediction of Alzheimer's disease and related dementia (ADRD) from spontaneous speech.
-
18 Feb 2025 1 repository listedCurrent Multimodal Sentiment Analysis (MSA) and Emotion Recognition in Conversations (ERC) methods based on pre-trained language models exhibit two primary limitations: 1) Once trained for MSA and ERC tasks, these…
-
21 Jan 2025 1 repository listedIn this paper, by introducing a Multi-Modality Collaborative Learning (MMCL) framework, we facilitate cross-modal interactions and capture enhanced and complementary features from modality-common and modality-specific…
-
16 Dec 2024 1 repository listed Syntology ran 0 of 6 samples · 6 unverifiedTo address these issues, we propose a Disentangled-Language-Focused (DLF) multimodal representation learning framework, which incorporates a feature disentanglement module to separate modality-shared and…
-
13 Dec 2024 1 repository listedDespite multimodal sentiment analysis being a fertile research ground that merits further investigation, current approaches take up high annotation cost and suffer from label ambiguity, non-amicable to high-quality…
-
10 Dec 2024 1 repository listed Syntology ran 5 of 10 samples · 5 unverifiedHowever, in real-world dynamic scenarios, the distribution of target data is always changing and different from the source data used to train the model, which leads to performance degradation.
-
19 Oct 2024 1 repository listedIn multimodal sentiment analysis, collecting text data is often more challenging than video or audio due to higher annotation costs and inconsistent automatic speech recognition (ASR) quality.
-
6 Oct 2024 1 repository listedMultimodal Sentiment Analysis (MSA) utilizes multimodal data to infer the users' sentiment.
-
30 Sep 2024 1 repository listed Syntology ran 10 of 14 samples · 4 unverifiedThe field of Multimodal Sentiment Analysis (MSA) has recently witnessed an emerging direction seeking to tackle the issue of data incompleteness.
-
11 Sep 2024 1 repository listedThe goal of this survey is to explore the current landscape of multimodal affective research, identify development trends, and highlight the similarities and differences across various tasks, offering a comprehensive…
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections