Papers › Learning Language-guided Adaptive Hyper-modality Representation for Multimodal...

Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis

9 Oct 2023arXiv:2310.05804archive 2025-07-28

Haoyu Zhang, Yu Wang, Guanghao Yin, Kejun Liu, Yuanyuan Liu, Tianshu Yu

Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved. To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales. With the obtained hyper-modality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA. In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Haoyu-ha/ALMT officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Multimodal Sentiment AnalysisSentiment Analysis

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Multimodal Sentiment Analysis CH-SIMS ALMT Acc-2 81.19 #2 of 2 Archive leaderboard report
Multimodal Sentiment Analysis CH-SIMS ALMT Acc-3 68.93 #2 of 2 Archive leaderboard report
Multimodal Sentiment Analysis CH-SIMS ALMT Acc-5 45.73 #2 of 2 Archive leaderboard report
Multimodal Sentiment Analysis CH-SIMS ALMT CORR 0.619 #2 of 2 Archive leaderboard report
Multimodal Sentiment Analysis CH-SIMS ALMT F1 81.57 #2 of 2 Archive leaderboard report
Multimodal Sentiment Analysis CH-SIMS ALMT MAE 0.404 #2 of 2 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSEI ALMT Acc-5 55.96 #15 of 15 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSEI ALMT Acc-7 54.28 #15 of 15 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSEI ALMT Corr 0.779 #15 of 15 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSEI ALMT MAE 0.526 #15 of 15 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSI ALMT Acc-5 56.41 #10 of 12 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSI ALMT Acc-7 49.42 #10 of 12 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSI ALMT Corr 0.805 #10 of 12 Archive leaderboard report
Multimodal Sentiment Analysis CMU-MOSI ALMT MAE 0.683 #10 of 12 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections