{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-language-guided-adaptive-hyper","title":"Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment Analysis","arxiv_id":"2310.05804","date":"2023-10-09","proceeding":null,"authors":["Haoyu Zhang","Yu Wang","Guanghao Yin","Kejun Liu","Yuanyuan Liu","Tianshu Yu"],"abstract":"Though Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved. To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales. With the obtained hyper-modality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA. In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism.","url_abs":"https://arxiv.org/abs/2310.05804v2","url_pdf":"https://arxiv.org/pdf/2310.05804v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-language-guided-adaptive-hyper","repo_url":"https://github.com/Haoyu-ha/ALMT","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"multimodal-sentiment-analysis","task_name":"Multimodal Sentiment Analysis"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multimodal-sentiment-analysis-on-ch-sims","task":"Multimodal Sentiment Analysis","dataset":"CH-SIMS","model":"ALMT","rank_in_archive_order":2,"of":2,"metrics":{"Acc-2":"81.19","Acc-3":"68.93","Acc-5":"45.73","CORR":"0.619","F1":"81.57","MAE":"0.404"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-sentiment-analysis-on-cmu-mosei-1","task":"Multimodal Sentiment Analysis","dataset":"CMU-MOSEI","model":"ALMT","rank_in_archive_order":15,"of":15,"metrics":{"Acc-5":"55.96","Acc-7":"54.28","Corr":"0.779","MAE":"0.526"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-sentiment-analysis-on-cmu-mosi","task":"Multimodal Sentiment Analysis","dataset":"CMU-MOSI","model":"ALMT","rank_in_archive_order":10,"of":12,"metrics":{"Acc-5":"56.41","Acc-7":"49.42","Corr":"0.805","MAE":"0.683"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.05804","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}