{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mmlatch-bottom-up-top-down-fusion-for","title":"MMLatch: Bottom-up Top-down Fusion for Multimodal Sentiment Analysis","arxiv_id":"2201.09828","date":"2022-01-24","proceeding":null,"authors":["Georgios Paraskevopoulos","Efthymios Georgiou","Alexandros Potamianos"],"abstract":"Current deep learning approaches for multimodal fusion rely on bottom-up fusion of high and mid-level latent modality representations (late/mid fusion) or low level sensory inputs (early fusion). Models of human perception highlight the importance of top-down fusion, where high-level representations affect the way sensory inputs are perceived, i.e. cognition affects perception. These top-down interactions are not captured in current deep learning models. In this work we propose a neural architecture that captures top-down cross-modal interactions, using a feedback mechanism in the forward pass during network training. The proposed mechanism extracts high-level representations for each modality and uses these representations to mask the sensory inputs, allowing the model to perform top-down feature masking. We apply the proposed model for multimodal sentiment recognition on CMU-MOSEI. Our method shows consistent improvements over the well established MulT and over our strong late fusion baseline, achieving state-of-the-art results.","url_abs":"https://arxiv.org/abs/2201.09828v1","url_pdf":"https://arxiv.org/pdf/2201.09828v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mmlatch-bottom-up-top-down-fusion-for","repo_url":"https://github.com/georgepar/mmlatch","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"multimodal-sentiment-analysis","task_name":"Multimodal Sentiment Analysis"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multimodal-sentiment-analysis-on-cmu-mosei-1","task":"Multimodal Sentiment Analysis","dataset":"CMU-MOSEI","model":"MMLatch","rank_in_archive_order":7,"of":15,"metrics":{"Accuracy":"82.4","MAE":"0.7"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2201.09828","atlas_url":"https://app.syntology.ai/?focus=2201.09828","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}